Principal DevOps / Platform Engineer — Kubernetes, AKS and High Availability
Бюджэт: $35.0 - $90.0
HOURLY / FULL_TIME
⭐ 4.89 (294)
United States
kubernetes, docker, devops, cicd, automated-deployment, automation-software-release
Preferred qualifications
- Experience: Expert
Engagement: Full-time contract, 40 hours per week
Duration: Long-term — this is a build, improve and own role, not a short migration project
Location: Remote
Timezone: Must overlap at least four working hours with both US Eastern
Location preference: Candidates based in the EU/EEA with existing authorization to work or contract there are strongly preferred. We cannot provide visa sponsorship.
Start: Immediate
About EmpowerID
EmpowerID is an established enterprise identity security and Identity Governance and Administration company. We have approximately 120 employees and more than 20 years of experience serving complex enterprise customers.
We are building the next generation of the EmpowerID Identity Fabric: a modern microservices platform spanning identity governance, identity provider services, AuthZEN-based authorization, workflow orchestration, identity graph services, analytics and AI-agent governance.
Our application stack includes:
Python and FastAPI microservices
Traefik
PostgreSQL and pgvector
Redis/Valkey
Redpanda/Kafka
Neo4j
ClickHouse
OpenBao
S3-compatible object storage
Docker Compose
Kubernetes, primarily Azure AKS
Helm, Terraform/Bicep and GitOps delivery
Developers run the complete platform locally through a large Docker Compose environment. We already have Helm charts and Kubernetes/AKS deployment capabilities in varying stages of maturity. This is not a greenfield “convert a Compose file to Kubernetes” assignment.
We need a senior hands-on platform engineer to assess what exists, strengthen it, eliminate remaining single points of failure, mature our Helm and infrastructure automation, prove recoverability and high availability, and then take long-term operational ownership of the platform.
Our primary SaaS environment runs on Azure AKS. We also need a portable Kubernetes deployment profile for enterprise customers running RKE2-class or comparable on-premises infrastructure.
You will work directly with the CEO and the engineering leads. A detailed reference architecture and reviewed operationalization plan already exist. We expect you to execute them carefully, improve them when implementation evidence demands it and challenge assumptions that do not survive contact with production.
What you will own
1. Assess and harden the current platform
Your first responsibility will be to establish an evidence-based baseline of the current Docker, Helm, Kubernetes and AKS environments.
Initial work includes:
Review and rationalize existing Helm charts and deployment automation
Remove remaining environment-specific and filesystem coupling
Rotate credentials and remove hardcoded secrets
Eliminate publicly exposed database and administrative ports
Correct fail-open readiness and health checks
Split cache, session/security-state and coordination workloads where required
Pin production images by digest and establish controlled release manifests
Validate resource requests, limits, disruption budgets and topology placement
Implement and test backup and clean-environment restore procedures
Produce an explicit risk register instead of assuming that existing Kubernetes deployment equals high availability
2. Mature the AKS and portable Kubernetes platforms
You will refine and complete the platform foundation, including:
Terraform or Bicep for AKS, networking, private endpoints, Azure Workload Identity, Key Vault, Front Door and load balancers
Reusable Helm library patterns for web services, scalable workers, fenced singleton workers, background processors and migration Jobs
Per-service charts composed through Argo CD ApplicationSets or an equivalent GitOps model
Ordered database migrations and deployment dependencies without recreating Compose-style depends_on behavior
Signed, digest-pinned release manifests with provenance, SBOMs and vulnerability scanning
Equivalent portable deployment patterns for RKE2-class and customer-managed Kubernetes environments
3. Improve data-layer availability and recoverability
You will help operate and mature the following production data services:
PostgreSQL: Azure Database for PostgreSQL Flexible Server and CloudNativePG
Redis/Valkey: separate cache and security/coordination planes with appropriate eviction, persistence and HA behavior
Redpanda/Kafka: multi-broker operation, replication factors, topic partitioning, ordering, consumer lag and outbox-based recovery
OpenBao: Raft HA, external auto-unseal, Kubernetes authentication, policies, Transit and secret lifecycle management
Neo4j: clustered or explicitly recoverable graph operation
ClickHouse: replicated operation through the Altinity operator and Keeper
Object storage: durable application artifacts, backup data and recovery material
We do not expect one person to begin as the world’s leading expert in every datastore. We do expect deep production experience with several of them, strong distributed-systems judgment and the ability to become operationally competent with the remainder.
4. Establish secure Kubernetes defaults
You will make secure operation the default rather than something each service team must rediscover:
Restricted Pod Security Admission
Non-root and read-only container patterns
NetworkPolicy segmentation
Private endpoints for data services
Workload Identity instead of static cloud credentials
Isolated node pools for untrusted or higher-risk execution
Image-signature admission
Certificate lifecycle management
Least-privilege Kubernetes and Azure RBAC
Controlled break-glass procedures with audit evidence
5. Build observable, measurable reliability
You will implement and operate:
Prometheus and Grafana
Loki
Tempo
OpenTelemetry collectors
Service and dependency dashboards
SLO and error-budget definitions
Multi-window burn-rate alerting
Database replication and backup-lag alerts
Redpanda consumer-lag alerts
Worker lease-age and fencing alerts
OpenBao sealed-state alerts
Certificate-expiration and key-rotation alerts
Capacity forecasting and cost visibility
6. Certify failure behavior
High availability will be proven, not declared.
You will create and execute repeatable failure drills covering:
Pod, node and availability-zone loss
PostgreSQL and other state-store failover
Ambiguous network partitions
Worker termination and lease expiry under load
Rolling Kubernetes and application upgrades
Schema migration interruption
Key and certificate rotation during live traffic
Backup restoration into a clean environment
Kafka drain, replay and transactional-outbox recovery
Application reconnection and retry behavior during failover
The results will be measured against ratified RPO, RTO and service-level objectives. Where safety and availability conflict, the behavior and operational decision process must be explicit.
7. Own what you build
After the platform is hardened, this becomes an ongoing platform engineering and SRE role.
You will own:
Production operational readiness
On-call participation and incident response
Capacity planning
Platform and datastore upgrades
Backup verification and restore drills
Security and certificate rotation
Reliability reviews
Runbook maintenance
Root-cause analysis
Automation of recurring operational work
Quarterly or agreed failure-drill cadence
AI-native engineering is required
We expect you to use modern AI coding and reasoning tools fluently as part of your daily engineering workflow.
You should be comfortable using tools such as Codex, Claude Code, GitHub Copilot or equivalent systems to:
Inspect unfamiliar repositories and infrastructure
Draft and refactor Terraform, Helm and automation code
Generate test matrices and failure-injection tooling
Review manifests and configuration changes
Investigate incidents
Produce and maintain operational documentation
Accelerate repetitive platform work
“AI-native” does not mean deploying unreviewed generated infrastructure. You must be able to explain how you:
Constrain an AI agent’s access
Keep production credentials and customer data out of prompts
Review generated changes
Validate infrastructure plans before applying them
Test generated failure and recovery automation
Preserve an auditable human approval boundary
Required experience
You should have:
At least seven years in DevOps, platform engineering, SRE or infrastructure engineering
At least five years operating production Kubernetes
Hands-on responsibility for a substantial multi-service production platform
Experience hardening or migrating a real VM, Compose or early-stage Kubernetes platform into a multi-node or multi-zone production environment
Deep Azure and AKS experience, including private clusters, availability zones, Workload Identity, Key Vault and managed PostgreSQL
Experience with non-Azure Kubernetes such as RKE2, kubeadm or another customer-managed distribution
Strong Terraform or Bicep experience
Ability to author reusable Helm library charts, not merely install third-party charts
GitOps experience with Argo CD or Flux
Experience ordering migrations, Jobs and application rollouts safely
Strong Linux, networking, DNS, TLS, storage and Kubernetes troubleshooting skills
Production experience operating at least three of the major data services in our stack
Hands-on Vault or OpenBao experience, including HA, auto-unseal, Kubernetes authentication and policy design
Experience restoring production data into a clean environment
Experience with observability, SLOs, alerting and incident response
Strong written and spoken English
You must also be able to explain:
Fencing tokens versus advisory locks
Lease expiration under network partitions and process pauses
Transactional outboxes
Idempotency keys and unknown transaction outcomes
At-least-once delivery and idempotent consumers
Why “exactly once” should not be used casually
Why a scheduled backup is not evidence that recovery works
Why Kubernetes replica counts alone do not create high availability
Particularly valuable experience
The following would be helpful:
Python and FastAPI
Traefik v3
cert-manager
KEDA
Prometheus and Grafana rule authoring
Cosign, SBOM generation and admission controllers
Redpanda or Kafka operator experience
CloudNativePG
Altinity ClickHouse Operator and Keeper
Neo4j clustering
Redis or Valkey HA
Identity, OAuth, OpenID Connect or authorization systems
Enterprise security or regulated-customer environments
What success looks like
Within the first 30–45 days, we expect:
A verified inventory of the current deployment and its failure domains
Reproduction of the existing AKS and portable Kubernetes deployments
Closure or ownership plans for immediate security and recoverability risks
Tested backup and restore evidence for the most critical state stores
A prioritized platform backlog tied to measurable operational outcomes
Within approximately six months, we expect:
Repeatable GitOps deployment into clean environments
Hardened and reusable Helm patterns
Clear AKS and portable Kubernetes profiles
Protected roots of trust
Measured state-store recovery behavior
Safe singleton-worker execution
Observable platform SLOs
Scripted failure and restoration drills
Runbooks that another qualified engineer can execute
Longer term, success means the platform remains secure, recoverable and operable without relying on tribal knowledge.
Selection process
The process consists of:
A technical interview with the CEO and platform leadership.
A paid, time-boxed pilot using an isolated copy of the platform.
Review of the implementation, evidence, documentation and reasoning—not merely whether the final command succeeded.
We value engineers who are precise, verify their own work, communicate clearly and are willing to say:
This part of the plan is wrong. Here is the evidence, the risk and the safer implementation.
To apply
Applications that do not answer the following questions will not be considered:
Describe a substantial Compose-, VM- or early-Kubernetes-to-production-Kubernetes project you led. How large was the platform, what failed or surprised you, and what would you do differently?
Describe how you would run Vault or OpenBao in HA without distributing static unseal material. How would workloads authenticate, and how would you recover the root of trust?
A worker writes to a database and must not execute concurrently with another instance. Describe your lease and fencing design. What happens during a network partition, long process pause, Kubernetes reschedule or database failover?
Which of PostgreSQL HA, Redis/Valkey, Kafka/Redpanda, OpenBao/Vault, Neo4j and ClickHouse have you operated in production? Describe the scale and your personal responsibility.
Give one concrete example of using an AI coding agent for infrastructure or production operations. What did it produce, how did you verify it and how did you prevent unsafe access or deployment?
Where are you located? Do you already have authorization to work or provide contracting services in the EU/EEA?
State your timezone, normal working hours, weekly availability, earliest start date and expected hourly or monthly rate.
Адкрыць заказ
AI-чарнавік адказу
Згенеруйце кароткі cover letter па гэтай вакансіі. Перад адпраўкай адрэдагуйце.
Увайдзіце, каб згенерыраваць AI-чарнавік.
Увайсці