← Вакансіі

Principal DevOps / Platform Engineer — Kubernetes, AKS and High Availability

Бюджэт: $35.0 - $90.0 HOURLY / FULL_TIME ⭐ 4.89 (294) United States

kubernetes, docker, devops, cicd, automated-deployment, automation-software-release

Preferred qualifications

  • Experience: Expert
Engagement: Full-time contract, 40 hours per week Duration: Long-term — this is a build, improve and own role, not a short migration project Location: Remote Timezone: Must overlap at least four working hours with both US Eastern Location preference: Candidates based in the EU/EEA with existing authorization to work or contract there are strongly preferred. We cannot provide visa sponsorship. Start: Immediate About EmpowerID EmpowerID is an established enterprise identity security and Identity Governance and Administration company. We have approximately 120 employees and more than 20 years of experience serving complex enterprise customers. We are building the next generation of the EmpowerID Identity Fabric: a modern microservices platform spanning identity governance, identity provider services, AuthZEN-based authorization, workflow orchestration, identity graph services, analytics and AI-agent governance. Our application stack includes: Python and FastAPI microservices Traefik PostgreSQL and pgvector Redis/Valkey Redpanda/Kafka Neo4j ClickHouse OpenBao S3-compatible object storage Docker Compose Kubernetes, primarily Azure AKS Helm, Terraform/Bicep and GitOps delivery Developers run the complete platform locally through a large Docker Compose environment. We already have Helm charts and Kubernetes/AKS deployment capabilities in varying stages of maturity. This is not a greenfield “convert a Compose file to Kubernetes” assignment. We need a senior hands-on platform engineer to assess what exists, strengthen it, eliminate remaining single points of failure, mature our Helm and infrastructure automation, prove recoverability and high availability, and then take long-term operational ownership of the platform. Our primary SaaS environment runs on Azure AKS. We also need a portable Kubernetes deployment profile for enterprise customers running RKE2-class or comparable on-premises infrastructure. You will work directly with the CEO and the engineering leads. A detailed reference architecture and reviewed operationalization plan already exist. We expect you to execute them carefully, improve them when implementation evidence demands it and challenge assumptions that do not survive contact with production. What you will own 1. Assess and harden the current platform Your first responsibility will be to establish an evidence-based baseline of the current Docker, Helm, Kubernetes and AKS environments. Initial work includes: Review and rationalize existing Helm charts and deployment automation Remove remaining environment-specific and filesystem coupling Rotate credentials and remove hardcoded secrets Eliminate publicly exposed database and administrative ports Correct fail-open readiness and health checks Split cache, session/security-state and coordination workloads where required Pin production images by digest and establish controlled release manifests Validate resource requests, limits, disruption budgets and topology placement Implement and test backup and clean-environment restore procedures Produce an explicit risk register instead of assuming that existing Kubernetes deployment equals high availability 2. Mature the AKS and portable Kubernetes platforms You will refine and complete the platform foundation, including: Terraform or Bicep for AKS, networking, private endpoints, Azure Workload Identity, Key Vault, Front Door and load balancers Reusable Helm library patterns for web services, scalable workers, fenced singleton workers, background processors and migration Jobs Per-service charts composed through Argo CD ApplicationSets or an equivalent GitOps model Ordered database migrations and deployment dependencies without recreating Compose-style depends_on behavior Signed, digest-pinned release manifests with provenance, SBOMs and vulnerability scanning Equivalent portable deployment patterns for RKE2-class and customer-managed Kubernetes environments 3. Improve data-layer availability and recoverability You will help operate and mature the following production data services: PostgreSQL: Azure Database for PostgreSQL Flexible Server and CloudNativePG Redis/Valkey: separate cache and security/coordination planes with appropriate eviction, persistence and HA behavior Redpanda/Kafka: multi-broker operation, replication factors, topic partitioning, ordering, consumer lag and outbox-based recovery OpenBao: Raft HA, external auto-unseal, Kubernetes authentication, policies, Transit and secret lifecycle management Neo4j: clustered or explicitly recoverable graph operation ClickHouse: replicated operation through the Altinity operator and Keeper Object storage: durable application artifacts, backup data and recovery material We do not expect one person to begin as the world’s leading expert in every datastore. We do expect deep production experience with several of them, strong distributed-systems judgment and the ability to become operationally competent with the remainder. 4. Establish secure Kubernetes defaults You will make secure operation the default rather than something each service team must rediscover: Restricted Pod Security Admission Non-root and read-only container patterns NetworkPolicy segmentation Private endpoints for data services Workload Identity instead of static cloud credentials Isolated node pools for untrusted or higher-risk execution Image-signature admission Certificate lifecycle management Least-privilege Kubernetes and Azure RBAC Controlled break-glass procedures with audit evidence 5. Build observable, measurable reliability You will implement and operate: Prometheus and Grafana Loki Tempo OpenTelemetry collectors Service and dependency dashboards SLO and error-budget definitions Multi-window burn-rate alerting Database replication and backup-lag alerts Redpanda consumer-lag alerts Worker lease-age and fencing alerts OpenBao sealed-state alerts Certificate-expiration and key-rotation alerts Capacity forecasting and cost visibility 6. Certify failure behavior High availability will be proven, not declared. You will create and execute repeatable failure drills covering: Pod, node and availability-zone loss PostgreSQL and other state-store failover Ambiguous network partitions Worker termination and lease expiry under load Rolling Kubernetes and application upgrades Schema migration interruption Key and certificate rotation during live traffic Backup restoration into a clean environment Kafka drain, replay and transactional-outbox recovery Application reconnection and retry behavior during failover The results will be measured against ratified RPO, RTO and service-level objectives. Where safety and availability conflict, the behavior and operational decision process must be explicit. 7. Own what you build After the platform is hardened, this becomes an ongoing platform engineering and SRE role. You will own: Production operational readiness On-call participation and incident response Capacity planning Platform and datastore upgrades Backup verification and restore drills Security and certificate rotation Reliability reviews Runbook maintenance Root-cause analysis Automation of recurring operational work Quarterly or agreed failure-drill cadence AI-native engineering is required We expect you to use modern AI coding and reasoning tools fluently as part of your daily engineering workflow. You should be comfortable using tools such as Codex, Claude Code, GitHub Copilot or equivalent systems to: Inspect unfamiliar repositories and infrastructure Draft and refactor Terraform, Helm and automation code Generate test matrices and failure-injection tooling Review manifests and configuration changes Investigate incidents Produce and maintain operational documentation Accelerate repetitive platform work “AI-native” does not mean deploying unreviewed generated infrastructure. You must be able to explain how you: Constrain an AI agent’s access Keep production credentials and customer data out of prompts Review generated changes Validate infrastructure plans before applying them Test generated failure and recovery automation Preserve an auditable human approval boundary Required experience You should have: At least seven years in DevOps, platform engineering, SRE or infrastructure engineering At least five years operating production Kubernetes Hands-on responsibility for a substantial multi-service production platform Experience hardening or migrating a real VM, Compose or early-stage Kubernetes platform into a multi-node or multi-zone production environment Deep Azure and AKS experience, including private clusters, availability zones, Workload Identity, Key Vault and managed PostgreSQL Experience with non-Azure Kubernetes such as RKE2, kubeadm or another customer-managed distribution Strong Terraform or Bicep experience Ability to author reusable Helm library charts, not merely install third-party charts GitOps experience with Argo CD or Flux Experience ordering migrations, Jobs and application rollouts safely Strong Linux, networking, DNS, TLS, storage and Kubernetes troubleshooting skills Production experience operating at least three of the major data services in our stack Hands-on Vault or OpenBao experience, including HA, auto-unseal, Kubernetes authentication and policy design Experience restoring production data into a clean environment Experience with observability, SLOs, alerting and incident response Strong written and spoken English You must also be able to explain: Fencing tokens versus advisory locks Lease expiration under network partitions and process pauses Transactional outboxes Idempotency keys and unknown transaction outcomes At-least-once delivery and idempotent consumers Why “exactly once” should not be used casually Why a scheduled backup is not evidence that recovery works Why Kubernetes replica counts alone do not create high availability Particularly valuable experience The following would be helpful: Python and FastAPI Traefik v3 cert-manager KEDA Prometheus and Grafana rule authoring Cosign, SBOM generation and admission controllers Redpanda or Kafka operator experience CloudNativePG Altinity ClickHouse Operator and Keeper Neo4j clustering Redis or Valkey HA Identity, OAuth, OpenID Connect or authorization systems Enterprise security or regulated-customer environments What success looks like Within the first 30–45 days, we expect: A verified inventory of the current deployment and its failure domains Reproduction of the existing AKS and portable Kubernetes deployments Closure or ownership plans for immediate security and recoverability risks Tested backup and restore evidence for the most critical state stores A prioritized platform backlog tied to measurable operational outcomes Within approximately six months, we expect: Repeatable GitOps deployment into clean environments Hardened and reusable Helm patterns Clear AKS and portable Kubernetes profiles Protected roots of trust Measured state-store recovery behavior Safe singleton-worker execution Observable platform SLOs Scripted failure and restoration drills Runbooks that another qualified engineer can execute Longer term, success means the platform remains secure, recoverable and operable without relying on tribal knowledge. Selection process The process consists of: A technical interview with the CEO and platform leadership. A paid, time-boxed pilot using an isolated copy of the platform. Review of the implementation, evidence, documentation and reasoning—not merely whether the final command succeeded. We value engineers who are precise, verify their own work, communicate clearly and are willing to say: This part of the plan is wrong. Here is the evidence, the risk and the safer implementation. To apply Applications that do not answer the following questions will not be considered: Describe a substantial Compose-, VM- or early-Kubernetes-to-production-Kubernetes project you led. How large was the platform, what failed or surprised you, and what would you do differently? Describe how you would run Vault or OpenBao in HA without distributing static unseal material. How would workloads authenticate, and how would you recover the root of trust? A worker writes to a database and must not execute concurrently with another instance. Describe your lease and fencing design. What happens during a network partition, long process pause, Kubernetes reschedule or database failover? Which of PostgreSQL HA, Redis/Valkey, Kafka/Redpanda, OpenBao/Vault, Neo4j and ClickHouse have you operated in production? Describe the scale and your personal responsibility. Give one concrete example of using an AI coding agent for infrastructure or production operations. What did it produce, how did you verify it and how did you prevent unsafe access or deployment? Where are you located? Do you already have authorization to work or provide contracting services in the EU/EEA? State your timezone, normal working hours, weekly availability, earliest start date and expected hourly or monthly rate.
Адкрыць заказ

AI-чарнавік адказу

Згенеруйце кароткі cover letter па гэтай вакансіі. Перад адпраўкай адрэдагуйце.

Увайдзіце, каб згенерыраваць AI-чарнавік.

Увайсці