MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps
Budget: $10.0 - $25.0
HOURLY / FULL_TIME
⭐ 4.81 (12)
United States
terraform, bash, kubernetes, containerization, amazon-web-services, google-cloud-platform, python, grafana
MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps
Summary
We are working on a proprietary AI/ML infrastructure engagement and are looking for an experienced SRE Architect to support our small technical team and share the hands-on engineering workload.
The environment includes 50+ Kubernetes clusters across AWS, GCP, on-premises infrastructure, and other cloud providers, supporting production GPU and AI/ML inference workloads.
The work will involve troubleshooting multi-cluster Kubernetes environments, managing node lifecycle activities, developing StackStorm auto-remediation workflows, maintaining Terraform and Flux CD configurations, debugging container runtimes, improving observability, and supporting structured incident response.
This is a hands-on role for someone who can investigate complex infrastructure problems, write automation, implement safe solutions, and clearly document technical findings.
Our Tech Stack
Kubernetes & Fleet Management: Kubernetes, Rancher, multi-cluster operations, node cordon/drain/reconfiguration
Automation: StackStorm/ST2, Python, Bash, event-driven auto-remediation
Infrastructure & GitOps: Terraform, Flux CD, Helm, Git-based infrastructure workflows
Cloud: AWS, GCP, on-premises Kubernetes; neocloud experience is a plus
Container Runtime: containerd, stargz, image caching, snapshotters, cgroups
GPU & AI Infrastructure: NVIDIA GPU Operator, DCGM, GPU workloads, AI/ML inference and model-serving platforms
Observability: Grafana, VictoriaMetrics, Prometheus-compatible alerting, observability-as-code, runbooks
Reliability: Service catalogs, SLI/SLO implementation, incident.io or similar incident-management platforms
Requirements
● Proven experience as a senior SRE, SRE Architect, Platform Engineer, or MLOps Infrastructure Engineer in large-scale production environments.
● Deep Kubernetes troubleshooting experience across multiple clusters, cloud providers, and on-premises environments.
● Strong experience managing Kubernetes node lifecycle activities, including cordoning, draining, reconfiguration, recovery, and safe workload rescheduling.
● Hands-on StackStorm/ST2 experience for operational automation and auto-remediation.
● Strong Terraform, Flux CD, Rancher, Python, and Bash experience.
● Experience debugging containerd, image-cache, stargz, snapshotter, and cgroup-related issues.
● Experience operating GPU workloads using NVIDIA GPU Operator and DCGM.
● Experience supporting AI/ML inference infrastructure or model-serving platforms.
● Ability to create Grafana dashboards, alerting rules, runbooks, and observability configurations, not only monitor existing dashboards.
● Structured on-call and incident-response experience, including root-cause analysis, remediation tracking, and postmortems.
● This role is not suitable for general cloud engineers without deep Kubernetes experience or engineers who rely primarily on manual operations without scripting and automation.
● Please include brief examples of your experience with large Kubernetes fleets, StackStorm, GPU infrastructure, containerd debugging, and Terraform/Flux CD when applying
Auf Upwork öffnen