← Lavori

MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps

Budget: $10.0 - $25.0 HOURLY / FULL_TIME ⭐ 4.81 (12) United States

terraform, bash, kubernetes, containerization, amazon-web-services, google-cloud-platform, python, grafana

MLOps/SRE Architect – Multi-Cloud Kubernetes, StackStorm, GPU Infrastructure & GitOps Summary We are working on a proprietary AI/ML infrastructure engagement and are looking for an experienced SRE Architect to support our small technical team and share the hands-on engineering workload. The environment includes 50+ Kubernetes clusters across AWS, GCP, on-premises infrastructure, and other cloud providers, supporting production GPU and AI/ML inference workloads. The work will involve troubleshooting multi-cluster Kubernetes environments, managing node lifecycle activities, developing StackStorm auto-remediation workflows, maintaining Terraform and Flux CD configurations, debugging container runtimes, improving observability, and supporting structured incident response. This is a hands-on role for someone who can investigate complex infrastructure problems, write automation, implement safe solutions, and clearly document technical findings. Our Tech Stack Kubernetes & Fleet Management: Kubernetes, Rancher, multi-cluster operations, node cordon/drain/reconfiguration Automation: StackStorm/ST2, Python, Bash, event-driven auto-remediation Infrastructure & GitOps: Terraform, Flux CD, Helm, Git-based infrastructure workflows Cloud: AWS, GCP, on-premises Kubernetes; neocloud experience is a plus Container Runtime: containerd, stargz, image caching, snapshotters, cgroups GPU & AI Infrastructure: NVIDIA GPU Operator, DCGM, GPU workloads, AI/ML inference and model-serving platforms Observability: Grafana, VictoriaMetrics, Prometheus-compatible alerting, observability-as-code, runbooks Reliability: Service catalogs, SLI/SLO implementation, incident.io or similar incident-management platforms Requirements ● Proven experience as a senior SRE, SRE Architect, Platform Engineer, or MLOps Infrastructure Engineer in large-scale production environments. ● Deep Kubernetes troubleshooting experience across multiple clusters, cloud providers, and on-premises environments. ● Strong experience managing Kubernetes node lifecycle activities, including cordoning, draining, reconfiguration, recovery, and safe workload rescheduling. ● Hands-on StackStorm/ST2 experience for operational automation and auto-remediation. ● Strong Terraform, Flux CD, Rancher, Python, and Bash experience. ● Experience debugging containerd, image-cache, stargz, snapshotter, and cgroup-related issues. ● Experience operating GPU workloads using NVIDIA GPU Operator and DCGM. ● Experience supporting AI/ML inference infrastructure or model-serving platforms. ● Ability to create Grafana dashboards, alerting rules, runbooks, and observability configurations, not only monitor existing dashboards. ● Structured on-call and incident-response experience, including root-cause analysis, remediation tracking, and postmortems. ● This role is not suitable for general cloud engineers without deep Kubernetes experience or engineers who rely primarily on manual operations without scripting and automation. ● Please include brief examples of your experience with large Kubernetes fleets, StackStorm, GPU infrastructure, containerd debugging, and Terraform/Flux CD when applying
Apri su Upwork