Kbeadm with GPU
Budget: -
HOURLY / PART_TIME
⭐ 0.00 (0)
Pakistan
kubernetes, terraform, amazon-web-services, azure-devops, devops, cicd, infrastructure-as-code, docker, amazon-ecs-for-kubernetes, cloud-migration, python, ansible, solution-architecture, distributed-computing, system-administration
Bevorzugte Qualifikationen
- Erfahrung: Fortgeschritten
The platform is built on a cluster-of-clusters Kubernetes architecture: a central management cluster provisions and reconciles child GPU clusters across multiple cloud providers. The Senior Platform Engineer builds and operates this management-to-child relationship through Terraform modules, Ansible playbooks, Kubernetes operators, and Go provider adapters that expose a uniform gRPC surface regardless of the underlying network fabric.
Three engineering principles govern everything this role builds:
Operator-first lifecycle management: Every stateful platform component — schedulers, GPU device plugins, cluster controllers, observability stack, container registry — is delivered as a Kubernetes operator backed by a CRD. No component manages its own lifecycle through scripts or manual operations.
Observability by default: OpenTelemetry instrumentation is embedded from the first line of code. NVLink saturation, ECC error rates, MIG utilisation, and provider API latencies must be visible in Prometheus and Grafana before any workload runs. Observability is not added retrospectively.
Infrastructure-as-code purity: All cloud resources — management plane, compute nodes, networking, storage — are expressed in Terraform. No manual cloud console operations. State is managed remotely; modules are reviewed by peers before apply.
________________________________________
Key Responsibilities and Deliverables
- Provision and operate the cloud management plane cluster on a confirmed stable Kubernetes version with documented upgrade criteria
- Author and maintain kubeadm/Ansible compute plane bootstrap: playbooks, inventory structure, GPU driver version pinning, and hardening templates
- Implement the platform cluster operator: controller reconciliation loop, CRD schema integration, and admission webhook for configuration validation
- Build the primary cloud provider adapter: EFA placement group configuration, shared filesystem CSI integration, and SLURM node registration
- Build the second cloud provider adapter: InfiniBand fabric configuration, kubeadm on provider VMs, distributed filesystem mount, and SLURM node registration
- Deploy and configure the SLURM operator via Helm; implement the cluster configuration-to-SLURM translation layer
- Implement MIG profile management integrated with the SLURM operator, with per-profile utilisation metrics exposed via OTel
- Deploy the OpenTelemetry Operator and configure collectors for per-node GPU counter collection across all providers
- Establish the mono-repo structure with CI/CD pipeline, task runner configuration, and component-level change detection rules
- Validate blueprint parity across both cloud providers before the cross-provider migration demo
- Harden both provider adapters for production reliability under design partner conditions
Auf Upwork öffnen
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Anmelden