Senior DevOps Engineer for Voice AI and AI Infrastructure
Budget: $20.0 - $30.0
HOURLY / PART_TIME
⭐ 5.00 (1)
USA
windows-azure, amazon-web-services, devops, kubernetes, docker, grafana, google-cloud-platform
We are looking for a highly capable DevOps or Platform Engineer based in Latin America to help us manage, scale, and improve the infrastructure behind our production Voice AI platform.
This is not a traditional CI/CD-focused DevOps role. We need someone with strong experience in cloud infrastructure, Kubernetes, real-time systems, networking, observability, and AI workloads.
Our platform handles real-time voice conversations and includes components such as SIP servers, LiveKit, Asterisk, WebSockets, AI model services, GPU infrastructure, Kubernetes, and cloud-native services.
Responsibilities
Review and improve our existing cloud and Kubernetes architecture
Manage production deployments across AWS and Azure
Improve reliability, scalability, and fault tolerance
Troubleshoot SIP, RTP, WebSocket, networking, and real-time media issues
Support Asterisk and LiveKit-based voice infrastructure
Set up and improve monitoring, alerting, logging, and incident response
Optimize infrastructure for high-concurrency voice workloads
Manage GPU-based AI inference workloads
Improve autoscaling, resource allocation, and cost efficiency
Review security, networking, secrets management, and access controls
Investigate production incidents and perform root-cause analysis
Automate deployments using Terraform, Helm, GitHub Actions, or similar tools
Work closely with our AI, backend, and voice engineering teams
Required Experience
Strong hands-on experience with Kubernetes, Docker, Helm, and Linux
Strong experience with AWS, Azure, or both
Experience with Terraform or other infrastructure-as-code tools
Experience with Prometheus, Grafana, Loki, OpenTelemetry, or similar tools
Strong understanding of networking, load balancing, DNS, TLS, and firewalls
Experience managing production systems with high concurrency
Strong troubleshooting and root-cause analysis skills
Experience with Python, Bash, Go, or another scripting language
Good written and spoken English
High availability and fast communication during critical incidents
Strongly Preferred
Experience with Voice AI, conversational AI, or real-time communication platforms
Experience with LiveKit, Asterisk, Kamailio, OpenSIPS, FreeSWITCH, or similar systems
Understanding of SIP, RTP, WebRTC, STUN, TURN, and WebSockets
Experience with GPU workloads, NVIDIA container runtime, and AI inference services
Experience scaling LLM, STT, TTS, or machine-learning infrastructure
Experience working with distributed systems and low-latency applications
Familiarity with Redis, PostgreSQL, Kafka, NATS, or similar technologies
Experience with on-premise or private-cloud deployments
Engagement
Freelance or long-term contract
Initially milestone-based
Potential for an ongoing engagement
Preference for candidates based in Mexico, Colombia, Argentina, Brazil, Chile, or other LATAM countries
Must overlap with US working hours
Must be responsive and available during agreed production support windows
Application Requirements
Please include:
Your location and time zone
Your availability per week
Your hourly or monthly rate
Examples of production infrastructure you have managed
Your experience with Kubernetes and high-concurrency systems
Any experience with SIP, WebRTC, LiveKit, Asterisk, Voice AI, or AI infrastructure
A brief description of the most difficult production incident you have resolved
Whether you are open to completing a paid technical assessment
Please do not apply if your experience is primarily limited to basic CI/CD pipelines, website hosting, or standard application deployments. We are looking for someone who can independently investigate complex production infrastructure and real-time communication issues.
Apri su Upwork