Senior AI Voice Engineer
Budget: $30.0 - $59.0
HOURLY / FULL_TIME
⭐ 0.00 (0)
Pakistan
artificial-intelligence, machine-learning, python, natural-language-processing, voice-over
Qualifiche preferite
- Esperienza: Esperto
We are looking for a Senior AI Voice Engineer with strong hands-on experience building real-time conversational voice systems and telephony infrastructure.
You will be responsible for developing a standalone AI voice-agent service that integrates with an existing SaaS platform and can autonomously make and manage real-world phone calls.
This is not a simple chatbot, LLM integration, or prompt-engineering project. The focus is on building a reliable, low-latency voice system capable of handling interruptions, IVRs, DTMF, voicemail, transfers, silence, external API calls, and other unpredictable telephony conditions.
Technology Stack
The expected technology stack includes:
Python
Pipecat
FastAPI
Twilio Voice / Twilio Media Streams
OpenAI Realtime API and/or OpenAI models
Speech-to-text and text-to-speech providers
WebSockets / REST APIs
Docker
Async Python
We are particularly interested in engineers who understand real-time audio pipelines and telephone systems, rather than developers whose experience is limited to integrating LLMs into conventional web applications.
What You Will Build
You will own the development of a standalone voice-agent service that communicates with our existing application through a well-defined API.
The system will need to:
Initiate and manage outbound AI-powered phone calls through Twilio
Maintain low-latency, real-time voice conversations
Handle natural turn-taking and interruptions
Detect and respond to barge-ins
Navigate IVR systems and send DTMF tones
Handle ringing, silence, hold music, voicemail, transfers, and unexpected disconnects
Invoke external application functions while a call is in progress
Collect and validate structured information from conversations
Support configurable prompts, voices, languages, and agent behaviors
Generate transcripts and structured call outcomes
Report call metrics, events, errors, and other operational data
Provide a reusable architecture for future voice-agent workflows
The system must be designed for production reliability, not simply to demonstrate a successful proof of concept.
What We Provide
For shortlisted candidates, we will provide:
Detailed technical specifications
Functional requirements and acceptance criteria
API documentation
Mock APIs and development servers
Synthetic test data
Agent prompts and expected outputs
Required third-party development credentials
Engineering support for integration-related questions
Development will take place in a standalone repository. Access to our production application or customer data will not be required.
Expected Deliverables
The project will include:
Production-ready Python voice-agent service
Modular and configurable agent architecture
Twilio telephony integration
Real-time audio pipeline
Dockerized development and deployment environment
Automated unit and integration tests
End-to-end voice-call testing
Error handling, retries, logging, and monitoring
Technical and deployment documentation
Final engineering handoff
The work will be organized into clearly defined milestones.
Ideal Candidate
You should have substantial experience with several of the following:
Pipecat or similar real-time voice frameworks
Twilio Voice and Media Streams
WebSocket-based audio streaming
OpenAI Realtime API
Deepgram, Cartesia, ElevenLabs, or comparable STT/TTS services
Voice activity detection
Barge-in and interruption handling
IVR and DTMF automation
Async Python
FastAPI
Function/tool calling
Real-time distributed systems
Production observability and monitoring
Retry and failure-recovery strategies
Docker and cloud deployment
Hands-on experience building production telephone agents is highly preferred.
General LLM, chatbot, RAG, or prompt-engineering experience alone is not sufficient for this role.
How to Apply
Please do not submit a generic or AI-generated proposal.
Instead, describe the most technically challenging real-time voice or telephony system you have personally built.
Please include:
What the system did
Your specific responsibilities
The technologies and services you used
How you handled real-time audio and interruptions
How you dealt with telephony failures or unusual call states
Any latency, scalability, or reliability challenges you encountered
If your experience is primarily with managed platforms such as Vapi, Retell, Bland, Synthflow, or similar services, please be transparent about that. Such experience is valuable, but this project requires building and controlling a significant portion of the underlying voice infrastructure directly.
Shortlisted candidates will receive the complete technical specification before the final scope, milestones, and pricing are agreed upon.
Apri su Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Accedi