← Trabalhos

Senior Real-Time Voice AI Engineer

Orçamento: $30.0 - $60.0 HOURLY / FULL_TIME ⭐ 4.27 (10) United States

python, automatic-speech-recognition, artificial-intelligence, natural-language-processing, machine-learning, node.js, typescript, api-integration, webrtc, postgresql-programming, websockets

Qualificações preferidas

  • Experiência: Especialista
Epsom Technologies | Voca AI Remote | Expert Level | Initial 8–12 Week Contract | Possible Extension About Epsom Technologies Epsom Technologies is building an integrated education-to-career technology platform serving students, parents/sponsors, advisors, coaches, recruiters, and administrators. One of our core products is Voca AI, designed to become the conversational intelligence layer of the Epsom ecosystem. Voca is not intended to be a simple chatbot or standalone voice bot. It combines conversational intake, structured student-profile creation, document interaction, personalized recommendations, internal-data retrieval, portal routing, human review, and role-based access across the wider Epsom platform. We already have an existing platform, application code, deployment environment, and partial Voca implementation. This is not a greenfield project. The Mission We are looking for a senior hands-on Real-Time Voice AI Engineer who has personally built, debugged, and shipped production conversational voice systems. Your first responsibility will not be to add features or rewrite the system. Your first responsibility will be to inspect the existing Voca implementation, establish technical ground truth, and determine what should be: KEEP · COMPLETE · FIX · REFACTOR · REPLACE · POSTPONE · REMOVE We specifically want an engineer who can inherit an imperfect existing system, understand what is actually happening in production, identify the highest-risk problems, and move it toward a reliable and maintainable production state. About the System You Will Inherit Voca's conversational layer is intended to provide a unified experience across: voice; text chat; live transcription; document interaction; structured data extraction; persistent student/session state; internal Epsom data retrieval; portal context; user-specific recommendations; authorized actions; human review. Voice, text, and documents must contribute to the same conversation state and Unified Student Profile, rather than functioning as disconnected products. We are especially interested in engineers who have embedded real-time voice into stateful production SaaS applications, rather than only building isolated telephone agents, demo bots, or simple STT → LLM → TTS pipelines. Immediate Responsibilities The selected engineer will begin by determining: What voice functionality is already implemented? What actually works reliably? What works only in demonstrations? How audio moves through the current system. How interim and final transcripts are handled. How conversation/session state is stored. How reconnect and resume behavior works. How the AI layer interacts with application data. How structured profile fields are updated. How authentication and authorization are enforced. How voice failures are currently logged and diagnosed. Where latency is introduced. Where data, state, security, or reliability problems exist. Which components should remain and which should change. You will then produce and help execute a prioritized production-hardening plan. Phase 1 Voice Experience The current Voca specification requires a live conversational experience in which: The user starts the microphone/session once. The user can continue speaking without clicking for every response. Live interim transcription appears while the user speaks. The final transcript is saved into the conversation. Voca responds in both voice and text. The user can continue naturally after Voca responds. Conversation state persists. Voca can route the user to the appropriate next portal/action. Voca can retrieve relevant internal information. Sensitive actions require appropriate confirmation. The system should also support natural turn-taking, including recognizing when the user begins speaking, pauses, or continues a multi-sentence response without being cut off prematurely. Session Persistence Live voice must connect to persistent application state. Relevant state may include: session ID; conversation history; interim and finalized transcripts; Voca responses; extracted structured fields; uploaded documents; completed intake steps; incomplete fields; current recommended action; resume-later state. A user who leaves and returns should be able to continue meaningfully rather than starting over. Reliability and Failure Handling The system must degrade gracefully when problems occur. Examples include: microphone permission denied; poor audio quality; low-confidence transcription; internet interruption; session reconnection; premature turn-ending; background noise; language changes; provider/API failures; lost or inconsistent application state. The current product requirements explicitly call for progress preservation and session recovery when connectivity is interrupted. We are looking for someone who treats these as production engineering problems, not simply API configuration tasks. Security and Authorization Voca may interact with sensitive student information, educational records, financial information, identity documents, visa-related information, and internal Epsom workflows. The system therefore requires strong application-level security. Relevant requirements include: role-based access control; restricted visibility of sensitive information; row-level security where appropriate; signed/expiring file access; access logging; audit history; secure document handling; consent tracking; restricted recruiter access; human review and confirmation for sensitive actions. Authorization must be enforced by the application/backend. An LLM prompt must never be the security boundary. Internal Data and Tool Use Voca should retrieve authorized information from Epsom systems rather than relying only on general AI knowledge. Relevant information may include: student profiles; intake answers; documents; university/program information; admission rules; service/package rules; document checklists; coaching resources; recruiter/job information; payment/package status; parent/sponsor action items; advisor information where authorized. When verified information is unavailable or outdated, Voca should not invent an answer. Multimodal and Product Integration Voca is also intended to operate as a floating assistant across the Epsom web and mobile ecosystem. Depending on permissions and implementation stage, it may interact with: the user's current page or portal; PDFs and Word documents; images/screenshots; links; uploaded records; screen sharing; page context; structured application data. Screen and document access must be explicit and permission-controlled. Experience building voice AI inside an authenticated, stateful, multimodal application is therefore highly valuable. Phase 2 Expertise — Strongly Preferred The following are not all required for the first production milestone, but strong candidates should understand them: interruption/barge-in; advanced turn-taking; Voice Activity Detection; low-latency streaming; multi-speaker conversations; multilingual voice interaction; real-time translation; provider failover; advanced observability; concurrency; latency optimization; voice-session cost optimization. The current product specification identifies advanced interruption handling, improved latency/turn-taking, multi-speaker behavior, and real-time multilingual conversations as later enhancements. Relevant Technical Experience Strong candidates may have experience with combinations of: WebRTC; WebSockets; streaming audio; Voice Activity Detection; streaming speech-to-text; streaming text-to-speech; OpenAI Realtime; Deepgram; ElevenLabs; LiveKit; Vapi; Twilio; AssemblyAI; OpenAI / Claude / Gemini APIs; structured tool calling; Node.js; TypeScript; Python; FastAPI; PostgreSQL; Supabase; Redis or equivalent state/session infrastructure; secure API design; real-time frontend applications; cloud deployment; observability and distributed tracing; CI/CD; production incident response. Experience with any particular vendor is not sufficient by itself. We care more about whether you understand the architecture underneath the vendor. What We Are NOT Looking For This is probably not the right role if your experience is primarily: prompt engineering; n8n or Make automation; basic chatbot development; text-only RAG; batch speech transcription; simple voice demos; connecting STT → LLM → TTS without owning the surrounding production system; configuring a commercial voice platform without understanding streaming behavior; architecture consulting without hands-on coding. Evidence Matters Candidates will be evaluated based on what they personally designed, coded, debugged, deployed, and operated. Do not present an employer's, agency's, or team's accomplishments as your own without clearly explaining your individual contribution. Where confidentiality permits, supporting evidence may include: public production products; personal GitHub repositories; sanitized architecture diagrams; sanitized screenshots; technical documentation; authorized code samples; pull requests; monitoring/deployment evidence; references; production metrics. We do not require confidential client source code or proprietary information. Engagement Initial engagement: approximately 8–12 weeks Potential extension: Yes Commitment: Approximately 20–40 hours/week depending on the initial technical assessment Location: Remote / Worldwide Experience level: Expert We prefer a hands-on individual engineer. If you are applying through an agency or team, disclose that clearly and identify the person who will personally perform the work. APPLICATION INSTRUCTIONS Start your proposal with exactly: VOICE PRODUCTION Then answer all screening questions below. Generic proposals or proposals that do not answer the questions may not be reviewed. Use of AI Tools You may use AI tools for spelling, grammar, formatting, or organizing your response. However, the technical experience, architecture decisions, production incidents, metrics, examples, and accomplishments you describe must reflect your actual experience. If generative AI materially assisted in drafting your answers, disclose that briefly at the end. Every substantive technical claim may be selected for an unannounced live follow-up question. You will be expected to explain and defend the technical details during a live interview without AI assistance. We are not evaluating polished writing. We are evaluating: authenticity · engineering reasoning · technical depth · personal contribution · production evidence Screening Questions 1. Production Voice System Describe the strongest real-time conversational voice system you personally helped ship to production. Include: what the product did; approximately when it operated; who actually used it; architecture; audio transport; STT technology/provider; LLM/orchestration layer; TTS technology/provider; session/state architecture; backend; deployment architecture; monitoring/observability; what you personally designed, coded, debugged, and deployed. Provide reasonable non-confidential evidence where possible. 2. Personal Contribution For the system above: What parts would another engineer who worked on that project say were NOT your work? Then clearly identify what was yours versus what was performed by: another engineer; your employer; an agency/team; infrastructure/DevOps; frontend/backend colleagues; another vendor. 3. Production Incident Describe one difficult real production incident involving one or more of the following: streaming audio; conversational AI; WebRTC/WebSockets; STT/TTS; session state; concurrency; distributed systems; AI orchestration. Explain it in this order: symptom → hypotheses → investigation → telemetry/evidence → root cause → fix → prevention Do not provide a theoretical example. 4. Latency For a voice system you actually operated, what latency did you measure or observe? Where available, provide: speech-end → transcript; transcript → first model output; time to first audio; overall perceived response latency; p50/p95; what you personally changed to improve latency. If you did not measure these values, say so rather than estimating. 5. Session Recovery A user has been speaking with the AI for approximately 20 minutes. Their internet connection disappears for eight seconds and then returns. Explain exactly: what state should survive; where that state should live; how reconnection should work; what happens to the transcript; what happens to an in-progress LLM response; what happens to an in-progress tool call; how duplicate operations are prevented; what the user should experience after reconnection. 6. Security and Authorization A conversational AI system has access to user-specific documents, structured database records, and internal tools. How do you prevent: User A from seeing User B's information; the model from requesting unauthorized information; an AI tool call from executing an unauthorized action; sensitive documents from reaching the wrong role? Explain where authorization is enforced technically. 7. Inherited System You inherit an existing voice implementation that looks impressive in demonstrations but has: occasional slow responses; lost conversation state; unreliable reconnects; inconsistent transcription; intermittent production failures. What do you investigate before deciding to replace anything? Describe your investigation order and what evidence you would collect. 8. Availability Please state: how many hours/week you can personally commit; earliest start date; other major active engagements; whether you will personally perform the engineering work; whether you are applying independently or through an agency/team. Selection Process Candidates who advance may go through: application/evidence review → technical interview → live unseen reasoning scenario → paid diagnostic using a sanitized portion of the existing Voca system → reference verification Finalists may be asked to assess existing implementation choices using: KEEP · COMPLETE · FIX · REFACTOR · REPLACE · POSTPONE · REMOVE We are looking for someone who can understand before rewriting, measure before replacing, and personally get an imperfect real-time AI system into reliable production.
Abrir na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Entrar