Urgent Hiring: Sr. Voice and Chat Ai Engineer
Költségvetés: -
HOURLY / FULL_TIME
⭐ 5.00 (13)
United States
Előnyben részesített képesítések
- Tapasztalat: Szakértő
Important Notice to Applicants:
Please note that we are only contacting and communicating with candidates through Upwork. Any applications or direct contact made outside of these channels, including emails, social media messages, direct messages to our CEO, or messages sent to our general company email, will not be considered and will be automatically declined.
About the Company:
We are a private U.S.-based company operating across multiple departments that support legal, staffing, and client-service operations. Our teams collaborate in dynamic, fast-paced environments focused on innovation, integrity, and client success.
In this role, you’ll work closely with leadership and cross-functional teams, making a real impact in operational, legal, and client-focused projects—all from the comfort of your home. Details about our company structure and associated teams will be introduced during the interview process
About the Role
We're building a multi-tenant AI receptionist and case intake platform for law firms. It answers inbound calls over SIP, whether they originate on the PSTN or from the firm's own SIP infrastructure. Each call loads that firm's tenant config and runs the right intake workflow, collecting and confirming structured case details. It applies the firm's qualification and routing rules without giving legal advice. From there it books consultations, warm-transfers to staff, or creates follow-up tasks. Intake records, transcripts, summaries, and recordings go to firm staff through controlled access.
We are looking for a highly autonomous Senior Voice AI Engineer to build this from the ground up and deliver a pilot-ready MVP within 90 days. This is a hands-on, zero-to-one role: you will own architecture, technology decisions, implementation, deployment, testing, and technical operation, working directly with stakeholders who provide product direction, domain input, and rapid feedback rather than a detailed specification. The MVP channel is inbound voice with no outbound dialing, but design the conversation and workflow layer so web chat and other text-based channels can reuse the same tenant rules, integrations, and guardrails. This is applied product engineering, not AI research.
What You Will Own
Voice and Conversation Engineering
Build and optimize the real-time voice pipeline — LiveKit Agents, Strands, or another well-justified architecture — with streaming STT, LLM, and TTS or speech-to-speech models, tuned for perceived latency, transcription accuracy, and naturalness.
Implement voice activity detection, endpointing, turn-taking, barge-in, interruption recovery, silence handling, and conversation repair for noisy lines, accents, incomplete answers, long pauses, and caller corrections.
Build deterministic workflow and state management — reliable tool calling, validation, retries, timeouts, idempotency, failure recovery — so critical intake actions never depend on prompt behavior alone.
Telephony:
- Configure inbound numbers, SIP trunks, dispatch rules, call routing, DTMF, and human transfers with providers such as LiveKit, Twilio, Telnyx, or Bandwidth.
- Handle dropped calls, provider failures, after-hours routing, duplicate calls, spam, and wrong numbers, with monitoring for call quality, transfer success, and provider performance.
- Multi-Tenant Product Architecture
- Design secure tenant isolation across data, prompts, workflows, integrations, credentials, recordings, and reporting.
- Build a reusable configuration model — practice areas, questions, qualification criteria, disclosures, transfer destinations, calendars, languages, business rules — versioned so production behavior can be reproduced and audited, and so scaling from pilot to many firms is configuration rather than custom implementation.
Integrations and Data:
- Build APIs and webhooks for calendars, CRMs, and legal case-management platforms such as Clio, Filevine, Lawmatics, Litify, or MyCase, with queues, retries, reconciliation, and audit trails.
- Convert conversations into validated, structured intake records and safely store call metadata, transcripts, summaries, recordings, tool activity, and outcomes.
Safety, Security, and Reliability:
- Implement guardrails against legal advice, promised outcomes, an implied attorney-client relationship, or operation outside approved workflows, existing clients, sensitive matters, unsupported practice areas, and callers who need a human.
- Build configurable disclosures and recording controls per firm and jurisdiction; protect caller data through encryption, least-privilege access, secret management, retention controls, audit logging, and review of model and vendor data-handling policies; and run production monitoring, alerting, tracing, and incident response.
Testing and Evaluation:
Build automated conversation tests and regression suites covering interruptions, silence, noisy audio, accents, ambiguous or invalid answers, provider failures, and unexpected caller behavior.
Set measurable standards for intake completeness, field accuracy, tool-call and transfer success, latency, safety violations, and cost per call; review real calls to improve prompts, workflow logic, models, and provider selection; and keep a repeatable process for safely deploying changes.
Required Experience:
5+ years of professional software engineering — the scope and quality of what you have shipped matters more than the number.
Hands-on experience building and launching at least one real-time, phone-based voice AI agent in production or a serious customer pilot. Demos alone are not sufficient.
Strong Python: asynchronous programming, streaming systems, API development, and production debugging.
LiveKit Agents or a comparable real-time voice framework (Pipecat, Vapi, Retell, or a custom WebRTC/WebSocket implementation), integrating STT, LLM, TTS, or speech-to-speech models from multiple providers.
Practical telephony experience — SIP, PSTN providers, call routing, DTMF, transfers, codecs — with deep understanding of latency, voice activity detection, endpointing, turn-taking, barge-in, and streaming tool execution.
Dependable agent workflows with explicit state, structured outputs, and deterministic business rules; multi-tenant SaaS with strong data isolation; cloud deployment and operations (containers, CI/CD, infrastructure as code, monitoring); and clear communication of technical tradeoffs to a non-specialist.
Preferred Experience:
Production LiveKit (Agents, Cloud, or self-hosted), Strands Agents, Amazon Bedrock, AgentCore, or Nova Sonic.
OpenAI Realtime, Deepgram, AssemblyAI, ElevenLabs, Cartesia, or similar providers.
Web chat or omnichannel conversational systems, multilingual voice, and legal technology, healthcare, financial services, or another privacy-sensitive industry.
A computer science degree is not required; demonstrated production work, technical judgment, and ownership matter more.
What Success Looks Like
Within 90 days, pilot law firms can route inbound calls to a production voice agent that:
Conducts a natural, professional intake conversation and accurately captures the information the firm needs.
Follows tenant-specific rules and disclosures, completes approved actions reliably, transfers callers to a human when appropriate, and avoids legal advice and unsupported claims.
Produces useful structured records and summaries, can be monitored, evaluated, debugged, and improved, and supports additional firms primarily through configuration rather than custom development.
Megnyitás Upworkön
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Bejelentkezés