Senior On-Prem Voice AI Engineer — LiveKit and NVIDIA Blackwell
Budget: $500.0
FIXED /
⭐ 4.18 (33)
Canada
model-optimization, pytorch, artificial-intelligence, machine-learning
Qualifiche preferite
- Esperienza: Esperto
We need a senior engineer to install, integrate and test a completely on-premises voice AI stack for a production customer-care agent.
Our server has:
2 × NVIDIA RTX PRO 6000 Blackwell GPUs
96 GB VRAM per GPU
Self-hosted LiveKit and SIP
An existing Python voice-agent runtime
No caller audio, transcripts, prompts, customer data or responses may be sent to external inference APIs.
Proposed Stack
NVIDIA Parakeet for streaming STT
Whisper large-v3-turbo as an accuracy fallback
Gemma 4 E4B as a fast turn interpreter
Gemma 4 26B-A4B as the main LLM
Kyutai streaming TTS
vLLM or SGLang
LiveKit Audio Turn Detector v1-mini
Silero VAD and local interruption handling
Docker, PostgreSQL, Redis and Qdrant
Prometheus, Grafana and OpenTelemetry
The final model selection may be adjusted based on measured accuracy and latency.
Responsibilities
Configure NVIDIA drivers, CUDA, PyTorch and NVIDIA Container Toolkit for Blackwell.
Deploy and optimize the local STT, LLM and TTS models.
Integrate all services with LiveKit Agents and our Python runtime.
Configure self-hosted LiveKit SIP, Redis, TURN, TLS, firewall and media ports.
Validate telephone codecs, sample rates and audio resampling.
Optimize product-name recognition using local aliases and phonetic matching.
Benchmark STT accuracy using real anonymized 8 kHz telephone audio.
Measure TTFA, barge-in latency, GPU usage and concurrent-call capacity.
Build reproducible Docker containers and production health checks.
Configure monitoring, automatic recovery and outbound network restrictions.
Provide complete installation and operations documentation.
Required Experience
Applicants should have hands-on experience with:
NVIDIA CUDA and GPU inference
NVIDIA Blackwell or recent enterprise GPUs
vLLM or SGLang
PyTorch and Hugging Face
Streaming STT and TTS
NVIDIA NeMo or Parakeet
Whisper
Gemma models
LiveKit, WebRTC and SIP
Python and FastAPI
Docker and Linux
Prometheus and Grafana
This is not a prompt-engineering or hosted-API project.
Deliverables
Fully working on-premises voice AI stack
Reproducible Docker deployment
Pinned dependency and model versions
STT accuracy report
TTFA and concurrency report
Monitoring dashboards
Security and outbound-traffic verification
Installation, operations and troubleshooting guides
Source code committed to our repository
Knowledge-transfer session
Apri su Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Accedi