← Jobs

Senior On-Prem Voice AI Engineer — LiveKit and NVIDIA Blackwell

Budget: $500.0 FIXED / ⭐ 4.18 (33) Canada

model-optimization, pytorch, artificial-intelligence, machine-learning

Bevorzugte Qualifikationen

  • Erfahrung: Experte
We need a senior engineer to install, integrate and test a completely on-premises voice AI stack for a production customer-care agent. Our server has: 2 × NVIDIA RTX PRO 6000 Blackwell GPUs 96 GB VRAM per GPU Self-hosted LiveKit and SIP An existing Python voice-agent runtime No caller audio, transcripts, prompts, customer data or responses may be sent to external inference APIs. Proposed Stack NVIDIA Parakeet for streaming STT Whisper large-v3-turbo as an accuracy fallback Gemma 4 E4B as a fast turn interpreter Gemma 4 26B-A4B as the main LLM Kyutai streaming TTS vLLM or SGLang LiveKit Audio Turn Detector v1-mini Silero VAD and local interruption handling Docker, PostgreSQL, Redis and Qdrant Prometheus, Grafana and OpenTelemetry The final model selection may be adjusted based on measured accuracy and latency. Responsibilities Configure NVIDIA drivers, CUDA, PyTorch and NVIDIA Container Toolkit for Blackwell. Deploy and optimize the local STT, LLM and TTS models. Integrate all services with LiveKit Agents and our Python runtime. Configure self-hosted LiveKit SIP, Redis, TURN, TLS, firewall and media ports. Validate telephone codecs, sample rates and audio resampling. Optimize product-name recognition using local aliases and phonetic matching. Benchmark STT accuracy using real anonymized 8 kHz telephone audio. Measure TTFA, barge-in latency, GPU usage and concurrent-call capacity. Build reproducible Docker containers and production health checks. Configure monitoring, automatic recovery and outbound network restrictions. Provide complete installation and operations documentation. Required Experience Applicants should have hands-on experience with: NVIDIA CUDA and GPU inference NVIDIA Blackwell or recent enterprise GPUs vLLM or SGLang PyTorch and Hugging Face Streaming STT and TTS NVIDIA NeMo or Parakeet Whisper Gemma models LiveKit, WebRTC and SIP Python and FastAPI Docker and Linux Prometheus and Grafana This is not a prompt-engineering or hosted-API project. Deliverables Fully working on-premises voice AI stack Reproducible Docker deployment Pinned dependency and model versions STT accuracy report TTFA and concurrency report Monitoring dashboards Security and outbound-traffic verification Installation, operations and troubleshooting guides Source code committed to our repository Knowledge-transfer session
Auf Upwork öffnen

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Anmelden