← Zakázky

Research AI Engineer for Voice Platform

Rozpočet: $15.0 - $25.0 HOURLY / FULL_TIME ⭐ 5.00 (3) United States

pytorch, deep-neural-networks, artificial-intelligence, machine-learning, python, natural-language-processing, artificial-neural-networks, deep-learning, tensorflow

Preferované kvalifikace

  • Zkušenost: Expert
We are looking for a Research AI Engineer / Voice AI Engineer with strong hands-on experience in ASR, TTS, LLMs, GPU optimization, and production-grade model deployment. Our goal is to build and operate a highly scalable, real-time voice AI platform capable of supporting 1000+ concurrent voice calls with low latency and high reliability. You will work on both AI research/model optimization and production engineering, including deploying open-source ASR, TTS, and LLM models, optimizing GPU inference, and extending models to support new languages and voices. What You'll Work On Deploy and optimize open-source ASR, TTS, and LLM models for production. Build highly scalable inference infrastructure capable of supporting 500+ concurrent voice calls. Work with LiveKit and real-time voice communication pipelines. Optimize end-to-end voice latency, including: Time to first audio Time to first token Streaming ASR latency TTS generation latency GPU utilization Concurrent request throughput Benchmark models under high concurrency and identify bottlenecks. Optimize inference using technologies such as vLLM and other modern GPU inference/runtime technologies. Implement batching, continuous batching, dynamic batching, KV-cache optimization, quantization, parallelism, and GPU memory optimization where appropriate. Deploy models across GPU servers and design production-grade inference architectures. Monitor and troubleshoot GPU utilization, memory consumption, latency, throughput, and failures. ASR & TTS Research We need someone who understands more than simply calling an existing model API. You should have experience with the underlying ASR/TTS models and training pipelines, including: Fine-tuning ASR models. Fine-tuning and training TTS models. Preparing and cleaning speech datasets. Audio preprocessing and feature extraction. Speaker/voice data preparation. Adding support for new languages. Improving pronunciation and multilingual performance. Adding or adapting new voices/speakers. Understanding phonemes, tokenizers, vocabularies, alignments, and language-specific challenges. Evaluating WER/CER for ASR. Evaluating TTS quality, speaker similarity, pronunciation, naturalness, and latency. Troubleshooting model hallucinations, pronunciation issues, repetitions, and audio artifacts. Production Infrastructure You should be comfortable taking a research model and turning it into a production-ready inference service. Experience with the following is highly desirable: vLLM LiveKit Docker Kubernetes Linux GPU servers NVIDIA CUDA PyTorch FastAPI / gRPC Redis / messaging systems Prometheus / Grafana or equivalent monitoring Distributed inference Load testing and concurrency testing CI/CD Cloud GPU infrastructure Experience with SGLang or similar high-performance LLM inference/runtime frameworks is also strongly preferred.
Otevřít na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Přihlásit