← Trabajos

AI Voice Agent Optimization – Reduce Latency in Transcription-to-Notification Pipeline

Presupuesto: $50.0 FIXED / ⭐ 5.00 (254) United Kingdom

computer-vision, automatic-speech-recognition, python, javascript, machine-learning, api, java

We have a working AI voice workflow that needs performance optimization, not a rebuild. The pipeline currently: records a voice note → transcribes it (STT) → generates a summary (LLM) → extracts action items (LLM) → sends notifications. All steps work correctly, but end-to-end response time is too slow and needs to feel near real-time. What we need: A review of the current architecture to pinpoint bottlenecks (sequential processing, model choice, queue/worker delays, notification delivery method, etc.) Concrete optimization recommendations with expected latency impact for each Implementation of the highest-impact fixes Before/after latency benchmarks so we can verify improvement Ideal Candidate: experience with async/streaming architectures, LLM pipeline optimization (parallelizing calls, prompt/model selection), and background job systems (Celery/Redis or similar). Please share a specific example of a latency problem you diagnosed and fixed, not just general experience.
Abrir en Upwork