AI Voice Agent Optimization – Reduce Latency in Transcription-to-Notification Pipeline
Budget: $50.0
FIXED /
⭐ 5.00 (254)
United Kingdom
computer-vision, automatic-speech-recognition, python, javascript, machine-learning, api, java
We have a working AI voice workflow that needs performance optimization, not a rebuild. The pipeline currently: records a voice note → transcribes it (STT) → generates a summary (LLM) → extracts action items (LLM) → sends notifications.
All steps work correctly, but end-to-end response time is too slow and needs to feel near real-time.
What we need:
A review of the current architecture to pinpoint bottlenecks (sequential processing, model choice, queue/worker delays, notification delivery method, etc.)
Concrete optimization recommendations with expected latency impact for each
Implementation of the highest-impact fixes
Before/after latency benchmarks so we can verify improvement
Ideal Candidate: experience with async/streaming architectures, LLM pipeline optimization (parallelizing calls, prompt/model selection), and background job systems (Celery/Redis or similar). Please share a specific example of a latency problem you diagnosed and fixed, not just general experience.
Open job