Need help fixing latency bottlenecks & context issues in our Langgraph/Qdrant RAG setup
Rozpočet: $400.0
FIXED /
⭐ 0.00 (0)
Pakistan
python
Looking for a developer to help audit and speed up our RAG pipeline and multi-agent workflows. We're running FastAPI on the backend with LangGraph for agent orchestration, Qdrant for vector retrieval and self-hosted vLLM endpoints. Everything works functionally, but as we've scaled up, we're hitting a few pain points that's hurting performance and response times:
- Our LangGraph multi-agent execution is taking way too long on complex steps and state feels heavy
- Longer retrieval passes sometimes blows past our context window or just starts pulling noisy chunks
- Sometimes function calling break during LLM outputs
- Response streaming via SSE stutters/blocks under concurrent user requests.
Looking for someone who can jump into the codebase and pinpoint where the bottlenecks are coming from and clean up the implementation.
Otvoriť na Upwork