← Вакансии

Need help fixing latency bottlenecks & context issues in our Langgraph/Qdrant RAG setup

Бюджет: $400.0 FIXED / ⭐ 0.00 (0) Pakistan

python

Looking for a developer to help audit and speed up our RAG pipeline and multi-agent workflows. We're running FastAPI on the backend with LangGraph for agent orchestration, Qdrant for vector retrieval and self-hosted vLLM endpoints. Everything works functionally, but as we've scaled up, we're hitting a few pain points that's hurting performance and response times: - Our LangGraph multi-agent execution is taking way too long on complex steps and state feels heavy - Longer retrieval passes sometimes blows past our context window or just starts pulling noisy chunks - Sometimes function calling break during LLM outputs - Response streaming via SSE stutters/blocks under concurrent user requests. Looking for someone who can jump into the codebase and pinpoint where the bottlenecks are coming from and clean up the implementation.
Открыть заказ