← Livefeed

Fix Critical Production Issues in Multi-Agent AI Platform — LangGraph + LangChain + FastAPI

Budget: $50.0 FIXED / ⭐ 0.00 (0) United Arab Emirates

Gewenste kwalificaties

  • Ervaring: Expert
We have a multi-agent AI platform running in production that is experiencing several critical issues affecting reliability and performance. The system is built on LangGraph for agent orchestration, LangChain for tool calling and chain management, FastAPI for the backend API layer, and PostgreSQL with pgvector for vector storage and retrieval. This is not a greenfield build. The platform exists, users are on it, and things are breaking in ways that need a senior engineer who has actually operated multi-agent systems in production before — not someone who has built demos or prototypes. What is broken and needs fixing: Agent workflows are behaving non-deterministically in production under real load. Certain tool calling sequences are failing silently without proper error propagation back to the orchestration layer. Memory and state management across multi-step agent workflows is losing context between steps in specific edge cases. The retrieval layer is returning irrelevant chunks in certain query patterns, degrading response quality significantly. Some agent loops are not terminating correctly and hitting maximum iteration limits when they should be resolving successfully. API response times are degrading under concurrent requests and we need diagnosis and optimization. What I need from you: First you will review the existing codebase and architecture and give me a clear written diagnosis of what is wrong and in what priority order it should be fixed. Then you will fix it. I do not want someone who will just read the code and recommend changes — I need someone who will own the fixes end to end and verify them against production behavior before handing back. You must have: Genuine hands-on experience operating LangGraph or LangChain agent systems in production with real users, not controlled test environments. Deep understanding of why multi-step agent workflows fail under real conditions and specifically what mechanisms prevent those failures. Experience debugging FastAPI performance issues under concurrent load. Familiarity with pgvector retrieval optimization including hybrid search, re-ranking, and chunk strategy tuning. Ability to add proper observability — logging, tracing, and audit trails on agent decisions — so we can see what is happening inside the system going forward. Please do not apply if you have only built agents in notebooks or demo environments. This system has real users and real consequences and I need someone who understands the difference between making something work once and making it work reliably every time. To apply please include a specific example of a production agent system you have operated, what broke in it, and what you changed to prevent that class of problem from happening again. Generic applications will not be considered. Budget is negotiable based on experience. I would rather pay the right rate for someone who solves this correctly than pay less for someone who partially fixes it.
Openen op Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Inloggen