← Jobs

AI Agent Developer (Python/FastAPI) – RAG, Pinecone & LLM Automation Expert Needed

Budget: $25.0 - $50.0 HOURLY / FULL_TIME ⭐ 5.00 (5) United States

python, artificial-intelligence, automation

We're hiring a senior AI Agent Developer to architect and build an intelligent, production-ready RAG (Retrieval-Augmented Generation) system with autonomous agent capabilities. This is a hands-on technical role — you'll own the backend architecture, agent logic, and vector search pipeline end-to-end. Core Responsibilities: Architect and build AI agents using LLM orchestration frameworks (LangChain, LlamaIndex, CrewAI, or AutoGen) Design and optimize RAG pipelines: document ingestion, chunking strategies, embedding generation, and retrieval tuning Implement vector search using Pinecone (namespace management, hybrid search, metadata filtering) Build scalable, async backend services using FastAPI Integrate multiple LLM providers (OpenAI GPT models, Anthropic Claude API, open-source models via Hugging Face) Develop agent tool-calling, function-calling, and multi-step reasoning/planning workflows Implement conversational memory (short-term and long-term) for agent context retention Set up prompt engineering and evaluation pipelines for output quality and reliability Containerize and deploy services (Docker, CI/CD, cloud infrastructure) Write clean, modular, well-tested Python code following best practices Highly Valued (Nice to Have): Experience with AI automation platforms (n8n, Zapier, Make.com) Familiarity with LangGraph or other agent state-machine frameworks Knowledge of MCP (Model Context Protocol) or tool-use standards Experience with Weaviate, Chroma, Qdrant, or FAISS (alternative vector DBs) Background in chatbot development, virtual assistants, or workflow automation Familiarity with LLM evaluation/observability tools (LangSmith, Weights & Biases) Experience with streaming responses and WebSocket implementations Azure AI Engineer or similar cloud AI certification What Makes a Strong Applicant: A portfolio or GitHub repo showing real RAG/agent implementations (not tutorials) Clear, specific answers about architecture decisions — not generic pitches Demonstrated understanding of production concerns (latency, cost, hallucination control, rate limits) Comfortable working async with clear milestone-based delivery To Apply, Please Include: A short summary of a RAG or AI agent system you've built (stack + outcome) Your specific experience with Pinecone and FastAPI How you approach reducing hallucinations / improving retrieval accuracy Availability and estimated timeline for this scope
Auf Upwork öffnen