← Live feed

AI Full Stack Developer

Budget: $15.0 - $35.0 HOURLY / PART_TIME ⭐ 0.00 (0) Israel

python, api-integration, artificial-intelligence, web-programming

Preferred qualifications

  • Experience: Intermediate
## About the Project I am building an AI-powered service for **finance companies** that automatically answers customer questionnaires based on: * Documents uploaded by customers * Information retrieved through web search * RAG (Retrieval-Augmented Generation) * LLM-based evaluation agents For each questionnaire, the system should return a **Yes / No / Unknown** flag for every question. ## The Problem The core system is already working, but we have a critical **consistency/reliability issue** with the evaluation agent. For example, I provide **30 questions** for the same uploaded document/package. When I run the exact same questions multiple times, the system sometimes returns a different number of answers. For example: **Run 1** * Yes: 14 * No: 10 * Unknown: 6 * Total: 30 **Run 2** * Yes: 12 * No: 11 * Unknown: 5 * Total: 28 The number of returned flags can change between runs, even though the input questions and documents are the same. The goal is to ensure that **every question reliably produces exactly one Yes/No/Unknown result**, while maintaining the accuracy of the evaluation. ## What I Need I am looking for an **expert AI Full-Stack Developer / LLM Engineer** who can quickly investigate the existing implementation, identify the root cause, and implement a reliable fix. You will need to investigate areas such as: * LLM response consistency * Prompt design and structured output * JSON/schema validation * Missing or duplicated responses * Agent/tool execution * RAG retrieval behavior * Web-search integration * Async/concurrent processing * Token/context limitations * Parsing and post-processing * Retry and error-handling logic * LLM temperature/model configuration * Evaluation-agent architecture * Ensuring one result per input question ### Expected Result If I provide **30 questions**, the system must always produce: **Exactly 30 evaluation results.** Each result must contain: ```text Question → Yes / No / Unknown ``` No questions should be silently dropped, duplicated, or omitted. If the LLM fails to produce a valid answer for a question, the application should handle the failure deterministically rather than returning an incomplete result set. ## Required Skills ### Must Have * Strong Python experience * Experience building production AI/LLM applications * RAG systems * LangChain or similar LLM frameworks * LLM agents * Prompt engineering * Structured LLM outputs / JSON schemas * OpenAI / Anthropic / AWS Bedrock or similar APIs * Debugging nondeterministic LLM behavior * Async Python and concurrent processing * API development, preferably FastAPI * Experience troubleshooting production AI systems ### Strong Plus * LLM evaluation frameworks such as Ragas or OpenAI Evals * Vector databases such as Pinecone, Weaviate, pgvector, or similar * Web-search/retrieval pipelines * LLM observability and tracing * OpenTelemetry * Experience with financial/compliance/enterprise AI systems * React/Next.js experience ## What You Will Do 1. Review the existing RAG/evaluation architecture. 2. Reproduce the inconsistent results. 3. Trace the complete pipeline from questionnaire → retrieval → evaluation agent → final response. 4. Identify why questions/results are being dropped, duplicated, or inconsistently generated. 5. Implement a robust solution. 6. Add validation to guarantee one result per question. 7. Add appropriate retries/fallback handling where necessary. 8. Test the system repeatedly using the same input. 9. Provide a clear explanation of the root cause and the changes made. ## Ideal Candidate You are a strong fit if you have experience debugging real production LLM systems, rather than only building simple chatbot applications. I am specifically looking for someone who understands that LLM applications can fail at multiple layers — model generation, structured output, agents, asynchronous execution, retrieval, parsing, and application logic — and can systematically identify the actual source of the problem. ## To Apply Please provide: 1. A brief description of your experience with RAG/LLM systems. 2. Examples of similar LLM consistency/debugging issues you have solved. 3. Which LLM frameworks/models you have worked with. 4. Your availability to start. 5. Your hourly rate or estimated fixed price.
Open job

AI proposal draft

Generate a short cover letter to copy into the offer. Says you are interested and ready to work.

Sign in to generate an AI proposal draft.

Log in