← Jobb

AI Full Stack Developer

Budget: $15.0 - $35.0 HOURLY / PART_TIME ⭐ 0.00 (0) Israel

python, api, api-integration, amazon-web-services, artificial-intelligence

Föredragna kvalifikationer

  • Erfarenhet: Medel
1. About the Project I am building an AI-powered service for finance companies that automatically answers customer questionnaires based on: * Documents uploaded by customers * Information retrieved through web search * RAG (Retrieval-Augmented Generation) * LLM-based evaluation agents For each questionnaire, the system should return a "Yes / No / Unknown" flag for every question. 2. The Problem The core system is already working, but we have a critical consistency/reliability issue with the evaluation agent. For example, I provide 30 questions for the same uploaded document/package. When I run the exact same questions multiple times, the system sometimes returns a different number of answers. For example: Run 1 * Yes: 14 * No: 10 * Unknown: 6 * Total: 30 Run 2 * Yes: 12 * No: 11 * Unknown: 5 * Total: 28 The number of returned flags can change between runs, even though the input questions and documents are the same. The goal is to ensure that every question reliably produces exactly one Yes/No/Unknown result, while maintaining the accuracy of the evaluation. 3. What I Need I am looking for an expert AI Full-Stack Developer / LLM Engineer who can quickly investigate the existing implementation, identify the root cause, and implement a reliable fix. You will need to investigate areas such as: * LLM response consistency * Prompt design and structured output * JSON/schema validation * Missing or duplicated responses * Agent/tool execution * RAG retrieval behavior * Web-search integration * Async/concurrent processing * Token/context limitations * Parsing and post-processing * Retry and error-handling logic * LLM temperature/model configuration * Evaluation-agent architecture * Ensuring one result per input question 3.1 Expected Result If I provide 30 questions, the system must always produce: Exactly 30 evaluation results. Each result must contain: ```text Question → Yes / No / Unknown ``` No questions should be silently dropped, duplicated, or omitted. If the LLM fails to produce a valid answer for a question, the application should handle the failure deterministically rather than returning an incomplete result set. 4. Required Skills 4.1 Must Have * Strong Python experience * Experience building production AI/LLM applications * RAG systems * LangChain or similar LLM frameworks * LLM agents * Prompt engineering * Structured LLM outputs / JSON schemas * OpenAI / Anthropic / AWS Bedrock or similar APIs * Debugging nondeterministic LLM behavior * Async Python and concurrent processing * API development, preferably FastAPI * Experience troubleshooting production AI systems 4.2 Strong Plus * LLM evaluation frameworks such as Ragas or OpenAI Evals * Vector databases such as Pinecone, Weaviate, pgvector, or similar * Web-search/retrieval pipelines * LLM observability and tracing * OpenTelemetry * Experience with financial/compliance/enterprise AI systems * React/Next.js experience 5. What You Will Do * Review the existing RAG/evaluation architecture. * Reproduce the inconsistent results. * Trace the complete pipeline from questionnaire → retrieval → evaluation agent → final response. * Identify why questions/results are being dropped, duplicated, or inconsistently generated. * Implement a robust solution. * Add validation to guarantee one result per question. * Add appropriate retries/fallback handling where necessary. * Test the system repeatedly using the same input. * Provide a clear explanation of the root cause and the changes made. 6. Ideal Candidate You are a strong fit if you have experience debugging real production LLM systems, rather than only building simple chatbot applications. I am specifically looking for someone who understands that LLM applications can fail at multiple layers — model generation, structured output, agents, asynchronous execution, retrieval, parsing, and application logic — and can systematically identify the actual source of the problem. 7. To Apply Please provide: 1. A brief description of your experience with RAG/LLM systems. 2. Examples of similar LLM consistency/debugging issues you have solved. 3. Which LLM frameworks/models you have worked with. 4. Your availability to start. 5. Your hourly rate or estimated fixed price.
Öppna på Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Logga in