AI Full Stack Developer
Budget: $15.0 - $35.0
HOURLY / PART_TIME
⭐ 0.00 (0)
Israel
python, api-integration, artificial-intelligence, web-programming
Föredragna kvalifikationer
- Erfarenhet: Medel
## About the Project
I am building an AI-powered service for **finance companies** that automatically answers customer questionnaires based on:
* Documents uploaded by customers
* Information retrieved through web search
* RAG (Retrieval-Augmented Generation)
* LLM-based evaluation agents
For each questionnaire, the system should return a **Yes / No / Unknown** flag for every question.
## The Problem
The core system is already working, but we have a critical **consistency/reliability issue** with the evaluation agent.
For example, I provide **30 questions** for the same uploaded document/package. When I run the exact same questions multiple times, the system sometimes returns a different number of answers.
For example:
**Run 1**
* Yes: 14
* No: 10
* Unknown: 6
* Total: 30
**Run 2**
* Yes: 12
* No: 11
* Unknown: 5
* Total: 28
The number of returned flags can change between runs, even though the input questions and documents are the same.
The goal is to ensure that **every question reliably produces exactly one Yes/No/Unknown result**, while maintaining the accuracy of the evaluation.
## What I Need
I am looking for an **expert AI Full-Stack Developer / LLM Engineer** who can quickly investigate the existing implementation, identify the root cause, and implement a reliable fix.
You will need to investigate areas such as:
* LLM response consistency
* Prompt design and structured output
* JSON/schema validation
* Missing or duplicated responses
* Agent/tool execution
* RAG retrieval behavior
* Web-search integration
* Async/concurrent processing
* Token/context limitations
* Parsing and post-processing
* Retry and error-handling logic
* LLM temperature/model configuration
* Evaluation-agent architecture
* Ensuring one result per input question
### Expected Result
If I provide **30 questions**, the system must always produce:
**Exactly 30 evaluation results.**
Each result must contain:
```text
Question → Yes / No / Unknown
```
No questions should be silently dropped, duplicated, or omitted.
If the LLM fails to produce a valid answer for a question, the application should handle the failure deterministically rather than returning an incomplete result set.
## Required Skills
### Must Have
* Strong Python experience
* Experience building production AI/LLM applications
* RAG systems
* LangChain or similar LLM frameworks
* LLM agents
* Prompt engineering
* Structured LLM outputs / JSON schemas
* OpenAI / Anthropic / AWS Bedrock or similar APIs
* Debugging nondeterministic LLM behavior
* Async Python and concurrent processing
* API development, preferably FastAPI
* Experience troubleshooting production AI systems
### Strong Plus
* LLM evaluation frameworks such as Ragas or OpenAI Evals
* Vector databases such as Pinecone, Weaviate, pgvector, or similar
* Web-search/retrieval pipelines
* LLM observability and tracing
* OpenTelemetry
* Experience with financial/compliance/enterprise AI systems
* React/Next.js experience
## What You Will Do
1. Review the existing RAG/evaluation architecture.
2. Reproduce the inconsistent results.
3. Trace the complete pipeline from questionnaire → retrieval → evaluation agent → final response.
4. Identify why questions/results are being dropped, duplicated, or inconsistently generated.
5. Implement a robust solution.
6. Add validation to guarantee one result per question.
7. Add appropriate retries/fallback handling where necessary.
8. Test the system repeatedly using the same input.
9. Provide a clear explanation of the root cause and the changes made.
## Ideal Candidate
You are a strong fit if you have experience debugging real production LLM systems, rather than only building simple chatbot applications.
I am specifically looking for someone who understands that LLM applications can fail at multiple layers — model generation, structured output, agents, asynchronous execution, retrieval, parsing, and application logic — and can systematically identify the actual source of the problem.
## To Apply
Please provide:
1. A brief description of your experience with RAG/LLM systems.
2. Examples of similar LLM consistency/debugging issues you have solved.
3. Which LLM frameworks/models you have worked with.
4. Your availability to start.
5. Your hourly rate or estimated fixed price.
Öppna på Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Logga in