AI Agent Evaluation Developer for RAG
Budget: $300.0
FIXED /
⭐ 4.83 (51)
United States
artificial-intelligence, python, natural-language-processing, typescript
Preferred qualifications
- Experience: Intermediate
# AI Agent Evaluation Developer for RAG Document Workflow POC
We are looking for a developer to build a small document agent and evaluate it using Braintrust or Plumloom.
We will provide one large PDF or Markdown document. The agent should:
* Answer questions about the document
* Search the document multiple times when needed
* Refine searches when evidence is incomplete
* Generate a structured report or brief with citations
* Capture the full agent trace
Google ADK is preferred, but LangChain or another open-source framework is acceptable. You are free to design the chat interface using AI tools of your choice.
## Evaluation Work
* Create 15 to 20 test cases
* Define expected answers, report sections, or behavior
* Evaluate correctness, grounding, completeness, citations, and search trajectory
* Review failures and identify their cause
* Turn important failures into regression tests
* Re-run evaluations after changes
* Show one before-and-after improvement
* Document what the evaluation platform measures well and where it needs improvement
## Deliverables
* Complete source code
* Setup instructions
* Working document workflow
* Exported traces
* Evaluation dataset and results
* Short findings report
*AI coding assistants are allowed. The developer must use their own model credentials. We will own all code, prompts, datasets, traces, evaluations, and project outputs.*
Duration: 1 week
---
*This will be a milestone-based project managed and paid through Upwork.*
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in