AI Agent and LLM Evals Course Content Creator
Presupuesto: $500.0
FIXED /
⭐ 4.83 (51)
United States
content-writing, english, writing
Cualificaciones preferidas
- Experiencia: Intermedio
# AI Agent & LLM Evaluation Course Content Creator
We are looking for an experienced AI developer to create a short, practical course on evaluating LLM applications and AI agents.
The core theme is **evaluation reliability**: how teams determine whether an evaluation result reflects real AI performance or noise from model variability, judge variability, weak datasets, or poor evaluation design.
This is a technical content role. You should have hands-on experience building or evaluating LLM applications, RAG systems, or agents.
## Course Topics
* Why traditional software testing alone is insufficient for LLM applications
* Defining multidimensional quality across correctness, relevance, grounding, safety, task completion, and other dimensions
* Designing evaluation criteria, metrics, rubrics, datasets, and representative test cases
* Measuring quality using evaluators, scores, and results across the selected dimensions
* **Evaluation reliability: understanding when evaluation results are trustworthy and when they may be misleading**
* LLM-as-a-judge, deterministic evaluators, human review, calibration, and trade-offs
* RAG evaluation: retrieval quality, context relevance, faithfulness, and answer quality
* Agent evaluation: outcomes, tool use, decisions, and trajectory quality
* Tracing failures to prompts, retrieval, models, tools, or application logic
* Turning failures into regression tests and measuring reliable improvement over time
## Deliverables
* Course outline and lesson scripts
* Practical evaluation examples and demos
* Sample datasets, rubrics, evaluators, and exercises
* Supporting diagrams or charts
## Requirements
* Hands-on LLM or agent development experience
* Strong understanding of LLM, RAG, and agent evaluation
* Understanding of evaluation reliability and LLM-as-a-judge variability
* Python & Experience with evaluation or observability tools
* Ability to explain technical concepts clearly
**Budget:** $500 fixed price
**Project Duration:** 10 calendar days
**Delivery:** Milestone-based through Upwork
Abrir en Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Entrar