← Missions

AI Agent and LLM Evals Course Content Creator

Budget: $500.0 FIXED / ⭐ 4.83 (51) United States

content-writing, english, writing

Qualifications préférées

  • Expérience : Intermédiaire
# AI Agent & LLM Evaluation Course Content Creator We are looking for an experienced AI developer to create a short, practical course on evaluating LLM applications and AI agents. The core theme is **evaluation reliability**: how teams determine whether an evaluation result reflects real AI performance or noise from model variability, judge variability, weak datasets, or poor evaluation design. This is a technical content role. You should have hands-on experience building or evaluating LLM applications, RAG systems, or agents. ## Course Topics * Why traditional software testing alone is insufficient for LLM applications * Defining multidimensional quality across correctness, relevance, grounding, safety, task completion, and other dimensions * Designing evaluation criteria, metrics, rubrics, datasets, and representative test cases * Measuring quality using evaluators, scores, and results across the selected dimensions * **Evaluation reliability: understanding when evaluation results are trustworthy and when they may be misleading** * LLM-as-a-judge, deterministic evaluators, human review, calibration, and trade-offs * RAG evaluation: retrieval quality, context relevance, faithfulness, and answer quality * Agent evaluation: outcomes, tool use, decisions, and trajectory quality * Tracing failures to prompts, retrieval, models, tools, or application logic * Turning failures into regression tests and measuring reliable improvement over time ## Deliverables * Course outline and lesson scripts * Practical evaluation examples and demos * Sample datasets, rubrics, evaluators, and exercises * Supporting diagrams or charts ## Requirements * Hands-on LLM or agent development experience * Strong understanding of LLM, RAG, and agent evaluation * Understanding of evaluation reliability and LLM-as-a-judge variability * Python & Experience with evaluation or observability tools * Ability to explain technical concepts clearly **Budget:** $500 fixed price **Project Duration:** 10 calendar days **Delivery:** Milestone-based through Upwork
Ouvrir sur Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Connexion