← Live-стрічка

Senior GenAI / LLM Evaluation Engineer – RAG, DeepEval & AI Quality

Бюджет: $12.0 - $25.0 HOURLY / PART_TIME ⭐ 5.00 (2) United Arab Emirates

python

Бажана кваліфікація

  • Досвід: Середній
We need a hands-on senior GenAI/LLM Evaluation Engineer to immediately design and implement an evaluation framework for an existing healthcare AI application. This is not a traditional QA role. The consultant will build LLM/RAG evaluation pipelines covering accuracy, faithfulness, groundedness, hallucination detection, guardrails, regression testing and clinical AI validation, using DeepEval, RAGAS, LangSmith, Promptfoo or similar tools. We specifically need someone who has personally built production LLM evaluation frameworks and can start immediately. Required: Python, LLM/RAG evaluation, automated eval pipelines, ground-truth datasets, regression testing and AI quality metrics. Healthcare/clinical AI experience is strongly preferred. Please provide an example of an LLM/RAG evaluation framework you personally designed and implemented.
Відкрити замовлення

AI-чернетка відгуку

Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.

Увійдіть, щоб згенерувати AI-чернетку.

Увійти