← Live-лента

Senior GenAI / LLM Evaluation Engineer – RAG, DeepEval & AI Quality

Бюджет: $12.0 - $25.0 HOURLY / PART_TIME ⭐ 5.00 (2) United Arab Emirates

python

Предпочтительная квалификация

  • Опыт: Средний
We need a hands-on senior GenAI/LLM Evaluation Engineer to immediately design and implement an evaluation framework for an existing healthcare AI application. This is not a traditional QA role. The consultant will build LLM/RAG evaluation pipelines covering accuracy, faithfulness, groundedness, hallucination detection, guardrails, regression testing and clinical AI validation, using DeepEval, RAGAS, LangSmith, Promptfoo or similar tools. We specifically need someone who has personally built production LLM evaluation frameworks and can start immediately. Required: Python, LLM/RAG evaluation, automated eval pipelines, ground-truth datasets, regression testing and AI quality metrics. Healthcare/clinical AI experience is strongly preferred. Please provide an example of an LLM/RAG evaluation framework you personally designed and implemented.
Открыть заказ

AI-черновик отклика

Короткий текст отклика для копирования в оффер: интерес + готовность работать.

Войдите, чтобы сгенерировать AI-черновик.

Войти