Senior GenAI / LLM Evaluation Engineer – RAG, DeepEval & AI Quality
Бюджэт: $12.0 - $25.0
HOURLY / PART_TIME
⭐ 5.00 (2)
United Arab Emirates
python
Пераважная кваліфікацыя
- Вопыт: Сярэдні
We need a hands-on senior GenAI/LLM Evaluation Engineer to immediately design and implement an evaluation framework for an existing healthcare AI application. This is not a traditional QA role.
The consultant will build LLM/RAG evaluation pipelines covering accuracy, faithfulness, groundedness, hallucination detection, guardrails, regression testing and clinical AI validation, using DeepEval, RAGAS, LangSmith, Promptfoo or similar tools.
We specifically need someone who has personally built production LLM evaluation frameworks and can start immediately.
Required: Python, LLM/RAG evaluation, automated eval pipelines, ground-truth datasets, regression testing and AI quality metrics.
Healthcare/clinical AI experience is strongly preferred.
Please provide an example of an LLM/RAG evaluation framework you personally designed and implemented.
Адкрыць заказ
AI-чарнавік адказу
Згенеруйце кароткі cover letter па гэтай вакансіі. Перад адпраўкай адрэдагуйце.
Увайдзіце, каб згенерыраваць AI-чарнавік.
Увайсці