Senior GenAI / LLM Evaluation Engineer – RAG, DeepEval & AI Quality
Budget: $12.0 - $25.0
HOURLY / PART_TIME
⭐ 5.00 (2)
United Arab Emirates
python
Preferred qualifications
- Experience: Intermediate
We need a hands-on senior GenAI/LLM Evaluation Engineer to immediately design and implement an evaluation framework for an existing healthcare AI application. This is not a traditional QA role.
The consultant will build LLM/RAG evaluation pipelines covering accuracy, faithfulness, groundedness, hallucination detection, guardrails, regression testing and clinical AI validation, using DeepEval, RAGAS, LangSmith, Promptfoo or similar tools.
We specifically need someone who has personally built production LLM evaluation frameworks and can start immediately.
Required: Python, LLM/RAG evaluation, automated eval pipelines, ground-truth datasets, regression testing and AI quality metrics.
Healthcare/clinical AI experience is strongly preferred.
Please provide an example of an LLM/RAG evaluation framework you personally designed and implemented.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in