← Live feed

Senior AI Agent Evaluation & Benchmarking Team Lead — 2 Hours/Day

Budget: $25.0 - $55.0 HOURLY / PART_TIME ⭐ 4.99 (30) Israel

python

Preferred qualifications

  • Experience: Expert
We are looking for a senior AI/LLM evaluation engineer to technically lead a team of four outstanding junior developers in a selective bootcamp project for a real global technology company. The company name, business use case, initial agents, and detailed evaluation tasks are confidential and will be shared with shortlisted candidates. The technical challenge is to build a reusable evaluation harness for API-based AI agents. The platform will execute multiple agents against a structured task suite, repeat experiments under controlled conditions, collect traces and operational metrics, score outcomes, classify failure modes, and compare results through a dashboard. The role requires more than experience building agents. We need someone who understands how to test whether agents actually work, how to build reproducible benchmarks, and how to avoid misleading conclusions from a small number of nondeterministic runs. You will provide approximately 1-2 hours of hands-on technical leadership per working day (Sun - Thu)
Open job

AI proposal draft

Generate a short cover letter to copy into the offer. Says you are interested and ready to work.

Sign in to generate an AI proposal draft.

Log in