Senior AI Agent / Software Engineer – Coding Agent Evaluation & Benchmarks
Budget: $15.0 - $25.0
HOURLY / PART_TIME
⭐ 0.00 (0)
United States
Preferred qualifications
- Location: Pakistan, India, Brazil, Colombia, Mexico, Argentina
- Experience: Intermediate
We are looking for an experienced Software Engineer / AI Agent Developer to help us create and evaluate challenging coding tasks for AI coding agents.
The work involves taking real open-source codebases and creating realistic software-engineering challenges such as bug fixes, feature implementations, refactoring, performance issues, and concurrency problems. You will prepare reproducible Docker-based environments, write task instructions, build automated tests and verifiers, and create reference solutions where needed.
You will also work with AI coding agents by running multiple agents against the same task and evaluating their implementations. This includes reviewing their code changes, patches, trajectories, test results, and engineering decisions, then determining whether the solution is actually correct and which implementation is stronger.
Example tasks you'll be working on
1. AI Coding Benchmark Creation
Create a difficult software-engineering benchmark from an existing repository. Define the problem, prepare the Docker/Harbor environment, write the requirements, implement automated verification, and validate that the task is challenging but solvable.
2. Agentic Coding Task Evaluation
Give the same coding task to multiple AI coding agents, review their solutions, and compare them based on correctness, edge cases, maintainability, architecture, regression risk, and overall software-engineering quality. Create objective rubrics to evaluate the results.
Required skills
Strong professional software-engineering experience
Python and/or JavaScript/TypeScript
Git/GitHub and Linux
Docker and containerized environments
Automated testing and debugging
Code review and software architecture
Ability to understand unfamiliar codebases
Experience with LLMs, AI agents, or agentic coding tools
Nice to have
LangChain / LangGraph
OpenAI, Claude, or Gemini APIs
AI coding agents
SWE-bench or similar benchmarks
Harbor
AI evaluation / benchmark development
Automated grading/verifier systems
CI/CD
Python/FastAPI
We are looking for a real software engineer with strong AI-agent experience, not someone focused only on basic n8n/Zapier automation. The ability to understand complex codebases, build reliable tests/verifiers, debug software, and critically evaluate AI-generated code is the most important part of this role.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in