← Joburi

Senior AI Agent / Software Engineer – Coding Agent Evaluation & Benchmarks

Buget: $15.0 - $25.0 HOURLY / PART_TIME ⭐ 0.00 (0) United States

Calificări preferate

  • Locație: Pakistan, India, Brazil, Colombia, Mexico, Argentina
  • Experiență: Intermediar
We are looking for an experienced Software Engineer / AI Agent Developer to help us create and evaluate challenging coding tasks for AI coding agents. The work involves taking real open-source codebases and creating realistic software-engineering challenges such as bug fixes, feature implementations, refactoring, performance issues, and concurrency problems. You will prepare reproducible Docker-based environments, write task instructions, build automated tests and verifiers, and create reference solutions where needed. You will also work with AI coding agents by running multiple agents against the same task and evaluating their implementations. This includes reviewing their code changes, patches, trajectories, test results, and engineering decisions, then determining whether the solution is actually correct and which implementation is stronger. Example tasks you'll be working on 1. AI Coding Benchmark Creation Create a difficult software-engineering benchmark from an existing repository. Define the problem, prepare the Docker/Harbor environment, write the requirements, implement automated verification, and validate that the task is challenging but solvable. 2. Agentic Coding Task Evaluation Give the same coding task to multiple AI coding agents, review their solutions, and compare them based on correctness, edge cases, maintainability, architecture, regression risk, and overall software-engineering quality. Create objective rubrics to evaluate the results. Required skills Strong professional software-engineering experience Python and/or JavaScript/TypeScript Git/GitHub and Linux Docker and containerized environments Automated testing and debugging Code review and software architecture Ability to understand unfamiliar codebases Experience with LLMs, AI agents, or agentic coding tools Nice to have LangChain / LangGraph OpenAI, Claude, or Gemini APIs AI coding agents SWE-bench or similar benchmarks Harbor AI evaluation / benchmark development Automated grading/verifier systems CI/CD Python/FastAPI We are looking for a real software engineer with strong AI-agent experience, not someone focused only on basic n8n/Zapier automation. The ability to understand complex codebases, build reliable tests/verifiers, debug software, and critically evaluate AI-generated code is the most important part of this role.
Deschide pe Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Autentificare