← Állások

Software Engineers Needed to Create Terminal-Bench Training Tasks for AI Agents

Költségvetés: $30.0 - $60.0 HOURLY / PART_TIME ⭐ 0.00 (0) United States

python, software-development, docker, linux, bash, devops, cicd

Előnyben részesített képesítések

  • Helyszín: Americas, Asia, Europe
  • Tapasztalat: Szakértő
Coastside Labs is hiring 4–5 strong software engineers to help build a large-scale dataset of challenging terminal-based tasks for training advanced AI coding agents. We are creating original training environments inspired by Terminal-Bench and similar agent benchmarks. The goal is to improve models at the underlying capabilities required to solve difficult terminal tasks, including debugging, Linux, Docker, infrastructure, backend engineering, databases, build systems, security, ML infrastructure, and other real-world software engineering workflows. This is not traditional data annotation or basic prompt writing. You will be building real software environments that an AI agent must inspect, debug, modify, and successfully complete. WHAT YOU’LL DO You will create Terminal-Bench-style training tasks that include: - A realistic software engineering problem - A reproducible Docker environment - Clear task instructions - A reference solution - Automated tests/verifiers that determine whether the task was solved correctly Example task areas include: - Debugging broken backend services - Linux and Docker configuration issues - CI/CD failures - Database and migration problems - Dependency and build failures - Networking and permissions issues - Bugs spanning multiple files or services - Security vulnerabilities - ML training and inference pipelines - Performance optimization - Infrastructure and systems debugging We are especially interested in tasks that require multiple steps of investigation, experimentation, and debugging rather than a single obvious code change. AI-ASSISTED TASK GENERATION We expect you to use AI coding tools heavily. The goal is not to manually build every task from scratch. Strong engineers should use models to accelerate the creation of: - Task ideas and variations - Docker environments - Starter repositories - Bugs and failure scenarios - Tests and verifiers - Reference solutions - Synthetic data and fixtures Your job is to ensure the resulting tasks are technically correct, realistic, challenging, reproducible, and properly verified. We are looking for engineers who can figure out how to create high-quality tasks at scale. WHAT MAKES A GOOD TASK A strong task should: - Resemble real software engineering work - Require meaningful terminal interaction - Require debugging or multi-step reasoning - Have a clear and objectively verifiable end state - Run reliably in a containerized environment - Be difficult because of genuine technical complexity - Include robust automated tests - Avoid obvious shortcuts or ways to game the verifier - Be original and meaningfully different from public Terminal-Bench tasks You may study Terminal-Bench to understand the format and difficulty level, but tasks must be original. Do not copy or lightly modify existing benchmark tasks. WHO WE’RE LOOKING FOR Ideal candidates have: - 3+ years of professional software engineering experience - Strong Linux and command-line skills - Strong Python and/or Bash - Experience with Docker - Experience writing automated tests - Strong debugging ability - Experience with Git - Ability to quickly understand unfamiliar codebases - Experience using modern AI coding tools Experience in the following areas is especially valuable: - Backend engineering - DevOps / SRE - Cloud infrastructure - Platform engineering - Distributed systems - Databases - Cybersecurity - Systems programming - ML infrastructure - Kubernetes - Python, Go, Rust, C/C++, Java, or TypeScript ENGAGEMENT We are initially hiring approximately 4–5 engineers. This will begin with a paid trial and can turn into substantial ongoing work for strong performers. Expected availability: 20–40 hours per week. We care about both quality and throughput. We want engineers who can use AI effectively to create many strong tasks rather than spending days manually building a single environment. Top performers may also help us build internal tooling for synthetic task generation, automated validation, agent rollouts, difficulty calibration, QA, and dataset generation. HOW TO APPLY Please answer the following: 1. How many years of professional software engineering experience do you have? 2. What are your strongest technical areas? 3. Share your GitHub, LinkedIn, portfolio, or examples of technical work. 4. Describe one difficult engineering problem you have personally debugged. 5. If you needed to create 20 different Terminal-Bench-style tasks using AI, how would you approach generating them efficiently while maintaining quality? 6. Briefly describe one Terminal-Bench-style task you would create and how you would verify that an AI agent solved it correctly. 7. Which AI coding tools do you currently use? 8. How many hours per week are you available? 9. What is your hourly rate?
Megnyitás Upworkön

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Bejelentkezés