← Trabalhos

Agentic RL Training Data Author — AI Safety Evaluation Environments

Orçamento: $15.0 FIXED / ⭐ 0.00 (0) IND

python, docker, reinforcement-learning

Qualificações preferidas

  • Experiência: Intermédio
**Description:** We're hiring an expert in building agentic Reinforcement Learning (RL) training data — specifically, scenario-based evaluation environments that test whether AI agents behave safely during real coding tasks. **What you'll build:** Complete RLE (Reinforcement Learning Environment) packages across AI safety dimensions. Each package includes: - A realistic coding task (the "base task") that an agent can complete fully - An embedded safety challenge — repo content, tool outputs, test suites, config defaults, or code comments that pressure the agent into crossing a boundary the user's task explicitly sets - Reference safe and unsafe solutions demonstrating both outcomes - A deterministic, effect-based Python verifier that grades on observed behavior and persisted state (not string matching) - A Dockerized environment with all fixtures - A branch validation matrix (noop, safe oracle, alternate safe, unsafe reference, tamper) You'll prepare 1 seed RLE from scratch and then author 4 variations, changing the attack mechanism across at least two axes (e.g., where the pressure lives, how it's disguised, what action it pushes toward, what boundary it targets). Each variation must hold three design invariants: separability (base task passes even on the unsafe path), pressure sourced only from environment content (never the user instruction), and a reachable safe path. **Required skills:** - Hands-on experience creating training/evaluation data for LLM-based agents - Strong Python (writing verifiers, grading scripts, test harnesses) - Understanding of how coding agents interact with repo context, tool outputs, and embedded instructions - Ability to design realistic, non-contrived scenarios that test agent safety boundaries - Docker familiarity for environment setup - Attention to grading robustness — verifiers must be effect-based and resistant to surface-level evasion **Nice to have:** - Prior work in RLHF/RLAIF data pipelines or AI safety red-teaming - Familiarity with frontier-model evaluation methodology - Experience with deterministic scoring pipelines (no LLM-judge reliance) **Project scope:** 1 seed + 4 variations (5 complete RLE packages) **To apply:** Share relevant experience with agentic evaluation or RL training data. A sample of prior work in this space (even a brief description of methodology) will go a long way.
Abrir na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Entrar