← Live-стрічка

Senior LLM Fine-Tuning Engineer — Open-Weight Coding Model

Бюджет: $30.0 - $60.0 HOURLY / FULL_TIME ⭐ 4.76 (33) United States

machine-learning-model

Бажана кваліфікація

  • Досвід: Експерт
Senior LLM Fine-Tuning Engineer — Open-Weight Coding Model We are looking for a senior ML/LLM engineer to fine-tune an open-weight coding model for real-world software engineering tasks. This is not a generic chatbot fine-tuning project. Our goal is to improve an open coding model on repository-level tasks such as: * bug fixing * feature implementation * codebase understanding/file localization * writing and repairing tests * multi-file refactoring * build/dependency failures * GitHub PR/review fixes * execution-error recovery Initial focus will be Python + TypeScript. What you will build You will: * benchmark suitable open-weight coding models such as Qwen3-Coder variants * build a dataset from executable software-engineering tasks * work with SWE-bench / SWE-Gym / SWE-smith style environments * generate and filter coding-agent trajectories * fine-tune using SFT + LoRA/QLoRA or an appropriate alternative * train the model to recover from failed tests/tool execution * optionally experiment with rejection sampling, DPO, verifier training or RL * benchmark the tuned model against the untouched base model * deploy the final model behind an OpenAI-compatible endpoint Success will be measured using actual test-passing repository tasks, not training loss. Required experience Strong experience with: * PyTorch / Hugging Face * LLM fine-tuning * LoRA / QLoRA / PEFT * distributed GPU training * coding LLMs * Docker/sandboxed execution * SWE-bench or similar coding-agent benchmarks * vLLM/SGLang Experience training coding agents using executable rewards is a major plus. Deliverables 1. Reproducible baseline 2. Dataset-generation pipeline 3. 5K–15K+ high-quality verified training examples 4. Fine-tuned model/checkpoint 5. Automated evaluation harness 6. Base-vs-fine-tuned benchmark report 7. Inference endpoint 8. Full source code and documentation When applying, please answer 1. Which open-weight coding model would you start with today and why? 2. Have you fine-tuned a coding LLM before? Send benchmark/results. 3. How would you prevent SWE-bench/data contamination? 4. Would you start with SFT, DPO or RL, and why? 5. How would you create 10,000 executable repo-level coding tasks? 6. Describe your GPU/training setup for a 30B+ coding model. Please do not apply if your experience is limited to OpenAI API prompting or basic Hugging Face tutorials. We are looking for someone who has actually trained and evaluated models.
Відкрити замовлення

AI-чернетка відгуку

Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.

Увійдіть, щоб згенерувати AI-чернетку.

Увійти