Senior LLM Fine-Tuning Engineer — Open-Weight Coding Model
Budżet: $30.0 - $60.0
HOURLY / FULL_TIME
⭐ 4.76 (33)
United States
machine-learning-model
Preferowane kwalifikacje
- Doświadczenie: Ekspert
Senior LLM Fine-Tuning Engineer — Open-Weight Coding Model
We are looking for a senior ML/LLM engineer to fine-tune an open-weight coding model for real-world software engineering tasks.
This is not a generic chatbot fine-tuning project.
Our goal is to improve an open coding model on repository-level tasks such as:
* bug fixing
* feature implementation
* codebase understanding/file localization
* writing and repairing tests
* multi-file refactoring
* build/dependency failures
* GitHub PR/review fixes
* execution-error recovery
Initial focus will be Python + TypeScript.
What you will build
You will:
* benchmark suitable open-weight coding models such as Qwen3-Coder variants
* build a dataset from executable software-engineering tasks
* work with SWE-bench / SWE-Gym / SWE-smith style environments
* generate and filter coding-agent trajectories
* fine-tune using SFT + LoRA/QLoRA or an appropriate alternative
* train the model to recover from failed tests/tool execution
* optionally experiment with rejection sampling, DPO, verifier training or RL
* benchmark the tuned model against the untouched base model
* deploy the final model behind an OpenAI-compatible endpoint
Success will be measured using actual test-passing repository tasks, not training loss.
Required experience
Strong experience with:
* PyTorch / Hugging Face
* LLM fine-tuning
* LoRA / QLoRA / PEFT
* distributed GPU training
* coding LLMs
* Docker/sandboxed execution
* SWE-bench or similar coding-agent benchmarks
* vLLM/SGLang
Experience training coding agents using executable rewards is a major plus.
Deliverables
1. Reproducible baseline
2. Dataset-generation pipeline
3. 5K–15K+ high-quality verified training examples
4. Fine-tuned model/checkpoint
5. Automated evaluation harness
6. Base-vs-fine-tuned benchmark report
7. Inference endpoint
8. Full source code and documentation
When applying, please answer
1. Which open-weight coding model would you start with today and why?
2. Have you fine-tuned a coding LLM before? Send benchmark/results.
3. How would you prevent SWE-bench/data contamination?
4. Would you start with SFT, DPO or RL, and why?
5. How would you create 10,000 executable repo-level coding tasks?
6. Describe your GPU/training setup for a 30B+ coding model.
Please do not apply if your experience is limited to OpenAI API prompting or basic Hugging Face tutorials.
We are looking for someone who has actually trained and evaluated models.
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj