AI / LLM Research Engineer
Budżet: $30.0 - $50.0
HOURLY / PART_TIME
⭐ 5.00 (2)
GBR
chemistry, biology, research-papers, economics, artificial-intelligence, machine-learning, data-science, data-analysis, artificial-neural-networks, natural-language-processing, data-mining, deep-learning
Preferowane kwalifikacje
- Typ talentu: Niezależny
- Doświadczenie: Ekspert
- Angielski: Biegły
- Job Success: 90%+
- Preferowany Rising Talent
- Min. zarobki: $100+
This is a time-boxed research and evaluation spike (~1 week), scoped as a standalone deliverable. It can run ahead of or alongside a longer engagement.
The engagement
Contract, remote. ~1 week, fixed scope.
Some working-hours overlap with the UK.
Day rate or fixed price — tell us what you prefer.
What we want to know
Can a different model — an open-weight model on Bedrock, or a small custom fine-tune — meaningfully improve quality at the same or lower cost per call? (Cost is already ~1¢ per story, so the bar is quality, not spend.)
What would it take to move part of this on-device on iOS and Android — especially the multimodal caption step — and what does that cost us in quality, engineering, and maintenance?
Frontier hosted models (GPT, Claude Opus/Sonnet) are out of scope as production options; we use one only as a quality-reference ceiling.
The work
Understand the current pipeline (structured photo metadata in → JSON story/caption out).
Build a small eval harness and a dataset of real, anonymised examples, plus a human-written “gold” reference set.
Benchmark a shortlist of candidates — open-weight LLMs for the story, open vision models for the caption — using an LLM-as-judge plus a short human blind rank.
Produce a $-per-1,000-calls and latency table from measured token counts.
Assess the on-device options (Apple Foundation Models, Gemini Nano / ML Kit, ExecuTorch / llama.cpp) against the same references.
Package the work for reproducibility: a documented, re-runnable harness and a written record of method, candidates, prompts, parameters, and results, so anyone on the team can re-run the comparison or extend it later.
Deliver a written go / no-go recommendation per feature, with effort estimates.
You should have
Hands-on experience evaluating and shipping LLM features — not just prompting, but structured evaluation, LLM-as-judge, and cost/latency analysis
Familiarity with the current open-weight landscape (Qwen, Llama, Mistral, DeepSeek, GLM, and vision models) and where to find reliable benchmarks
AWS Bedrock experience, or equivalent with another hosted-inference platform
Comfort reading a Python codebase and writing a clean eval script
Strong communication and collaboration skills for working in an interdisciplinary team, and a habit of documenting your method so results can be reproduced and built on
Bonus: LoRA / small-model fine-tuning; on-device inference on mobile (Core ML, ExecuTorch, MLC, llama.cpp); multimodal models
Deliverables
Eval harness and dataset, committed to our repo, with a README that lets someone else re-run it from scratch
A results document: method, candidates, prompts and parameters used, scored comparison per feature, cost/latency, sample outputs, multilingual-capability notes, and an on-device feasibility summary
A clear recommendation and next-step effort estimate
To apply
Send your rate, availability, and a short example of a model evaluation or comparison you’ve done — ideally something where you had to weigh quality against cost.
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj