← Live-стрічка

AI / LLM Research Engineer

Бюджет: $30.0 - $50.0 HOURLY / PART_TIME ⭐ 5.00 (2) GBR

chemistry, biology, research-papers, economics, artificial-intelligence, machine-learning, data-science, data-analysis, artificial-neural-networks, natural-language-processing, data-mining, deep-learning

Бажана кваліфікація

  • Тип виконавця: Незалежний
  • Досвід: Експерт
  • Англійська: Вільний
  • Job Success: 90%+
  • бажаний Rising Talent
  • Мін. заробіток: $100+
This is a time-boxed research and evaluation spike (~1 week), scoped as a standalone deliverable. It can run ahead of or alongside a longer engagement. The engagement Contract, remote. ~1 week, fixed scope. Some working-hours overlap with the UK. Day rate or fixed price — tell us what you prefer. What we want to know Can a different model — an open-weight model on Bedrock, or a small custom fine-tune — meaningfully improve quality at the same or lower cost per call? (Cost is already ~1¢ per story, so the bar is quality, not spend.) What would it take to move part of this on-device on iOS and Android — especially the multimodal caption step — and what does that cost us in quality, engineering, and maintenance? Frontier hosted models (GPT, Claude Opus/Sonnet) are out of scope as production options; we use one only as a quality-reference ceiling. The work Understand the current pipeline (structured photo metadata in → JSON story/caption out). Build a small eval harness and a dataset of real, anonymised examples, plus a human-written “gold” reference set. Benchmark a shortlist of candidates — open-weight LLMs for the story, open vision models for the caption — using an LLM-as-judge plus a short human blind rank. Produce a $-per-1,000-calls and latency table from measured token counts. Assess the on-device options (Apple Foundation Models, Gemini Nano / ML Kit, ExecuTorch / llama.cpp) against the same references. Package the work for reproducibility: a documented, re-runnable harness and a written record of method, candidates, prompts, parameters, and results, so anyone on the team can re-run the comparison or extend it later. Deliver a written go / no-go recommendation per feature, with effort estimates. You should have Hands-on experience evaluating and shipping LLM features — not just prompting, but structured evaluation, LLM-as-judge, and cost/latency analysis Familiarity with the current open-weight landscape (Qwen, Llama, Mistral, DeepSeek, GLM, and vision models) and where to find reliable benchmarks AWS Bedrock experience, or equivalent with another hosted-inference platform Comfort reading a Python codebase and writing a clean eval script Strong communication and collaboration skills for working in an interdisciplinary team, and a habit of documenting your method so results can be reproduced and built on Bonus: LoRA / small-model fine-tuning; on-device inference on mobile (Core ML, ExecuTorch, MLC, llama.cpp); multimodal models Deliverables Eval harness and dataset, committed to our repo, with a README that lets someone else re-run it from scratch A results document: method, candidates, prompts and parameters used, scored comparison per feature, cost/latency, sample outputs, multilingual-capability notes, and an on-device feasibility summary A clear recommendation and next-step effort estimate To apply Send your rate, availability, and a short example of a model evaluation or comparison you’ve done — ideally something where you had to weigh quality against cost.
Відкрити замовлення

AI-чернетка відгуку

Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.

Увійдіть, щоб згенерувати AI-чернетку.

Увійти