← Live-лента

Python/AI Engineer — Review & Harden a Production LLM Agent (Ongoing, Part-Time)

Бюджет: $10.0 - $12.0 HOURLY / FULL_TIME ⭐ 4.69 (7) India

python, machine-learning, data-analysis, artificial-intelligence, data-science, api

Preferred qualifications

  • Experience: Expert
Python/AI Engineer — Review & Harden a Production LLM Agent (Ongoing, Part-Time) Description What this is We run a production LLM agent that answers operational questions for business users — it pulls from a live database, runs calculations, and returns numeric answers people make decisions on. It works. Roughly half of production traffic now goes through fast deterministic paths at sub-500ms with a clean abort rate. The other half is where we need you. The actual problem We have a strong Python developer building this. What we don't have is a second person whose job is to ask "is this actually correct, and can we prove it?" before anything ships. That gap has produced a repeatable set of failures: Deterministic work handed to the model. The agent has hand-summed a 32-row list and returned a total that was off by ~$27. Arithmetic that belongs in code went to a language model. Loop budget with no hard terminal. In one case the correct answer was already computed at step 2. The loop kept going, made an unnecessary call, and reported a number that was ~40x wrong. The system prompt already said "stop once you have a routed answer" — prompt wording didn't hold the line. Guardrails nobody calibrated. A self-check we added to catch wrong numbers now false-blocks a majority of our key metrics because of a 0–1 vs 0–100 scale mismatch nobody tested against real data. A success metric that lies in both directions. Our current "did this turn succeed" flag marks complete, correct answers as failures, and marks real refusals as fine. We have no trustworthy scoreboard. Things built and never proven live. An evaluation harness that has never once been run end-to-end. A change switched on in production while its own notes still said "untested." None of these are exotic. They're the ordinary failure modes of agentic LLM systems that nobody is reviewing at depth. What you'd actually do You are not the primary implementer — our developer stays hands-on-keyboard. You are the person who decides what gets fixed first, reviews the work before it merges, and can be held to a number. Concretely, week to week: Read PRs at real depth — async Python, control flow, tool-call loops, error paths — and block merges that can't be proven correct Decide which work belongs in deterministic code vs. the model, and enforce that boundary in code (hard terminals, not prompt wording) Recalibrate our guardrails against real production data instead of assumed ranges Define and stand up an honest success metric, then run it — a baseline computed on real traffic, not a harness that exists but never executes Own a small number of headline metrics: answer correctness, abort rate, p50/p95 latency, cost per resolved turn Write short, direct weekly notes: what moved, what's still broken, what you recommend next Who this fits 4+ years Python, and you're genuinely comfortable in async code — not "I've used asyncio," but you can spot a bug in a concurrent tool-call loop during review You've shipped and maintained an LLM agent in production, not just a demo. You know what a loop budget costs and why retries get expensive You've built or repaired an evaluation setup and can explain how you knew your numbers were real You can read someone else's code and say "this won't hold" with a specific reason You write clearly and briefly in English. Most of this role is written judgment Who this does not fit If your instinct for fixing a wrong number is to rewrite the system prompt, this isn't the role If you've only worked on greenfield demos, you'll find this frustrating If you need a fully specified ticket to start, this seat is the opposite of that Engagement Ongoing, part-time: 10–20 hours/week Overlap of at least 3 hours/day with 9am–6pm IST Starts with a paid 1-week trial: read the codebase, produce a prioritized remediation plan, and compute an honest baseline on our current traffic. If the plan is sharp, we continue indefinitely. Long-term relationship expected — we're not shopping for a one-off audit To apply Skip the template. Answer the screening questions. Applications that open with "I am a passionate full-stack developer" get archived unread. Skills Python, Asyncio, LLM, OpenAI API, Anthropic Claude, Prompt Engineering, AI Agent Development, Model Evaluation, Code Review, PostgreSQL Scope Ongoing project · More than 6 months · Part time (less than 30 hrs/week) · Intermediate · Hourly · $10.00–$12.00/hr · Worldwide · Freelancers only
Открыть заказ

AI-черновик отклика

Короткий текст отклика для копирования в оффер: интерес + готовность работать.

Войдите, чтобы сгенерировать AI-черновик.

Войти