Python + LLM Engineer — Build Autonomous Agents for a Live Freight-Tech Platform
Orçamento: $35.0 - $80.0
HOURLY / NOT_SURE
⭐ 4.88 (3)
United States
python, web-programming
Qualificações preferidas
- Experiência: Intermédio
We run a track-and-trace platform used by freight brokerages to keep tabs on live truckload shipments. We’ve just shipped the first two autonomous agents into production and they work. We want one excellent engineer to own that surface and take it from “two agents that work” to “a fleet of agents that customers pay for and trust.”
This is not a research project and it’s not a prototype. The agents email real customers about real trucks today. Everything you build ships within days of writing it.
What the agents actually are
Our agent service is a standalone Flask + Celery app with a deliberately small, generic core and a per-agent-type extension seam. Two shapes exist:
Proactive agents — a scanner sweeps the live load board, opens a “goal” per shipment that needs attention, and a dispatcher ticks each goal through a state machine on a poll clock. The first one watches for trucks that are late to a pickup appointment and alerts the account manager. Idempotency is enforced at the database, not in code, so a retry or a racing worker can never double-send.
Reactive agents — anyone at an enrolled customer emails a dedicated status address on our domain with a load number and gets back a real answer: where the truck is, whether it’s on time, what happens next. It runs an authentication and safety gate ladder (envelope sender verification, SPF, loop suppression, rate limiting, tenant scoping), uses an LLM for exactly one narrow job — pulling the reference number and the intent out of free-form email — and computes every fact itself. Replies to the answer feed back into that customer’s agent memory.
The near-term roadmap is roughly: move inbound processing onto the worker queue, add outbound driver SMS and voice as agent steps, build agents for proof-of-delivery collection, ETA confirmation and detention risk, wire per-agent observability so we can prove value to a customer with numbers, and make the whole thing configurable enough that we can turn a new agent on for a new brokerage without a deploy.
What you’d be doing
Designing and shipping new agent types end to end — the scanner or trigger, the state graph, the config schema, the templates, the tests, the migration, the deploy.
Connecting agents to the rest of the product: the admin UI in the main app, our outreach channels (email, SMS, AI voice), the workflow dashboards, and per-tenant configuration.
Making the LLM parts reliable. Our hard-won lesson so far is that the model is rarely the problem — the win comes from letting the database validate what the model produced, keeping prompts narrow, and never letting model output influence tenancy or authorization.
Instrumenting everything. Every reason an agent didn’t act is a distinct, countable outcome in our data. That’s what makes a Monday morning review quantitative instead of anecdotal.
Owning correctness. A wrong ETA in a customer’s inbox is worse than no email at all, and we build accordingly: facts are computed, judgements are read from a single source of truth, and the agent degrades honestly when it doesn’t know something.
Stack
Python 3.13, Flask, Celery (Beat + workers), PostgreSQL (primary + read replica, raw SQL and SQLAlchemy, hand-written migrations — no Alembic), Redis/Valkey, OpenAI (with instructor for structured output), SendGrid (outbound + Inbound Parse), Twilio, Vapi, Docker, AWS (ECS Fargate, ECR, ALB), GitHub Actions, Terraform/Terragrunt.
You do not need to have used all of these. You do need to be genuinely strong in Python and comfortable in a codebase where SQL, queues and production infrastructure are all part of the job.
Who we’re looking for
Required
Deep, practical Python. You’ve built and operated background-job systems — you know what at-least-once delivery means for your code and you design for it.
Real SQL. You can read a query plan, you know why an index isn’t being used, and you’re comfortable writing raw SQL against a large table without taking down a read replica.
You’ve shipped something with an LLM in it that had to be right, not just impressive — and you have opinions about where the model belongs in the design and where it absolutely doesn’t.
Production ownership. You’ve deployed, watched logs, found the thing that broke, and fixed it.
Excellent written English. Most of our coordination is written and async, and our internal docs are a real artifact we maintain.
Reliable and reachable. Not always-on — but when something is broken in production and we message you, we need to hear back the same working day, and we need to be able to hand you the keys while our CTO is on a plane.
Strong plus
Logistics, freight, supply chain or TMS experience (McLeod, Turvo, Revenova, Mastery, Edge — or any of the ecosystem). This is a domain where the details matter enormously and someone who already knows what a stop, a leg, detention and a POD are will move much faster.
Email infrastructure at the protocol level — SPF/DKIM, MIME, threading via References headers, rendering HTML that survives Outlook.
AWS/ECS and infrastructure-as-code.
What won’t work here
Agent-framework enthusiasm without systems fundamentals. We are not looking for someone to bolt LangChain onto this. Our core is a few hundred lines and stays that way on purpose.
Anyone who needs a fully specified ticket to start. You’ll get context, constraints and a goal.
Anyone who marks work “done” without having run it.
How we work
Small team, no ceremony, direct access to the founder. You’ll be in the same chat as the person making product decisions.
Fast. Features go from conversation to production in days, not sprints.
Written-down. We keep living documentation of how each subsystem actually behaves, including its gotchas and its known-broken parts. When you find one stale, you fix it in the same change.
Honest about tradeoffs. There are things in this codebase that are deliberately deferred, and they’re written down as deferred rather than pretended away. We expect the same from you.
We care about latency and about correctness, and we’ll push back on cleverness that costs either one.
Terms
5–10 hours a week to start, growing from there. We’re deliberately starting small so we can both find out fast whether this is a fit. The scope of the agent work is much bigger than 10 hours, and we’d like to get there with you.
Any timezone for the right person. Nearly all of our coordination is written and async. We’ll want the occasional live call and we’ll find a time that works — we’re not going to ask you to keep US hours.
Coverage while our CTO is away. Once you’re up to speed, part of this role is being the person who watches the agents when he’s travelling or on vacation: checking that scans and dispatches are running, triaging what breaks, and making sure nothing wrong goes out to a customer. It’s occasional, planned in advance, and it is not a pager in the middle of your night — but it is real responsibility, and it’s a big part of why we’re hiring someone for the long term rather than per-project. We’ll build up to it; nobody covers a system they don’t understand yet.
This is a long-term role for the right person. If the work is good, we want to move to a larger ongoing engagement and hand you this whole surface area to own. We’re saying that up front because we mean it, not as a carrot.
Abrir na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Entrar