ML engineer — object detection + LLM agents for inspection software
Buget: $20.0 - $50.0
HOURLY / PART_TIME
⭐ 5.00 (3)
USA
pytorch, machine-learning, python, tensorflow, robotics, computer-vision
Calificări preferate
- Experiență: Intermediar
Also attached the same do cofr better readablity.
ML engineer — object detection + LLM agents for inspection software
Fixed price · Phase 1 estimated 40–50 hours · [BUDGET]
What we do
We build automated inspection software for facility infrastructure. A robot captures imagery of industrial equipment
on a schedule, and our system turns that into inspection records that our clients hand to auditors and insurers.
These records are compliance documents. That shapes everything below and it is the most important thing to
understand about this role.
The work — two parts
Part 1 — object detection. Train and deliver a detection model for objects in an indoor industrial environment.
Equipment cabinets, doors and latches, cable bundles, obstructions on the floor, safety equipment, labels and
nameplates.
Part 2 — LLM agent layer. An agent that produces narrative summaries from structured inspection data, and answers
questions against a set of inspection records.
Phase 1 is one detection class, delivered end to end. Data pipeline, training, evaluation, exportable inference. If that
lands well, we extend class by class, and the agent work follows.
The constraint that makes this role unusual
Our output goes to auditors. A model that is right most of the time is not sufficient unless it also knows when it is
not sure.
Concretely:
• Every detection carries a calibrated confidence, and the system must be able to abstain rather than guess. A
confident wrong detection is worse for us than an abstention — that is not a slogan, it is the requirement
• Model outputs are findings, not verdicts. The model says what it sees. Whether that constitutes an issue is
decided downstream by deterministic rules. Do not build the classification into the model
• The agent never produces a number. It writes prose about values that come from the structured record. If a
client asks why a report says something, the answer must trace to data, not to a model
If that framing sounds restrictive rather than interesting, this is not the right role for you. It is the defining
characteristic of the work.
Part 1 — detection, in detail
You will receive: an anonymised image set, a class list with written definitions, a seed annotation set showing our
labelling standard, and the output schema.
You will deliver: a trained model, inference code, and an evaluation.
Requirements:
1. Data pipeline. Ingest, augment, split, train, evaluate. Reproducible from a config.
2. The split is by capture session, not random. Our imagery comes from continuous robot runs, so consecutive
frames are near-identical. A random split leaks near-duplicates into validation and produces a number that means
Upwork posting — ready to publish
Page 2 of
nothing. Whole sessions to train, whole sessions to validate, at least one session held back and never touched until
final evaluation.
3. Small object performance. Some target classes are small — latches, labels — and often near the frame edge. This
is the hard part and we know it.
4. Calibrated confidence. Raw detector confidence is not a probability and we need it to behave like one. Tell us how
you approach that.
5. Edge deployment. Inference runs on an NVIDIA Jetson Orin NX. TensorRT, FP16. INT8 only with a documented
accuracy comparison.
6. Versioned. Every detection records which model produced it. Historical outputs must be traceable.
Part 2 — agents, in detail
Only after Part 1, and possibly a separate engagement.
Narrative generation. Given a structured inspection record, produce readable prose summarising it. Every figure and
every claim traceable to a field in the input.
Question answering. Over a set of inspection records. Grounded, with citations back to specific records.
The requirement: the agent must be structurally incapable of stating a number that is not in the data. Tell us how
you would guarantee that rather than hope for it.
Skills — must have
Object detection, end to end. Not just training a model — the whole path. Data pipeline, augmentation, training,
evaluation, and getting it into something that runs.
PyTorch. Fluent, not tutorial-level.
A detection framework you know well. YOLO-family, Detectron2, MMDetection — we do not mind which. We mind
that you have taken one of them to production.
Edge deployment. TensorRT, ONNX, and quantisation. You should have measured what quantisation cost you rather
than assuming it was fine.
Evaluation you can defend. mAP at different IoU thresholds, precision-recall tradeoffs, per-class analysis. Knowing
why a headline number can be misleading matters more here than pushing it higher.
Confidence calibration. Temperature scaling, reliability diagrams, or an equivalent approach. Understanding that raw
detector confidence is not a probability is close to a hard requirement for this role.
Annotation practice. Labelling standards, inter-annotator agreement, and what noisy labels do to a model.
Reproducibility. Config-driven training, seeded runs, versioned outputs. Someone else has to be able to rerun what
you did and get the same result.
Python engineering. Readable, documented, packaged. This is deliverable code, not a notebook.
Clear written English. You will be documenting decisions and failure modes, and we will be reading them rather than
sitting next to you.
Skills — nice to have
Upwork posting — ready to publish
Page 3 of
LLM and agent work. Grounded generation, structured output, constrained decoding, RAG. If you have this as well as
the detection skills, say so — Part 2 is real work and we would rather not split it across two people.
Small object detection specifically. Tiling, high-resolution inference, anchor tuning. Some of our targets are small and
near the frame edge.
Fisheye or wide-angle imagery. Distortion models, undistortion, and what happens to detection accuracy at the
frame edges.
Thermal or radiometric imagery. Rare, and valuable to us. Especially understanding why standard RGB
augmentation is wrong on radiometric data.
Active learning or human-in-the-loop labelling. Our labelled set is small and will grow. Knowing how to choose what
to label next is worth a lot.
Robotics or ROS exposure. Our data comes off a mobile robot. Not required, but it shortens conversations.
Industrial, facility or infrastructure domain experience. Data centres, electrical equipment, industrial inspection.
Regulated or safety-critical software background. Automotive, medical, aerospace, anything where you had to prove
what you built rather than demonstrate it. That is unusually relevant here.
Tech
Python. PyTorch. Beyond that, your call — but justify it.
Detection is likely YOLO-family for edge inference, but if you think otherwise, say why. Agent work is framework-
agnostic.
No access to our systems is required or provided. You build against the data we supply. We integrate on our side.
Deliverables — Phase 1
1. Data pipeline, config-driven, reproducible
2. Trained model with exported weights
3. Inference code taking an image and returning detections with confidence
4. Evaluation on the held-out session, per-class metrics
5. TensorRT export with on-device timing
6. README covering training, evaluation, adding a class, and known failure modes
Acceptance — Phase 1
• Trains from the supplied data with a documented command
• Meets an agreed mAP threshold on the held-out session, set together once you have seen the data
• Abstains rather than guessing below a defined confidence threshold
• Runs within a latency budget on Orin NX, measured on-device
• Another engineer can add a class from the README alone
• Failure modes documented. Where it breaks, on what, and what would fix it
About us
Upwork posting — ready to publish
Page 4 of
Small engineering team, US-based, building robotics and inspection software. Clear scope, direct communication,
prompt payment. Ongoing work if this goes well — additional detection classes, the agent layer, and eventually
detection against 3D reconstructed geometry rather than raw frames.
To apply
Tell us briefly about relevant work — models you have trained and deployed, and any agent or LLM work if you have
done it. A link or a short description of something you shipped is worth more than a list of frameworks.
We are happy to discuss scope and approach
Deschide pe Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Autentificare