MLOps platform Wow moment demo
Бюджет: $1500.0
FIXED /
⭐ 0.00 (0)
Romania
computer-vision, machine-learning, artificial-intelligence, python, data-science, artificial-neural-networks, deep-learning, algorithm-development, adaptive-algorithms
Предпочитана квалификация
- Тип изпълнител: Независим
- Локация: Europe, Americas
- Опит: Средно ниво
- Английски: Свободен
AI perception engineer + video producer: record a 4-minute product demo
We're looking for an expert CV/ML practitioner with a good taste for video, to record a demo of our product. That's a rare combination and it's deliberate: you need to understand the result to film it well, and you need to be able to cut a video that a skeptical senior engineer watches to the end.
We're building WeightsLab, a tool that shows what a model is learning at the level of individual training samples, so a person can act on it mid-training instead of waiting for the next run.
https://github.com/GrayboxTech/weightslab
The high-level vision, which everything below serves: better insight plus twenty minutes of human judgment beats countless automated sweeps.
---
The flow we have (and its already showcased as the first gif in our github repo)
1. Train 4 epochs on a road segmentation dataset (~8k driving images).
2. Inspect live per-sample signals BCE, Dice, combined loss.
3. The turn. Sort by loss ascending, not descending. The lowest-loss training samples are near-empty ground truth: an image that's 2% road is trivially easy, so low loss means low signal, not correctness. This inverts what every ML engineer expects, and it is the heart of the video.
4. Rewind the weights to the epoch-2 checkpoint same weights, same optimizer state, so anything that differs from here is the data and not the seed.
5. Discard the questionable 4–5% of training samples.
6. State the prediction out loud, before resuming: false positives will drop, false negatives will stay flat because we removed samples that taught the model nothing is here, not samples that taught it what a road looks like.
7. Resume training for one epoch.
8. The result: FP roughly halves. FN flat, exactly as called.
9. The control. The same number of samples dropped at random, from the same checkpoint. Nothing moves. This has to be in the video without it, the result is arguable.
The task, the order, the metrics and the modality are all open. 3D lidar detection and 2D object detection are live use cases for us, and a compelling result in either is worth more to us than a polished version of what we already have. What has to survive is the shape: a human sees something the aggregate metric hid, predicts what will change, changes the data, and is right within one epoch.
---
Deliverables
1. A 3–4 minute narrated video showing the product.
2. A 60–90 second cut for LinkedIn, from the same footage.
3. The transcript, as text.
4. Checkpoints, and the dataset if a custom one is used.
5. The scripts or commands needed to reproduce every number that appears on screen.
---
What great looks like
Good means the video is clear, every number on screen is reproducible, and it looks like an engineer showing you something rather than an ad no stock music, no motion graphics, no logo stings.
Great means it contains two specific sentences, and they land.
1. A sentence that makes a team install it. It names something true about their dataset that they can't currently see and would want to check within the hour a claim about their data, not about our product. Place it early; it's the hook, not the conclusion. The test: if the reaction is "that's nice," it failed. If it's "wait, how much of my training set looks like that?", it worked.
Fails: "WeightsLab gives you deep visibility into your training data." That's a sentence about us.
2. A sentence that connects what just happened to why it matters generally. The video shows one result, on one dataset, on one task. This sentence says what class of thing that result is an instance of, without overclaiming it into meaninglessness. Place it at the end, once the viewer has seen the evidence. The shape to capture: a person saw what the aggregate metric was hiding, predicted what would change, and was right in one epoch the alternative was more runs.
Fails: "This shows how WeightsLab can improve your models." Unfalsifiable, and true of every tool ever built.
Both sentences should appear verbatim and called out in the transcript. They're the two lines we'll reuse everywhere else, and they're most of what we're paying for.
---
Working conditions
This is a live product under active development, not a canned recording. Expect slow-loading panels, occasional refreshes, and multiple takes. Training runs about 4.5 minutes on screen and the edit has to handle that rather than sit in it. We'll give you the full known-issues list on day one. If working around rough edges frustrates you, this isn't the job.
Timeline and budget: couple of days, 1500$. Tell us your rate and your availability.
---
To apply
Letter of intent outlining the experience or insights that makes you the right fit.
Отвори в Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Вход