Video Evaluation, Verification & Data Engineer
Бюджет: $80.0 - $140.0
HOURLY / NOT_SURE
⭐ 4.81 (59)
United States
machine-learning, python, etl-pipelines, computer-vision, data-science
Предпочтительная квалификация
- Опыт: Эксперт
We are hiring a senior evaluation and data engineer for a focused 60-day generative-video evidence sprint. You will build the data controls, evaluation harness, blinded review process, verifier calibration, and evidence package used to determine whether model or conditioning changes produce real improvement.
This is not a general data-labeling role, a compliance-writing role, or a request for someone to score a few polished demos. We need a hands-on engineer who can make video-model results reproducible, comparable, auditable, and difficult to game.
What you will own:
- Create versioned dataset manifests with stable IDs, checksums, source lineage, split assignments, permitted-use fields, and exception tracking.
- Define clean development, validation, and sealed holdout sets, then detect duplicates, contamination, identity leakage, and post-result test changes.
- Build a frozen evaluation harness for identity consistency, appearance continuity, scene and action adherence, temporal stability, critical failures, latency, and cost.
- Combine deterministic hard gates, automated scoring, and blinded human review without hiding disagreements between them.
- Design reviewer instructions, adjudication rules, sampling plans, calibration cases, and inter-rater agreement checks.
- Preserve immutable generation manifests, evaluator versions, scores, reviewer decisions, exceptions, and evidence links for every evaluated asset.
- Calibrate verifier thresholds against human judgments and document false positives, false negatives, uncertainty, and known blind spots.
- Produce decision-ready before-and-after evidence and a complete technical handover, including code, schemas, runbooks, dashboards, and evaluation documentation.
Operating boundaries:
- Proprietary data and generated artifacts remain inside a controlled private environment.
- Rights, consent, provenance, permitted use, retention, and revocation state must be represented in the data and evidence workflow.
- The implementation team cannot change test cases, thresholds, routes, or evaluator versions after seeing results without creating a new declared experiment.
- Missing evidence, rights uncertainty, holdout leakage, or incomplete manifests must cause a stop, quarantine, or explicit exception, not a silent pass.
- You will build and operate the evidence machinery, but final acceptance remains with a designated internal authority.
Required experience:
- Built production or research-grade evaluation pipelines for generative video, computer vision, multimodal models, image generation, or a closely related domain.
- Strong Python and data-engineering ability, including dataset versioning, manifests, validation, experiment tracking, and reproducible analysis.
- Experience designing held-out evaluation, blinded human review, annotation quality controls, adjudication, and regression testing.
- Able to distinguish model quality, data quality, reviewer disagreement, system failure, and policy failure in the final evidence.
- Comfortable challenging impressive demos when the underlying comparison, data split, or evidence is weak.
Strong pluses:
- Identity or face consistency evaluation, temporal video metrics, scene or action adherence, video-quality assessment, or production-usability scoring.
- Experience with tools such as FiftyOne, CVAT, Label Studio, W&B, MLflow, DVC, or equivalent internal systems.
- Data provenance, consent, media rights, model governance, or high-integrity evaluation work.
- Designed evaluation gates that stopped a bad model or system release.
To apply, answer every question below. Applications that skip them will not be reviewed.
1. Describe the most relevant generative-model or computer-vision evaluation system you personally built. What data did it cover, what did you own, and what release or research decision did it support?
2. How would you test character identity and temporal consistency without allowing the implementation team to tune against the holdout set?
3. Describe one case where leakage, reviewer bias, a misleading metric, or an evaluation bug changed your initial conclusion.
4. How would you combine automated scores and blinded human review when they disagree?
5. Share one sanitized artifact you can walk through live, such as evaluation code, a dataset manifest, reviewer rubric, dashboard, calibration report, or experiment record.
6. Can you commit at least 30 hours per week for 60 days, and will you personally perform the core work?
Открыть заказ
AI-черновик отклика
Короткий текст отклика для копирования в оффер: интерес + готовность работать.
Войдите, чтобы сгенерировать AI-черновик.
Войти