ML/Data Engineer — Audience Intelligence Platform (Python, BigQuery, dbt, LLM APIs, XGBoost)
Költségvetés: $40.0 - $80.0
HOURLY / NOT_SURE
⭐ 0.00 (0)
United Kingdom
python, bigquery, data-science, machine-learning, data-analysis, google-cloud-platform
Előnyben részesített képesítések
- Tapasztalat: Középhaladó
We're a UK media & IP company building an internal decision-intelligence platform that analyses creator audiences at scale — engagement patterns, cross-platform presence, fandom behaviour — to support investment decisions in creator-led IP. The commercial thesis is confidential; the engineering is described fully below and an NDA covers the rest after hire.
We're looking for a hands-on ML/data engineer to build the platform end to end as a contractor, working directly with the founder and our Head of IP Intelligence. The stack and architecture are already specified in a detailed PRD with numbered requirements and acceptance criteria — you'll be building to a real spec, not guessing at scope.
What you'll build (four workstreams):
Collection pipelines — YouTube Data API connector (channels, video stats time-series, comment corpora) with quota management and audit logging; a video transcription pipeline (self-hosted Whisper or managed API — your recommendation); scheduled public-data snapshots (cross-platform follower counts, Google Trends, Wikipedia pageviews); all landing in BigQuery.
Feature engineering — a versioned signal library in dbt: engagement ratios benchmarked against size-matched cohorts, LLM-based comment classification at corpus scale (batch API, pinned prompt/model versions), transcript-based content analysis, and point-in-time correct features (every feature computable "as of" an arbitrary historical date — this is a hard requirement, it powers our backtesting).
Scoring & backtesting — a calibrated gradient-boosted scoring model (XGBoost or LightGBM) over the engineered features with confidence intervals and driver attribution; an evaluation harness that tests the score against baseline heuristics on a historical corpus (AUC, precision-at-k); LLM-generated evidence reports from templates; per-run cost tracking with hard monthly spend caps.
Review UI & decision log — an internal tool (Retool or lightweight web app, your call) showing scores, evidence, comparator cohorts and time-series per candidate, with a decision-capture form writing an immutable, append-only decision log to the warehouse.
Engagement shape: ~60–80 working days across 6–9 months, front-loaded (roughly full-time for the first 6–8 weeks, then 2–3 days/week). Fully remote; we're UK-based, so at least 3–4 hours of overlap with UK working hours. Start within 2–3 weeks.
How we work: everything in company-owned GitHub and GCP from day one; full IP assignment on payment (non-negotiable); CI with dbt tests on every model; documentation is a deliverable, not an afterthought — the acceptance bar is that a second engineer can navigate the platform cold. Milestone-based acceptance against the PRD checklist.
You're a fit if you have:
3+ years building data/ML pipelines in Python with a cloud warehouse (BigQuery strongly preferred) and dbt in production
Real experience with LLM APIs in batch analytical workflows (classification, extraction, structured output) — including prompt versioning and cost control, not just chat integrations
Classical ML depth: gradient boosting, calibration, honest evaluation methodology (you can explain why precision-at-k matters more than accuracy here)
API integration craft: pagination, rate limits, backoff, idempotency — you've fought the YouTube Data API or similar and won
Evidence of documentation quality (a repo, dbt docs site, or technical writing you can share)
Nice to have: point-in-time/as-of feature engineering (feature stores, backtesting systems, quant background welcome); scraping platforms (Apify or similar); Retool; social/creator/audience data experience.
Not a fit: deep-learning research profiles (there's no model training beyond classical ML here), pure BI/dashboard specialists, or agencies proposing teams — we want one accountable individual.
To apply, answer the screening questions below. We'll shortlist within a week, hold one 60-minute technical conversation (walking through your past work and our architecture — no take-home test), check one reference, and start on a paid pilot milestone.
Megnyitás Upworkön
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Bejelentkezés