← Zakázky

ML/Data Engineer — Audience Intelligence Platform (Python, BigQuery, dbt, LLM APIs, XGBoost)

Rozpočet: $40.0 - $80.0 HOURLY / NOT_SURE ⭐ 0.00 (0) United Kingdom

python, bigquery, data-science, machine-learning, data-analysis, google-cloud-platform

Preferované kvalifikace

  • Zkušenost: Středně pokročilý
We're a UK media & IP company building an internal decision-intelligence platform that analyses creator audiences at scale — engagement patterns, cross-platform presence, fandom behaviour — to support investment decisions in creator-led IP. The commercial thesis is confidential; the engineering is described fully below and an NDA covers the rest after hire. We're looking for a hands-on ML/data engineer to build the platform end to end as a contractor, working directly with the founder and our Head of IP Intelligence. The stack and architecture are already specified in a detailed PRD with numbered requirements and acceptance criteria — you'll be building to a real spec, not guessing at scope. What you'll build (four workstreams): Collection pipelines — YouTube Data API connector (channels, video stats time-series, comment corpora) with quota management and audit logging; a video transcription pipeline (self-hosted Whisper or managed API — your recommendation); scheduled public-data snapshots (cross-platform follower counts, Google Trends, Wikipedia pageviews); all landing in BigQuery. Feature engineering — a versioned signal library in dbt: engagement ratios benchmarked against size-matched cohorts, LLM-based comment classification at corpus scale (batch API, pinned prompt/model versions), transcript-based content analysis, and point-in-time correct features (every feature computable "as of" an arbitrary historical date — this is a hard requirement, it powers our backtesting). Scoring & backtesting — a calibrated gradient-boosted scoring model (XGBoost or LightGBM) over the engineered features with confidence intervals and driver attribution; an evaluation harness that tests the score against baseline heuristics on a historical corpus (AUC, precision-at-k); LLM-generated evidence reports from templates; per-run cost tracking with hard monthly spend caps. Review UI & decision log — an internal tool (Retool or lightweight web app, your call) showing scores, evidence, comparator cohorts and time-series per candidate, with a decision-capture form writing an immutable, append-only decision log to the warehouse. Engagement shape: ~60–80 working days across 6–9 months, front-loaded (roughly full-time for the first 6–8 weeks, then 2–3 days/week). Fully remote; we're UK-based, so at least 3–4 hours of overlap with UK working hours. Start within 2–3 weeks. How we work: everything in company-owned GitHub and GCP from day one; full IP assignment on payment (non-negotiable); CI with dbt tests on every model; documentation is a deliverable, not an afterthought — the acceptance bar is that a second engineer can navigate the platform cold. Milestone-based acceptance against the PRD checklist. You're a fit if you have: 3+ years building data/ML pipelines in Python with a cloud warehouse (BigQuery strongly preferred) and dbt in production Real experience with LLM APIs in batch analytical workflows (classification, extraction, structured output) — including prompt versioning and cost control, not just chat integrations Classical ML depth: gradient boosting, calibration, honest evaluation methodology (you can explain why precision-at-k matters more than accuracy here) API integration craft: pagination, rate limits, backoff, idempotency — you've fought the YouTube Data API or similar and won Evidence of documentation quality (a repo, dbt docs site, or technical writing you can share) Nice to have: point-in-time/as-of feature engineering (feature stores, backtesting systems, quant background welcome); scraping platforms (Apify or similar); Retool; social/creator/audience data experience. Not a fit: deep-learning research profiles (there's no model training beyond classical ML here), pure BI/dashboard specialists, or agencies proposing teams — we want one accountable individual. To apply, answer the screening questions below. We'll shortlist within a week, hold one 60-minute technical conversation (walking through your past work and our architecture — no take-home test), check one reference, and start on a paid pilot milestone.
Otevřít na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Přihlásit