← İşler

Data Engineer

Bütçe: - HOURLY / FULL_TIME ⭐ 3.89 (11) United Kingdom

hadoop, sas, computer-networking, sap

Tercih edilen nitelikler

  • Deneyim: Orta
UK Job Postings Dataset I run a UK contractor job board and an accompanying market-intelligence platform that publishes rate and hiring-volume data for contractors. I've bought a large historical dataset of UK job postings to power it, and I'm looking for someone to build the pipeline and classification rules that turn the raw feed into something the site can query. The front end is already built; this is the layer underneath it. What I have. A commercial feed of UK job postings, one gzip JSON Lines file per day covering every day from January 2020 to present — roughly 13–16k postings a day, 30m+ rows in total. Each record holds a raw job title, the full body text, company, location, salary and a set of tag objects. Each daily file contains only postings first seen that day, so there's no overlap between days and no re-scrape. What I want from it. A clean posting-level table plus the classification rules that make it usable: contract versus permanent, a trustworthy normalised rate, job title mapped to my existing taxonomy of 396 terms, a proper UK city and region hierarchy, seniority, and IR35 status. On top of that, materialised aggregates with minimum-sample thresholds, which is what my front end reads from. The front end is already built, this is the layer underneath it. How it's delivered. The vendor drops the daily files into AWS, so ingest needs to pull from there on a schedule and be idempotent, resumable, and incremental rather than a full rebuild each time. Schedule can be month;y. How I think we should handle it. DuckDB with Python for the ingest and rule logic, running on a single machine. I'm open to argument on that, but it's the working assumption. Potential issues. Salary numbers can't be trusted and the salary text is the real source — one major board returns its own filter buckets as though they were real values, so "Competitive salary" comes back as £10,000 to £100,000. A contract-type field exists and is reliable when populated, but only around 11% of postings carry a value that distinguishes contract from permanent; the rest are unlabelled rather than permanent, so status has to be inferred from title and body text without defaulting the majority the wrong way. The vendor's own normalised job name and industry fields are too coarse to classify from, the record schema varies between rows, and the documentation doesn't match what's actually delivered.
Upwork'te aç

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Giriş yap