Upwork Job Post / Historical Competitor Price Data Extraction
Budget: $250.0
FIXED /
⭐ 4.98 (12)
South Africa
requests, html, regular-expressions
Preferred qualifications
- Experience: Intermediate
I need a freelancer to extract historical pricing data for a defined set of target products across four e-commerce websites, using the Internet Archive's Wayback Machine (and other web archives where needed) — going back as far as records exist.
This is a backward-looking / archival data project, not a live monitoring project. I want to reconstruct a price history timeline: what these sites charged for specific products, and how those prices changed, from whenever archive records begin up to today.
SCOPE
- 4 target websites, all within the same retail category. One of the four is no longer operating and is only reachable via web archives. Full URLs will be shared directly with the freelancer I engage (not published publicly for confidentiality reasons).
- A defined list of target products (specific branded product lines/SKUs within a single retail category). The full product list will be shared with the freelancer I engage.
WHAT I NEED DELIVERED
1. A coverage report — for each site, how far back does archive data actually go, and how complete/sparse is it? I want honesty here, not just a data dump — if a site has poor coverage, tell me.
2. A structured spreadsheet (CSV or XLSX) with columns: date, competitor, product_name, SKU/variant, price, source (URL + archive snapshot link), notes.
3. Missing data marked clearly — use "n/f" (not found) rather than estimating or interpolating. Every number in this file needs to be a real, sourced price.
4. Methodology notes — which archive sources you used (Wayback Machine, archive.today, Common Crawl, etc.) and why, per site.
5. Optional add-on, quote separately — a simple forward-looking scraper that continues logging these same prices going forward, so historical and future data connect into one continuous timeline.
TECHNICAL STARTING POINT
I already have a basic Python starter kit (Wayback Machine CDX API probe + scraper scaffold) I can share once engaged — you're welcome to use it as a base, rebuild it, or use your own approach/tools. What matters is the result: accurate, sourced, dated pricing data going as far back as possible.
REQUIRED SKILLS
- Python (requests/urllib + BeautifulSoup or lxml)
- Direct hands-on experience with the Wayback Machine CDX API specifically (not just general scraping)
- Regex / structured data extraction from HTML
- Data cleaning with pandas or similar
PREFERRED, NOT REQUIRED
- Experience with archive.today or Common Crawl as a fallback source
- E-commerce site experience (product schema, JSON-LD pricing markup)
- Prior work on price-history reconstruction or competitor price monitoring projects
IDEAL FREELANCER
- Comfortable telling me what's not possible, not just what is — I'd rather know a site has thin archive coverage than get a spreadsheet padded with guesses
- Can document their sourcing clearly enough that I can verify any given number myself
BUDGET & TIMELINE
Please quote as a fixed price for the full 4-site historical backfill, based on the scope above. If you believe an hourly arrangement makes more sense given uncertainty in archive coverage, propose a not-to-exceed estimate with milestones (e.g., "coverage report" as milestone 1, before committing to full extraction).
Target timeline: 1–3 weeks depending on archive depth and site complexity.
SCREENING QUESTIONS
1. Have you worked with the Wayback Machine CDX API or similar historical archive tools before? Briefly describe a past project.
2. Based on general experience, what's your rough sense of how deep Wayback Machine coverage typically goes for small-to-mid-size e-commerce sites?
3. What's your approach when a site has poor/no archive coverage — do you have fallback methods?
4. Please quote a fixed price and rough timeline for this scope.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in