← Обяви

Python scraper for TrustedHousesitters house sit listings (structured export)

Бюджет: $10.0 FIXED / ⭐ 5.00 (3) United States

api-integration, data-extraction, crawlers, data-mining, data-scraping, python

Предпочитана квалификация

  • Опит: Средно ниво
Project overview I need a reliable scraper for TrustedHousesitters.com that collects publicly available house sit listings into clean, structured files (CSV and JSON). Goal is market/research-style data: locations, dates, pets, amenities, and listing metadata — not private contact info. Data to collect (per listing) Where available on public listing pages / search results: Listing URL / listing ID Title / short description Location (city, region/state, country) Start date / end date (or date range) Duration (if shown) Pets (types, counts, notes) Home type / property details Amenities / responsibilities Photos (URLs only) Host/sitter-facing public profile fields if visible without login Posted/updated date (if available) Any public tags (e.g. family home, remote, long-term) Do not collect: emails, phone numbers, full street addresses, or anything behind a login/paywall that requires bypassing authentication. Scope Scrape search/browse results for house sits (all locations or filterable by country/region — we’ll confirm targets at kickoff). Open each listing and extract the fields above. Export to CSV + JSON with a clear schema. Handle pagination, rate limiting, retries, and basic anti-bot resilience without credential stuffing, CAPTCHA farms, or ToS-violating account abuse. Deliver a short README: how to run, dependencies, config, and field definitions. Optional (quote separately): scheduled re-runs / “new listing” monitoring. Tech preferences Python preferred (requests/httpx, Playwright/Selenium if needed, BeautifulSoup/lxml/parsel) Clean, documented code I can re-run myself Configurable delays / concurrency Idempotent scraping (dedupe by listing ID/URL) Logging for failed pages Deliverables Working scraper script(s) Sample output (CSV + JSON) from a real run Field schema / data dictionary README with setup + run instructions Brief notes on limitations (login walls, blocked regions, missing fields) Requirements for proposals Please include: Relevant scraping experience (similar travel/marketplace sites) Your approach (API inspection vs HTML, tooling) Timeline and fixed price for: MVP: one country/region + full field extract Full: broader geographic coverage Confirmation you will only scrape public data and respect reasonable rate limits Example of a past scraping deliverable (anonymized OK) Notes Site structure and access can change; please budget time for maintenance if we do ongoing runs. Prefer quality and clean schema over maximum volume on day one.
Отвори в Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Вход