Senior Data Engineer for Web Scraping
Budget: -
HOURLY / FULL_TIME
⭐ 0.00 (0)
United States
python, machine-learning, selenium, scrapy-framework
Qualifiche preferite
- Esperienza: Esperto
Title: Senior Data Engineer – Distributed Data Pipeline & Spatial Database Architecture
Description:
We are building an in-house real estate data engine to aggregate, clean, and enrich property records at scale across the US. We need an expert Data Engineer to architect a high-throughput data ingestion pipeline and spatial database that processes public records and integrates enriched data into our CRM APIs.
Responsibilities:
Design and deploy scalable data ingestion workflows for public county assessor portals, open datasets, and GIS APIs.
Implement robust network request management, rate limiting, and session handhandling to ensure high uptime and pipeline resilience.
Architect a PostgreSQL/PostGIS and Elasticsearch infrastructure optimized for complex spatial queries, property searches, land-use filtering, and Census data integration.
Normalize and clean large-scale unstructured/semi-structured real estate datasets.
Requirements:
Proven experience building high-throughput, distributed data pipelines and web scrapers using Python (Scrapy, Playwright, or similar).
Expertise with spatial databases (PostgreSQL/PostGIS) and Elasticsearch.
Deep understanding of HTTP request handling, IP rotation strategies, user-agent management, and distributed crawler architecture.
Prior experience working with US real estate, GIS, or public record datasets.
Apri su Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Accedi