Senior Data Engineer for Web Scraping
Бюджет: -
HOURLY / FULL_TIME
⭐ 0.00 (0)
United States
python, machine-learning, selenium, scrapy-framework
Preferred qualifications
- Experience: Expert
Title: Senior Data Engineer – Distributed Data Pipeline & Spatial Database Architecture
Description:
We are building an in-house real estate data engine to aggregate, clean, and enrich property records at scale across the US. We need an expert Data Engineer to architect a high-throughput data ingestion pipeline and spatial database that processes public records and integrates enriched data into our CRM APIs.
Responsibilities:
Design and deploy scalable data ingestion workflows for public county assessor portals, open datasets, and GIS APIs.
Implement robust network request management, rate limiting, and session handhandling to ensure high uptime and pipeline resilience.
Architect a PostgreSQL/PostGIS and Elasticsearch infrastructure optimized for complex spatial queries, property searches, land-use filtering, and Census data integration.
Normalize and clean large-scale unstructured/semi-structured real estate datasets.
Requirements:
Proven experience building high-throughput, distributed data pipelines and web scrapers using Python (Scrapy, Playwright, or similar).
Expertise with spatial databases (PostgreSQL/PostGIS) and Elasticsearch.
Deep understanding of HTTP request handling, IP rotation strategies, user-agent management, and distributed crawler architecture.
Prior experience working with US real estate, GIS, or public record datasets.
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти