← Обяви

Build a Repeatable Parcel and GIS Ingestion Pipeline (Python, PostGIS, ArcGIS REST)

Бюджет: $22000.0 FIXED / ⭐ 0.00 (0) United States

python, postgresql, postgis, etl

Предпочитана квалификация

  • Локация: United States
  • Опит: Експерт
Summary We are a US data company that works with public land and property records. We are hiring one senior engineer to build our ingestion pipeline from the ground up: first a working ingest of one full US state's parcel data (Washington State), then a repeatable process that refreshes on a schedule, and finally a documented playbook that lets our team add new states, counties, and cities without you. This is a fixed price, milestone based project with a short paid test up front. Estimated duration is 4 to 6 weeks at full time or near full time effort. There is a strong chance of ongoing work for the right person, because we add jurisdictions continuously after this project. Project specifics (target design, data sources, repository) are shared with shortlisted candidates after a signed NDA. What you will build The foundation. A Python codebase with Docker based local development (PostgreSQL + PostGIS, S3 compatible object storage), a catalog of sources and datasets, CI, and a raw file archive with checksums and acquisition metadata. The first ingest, one full state. Connectors for two delivery protocols: bulk file download from a state GIS program and an ArcGIS REST Feature Service from a county. Parse to Parquet and GeoParquet, run data quality checks (file integrity, schema changes, CRS and geometry validity, record counts by jurisdiction), and load a defined set of parcel attributes into PostGIS with full traceability from each loaded record back to its source file. Attributes include parcel boundary, parcel identifier, situs address, acreage, legal description, land use code, assessed value and tax record by year, and building and improvement records where the source provides them. Owner name and mailing address fields are routed by your code into a separate restricted schema; you will develop and test against a redacted copy of the source data, and we run the production load ourselves. The repeatable refresh. Scheduled runs that pull each new release, detect what changed, hold failed or anomalous runs for review instead of touching live tables, and produce a change report per run with freshness metadata per source. Internal parcel identifiers that stay stable when a county changes its parcel numbering or splits and merges parcels. The playbook for adding jurisdictions. A connector template, a fixture based conformance test suite, a written onboarding checklist for a new state, county, or city, and a demonstration: you onboard one additional county using only the playbook, and it passes conformance. What is out of scope Housing and development data of any kind (permits, plats, entitlements, new construction, builder or listing data). Recorded documents (deeds, mortgages, liens). Zoning, environmental, and utility layers (separate projects). Any paid or licensed data source. Any public API, product integration, AI features, or Esri hosted services. What we provide A written target design and acceptance criteria for every milestone (under NDA), a GitHub repository, hosting and object storage, and a single decision maker who responds within one business day. Requirements 5+ years building production data pipelines in Python. Real PostGIS experience: geometry types, CRS transformations, spatial indexes, validity repair. Hands on with GDAL/OGR, GeoPandas or Fiona, and shapefile, File Geodatabase, and GeoPackage formats. Experience extracting from ArcGIS REST Feature Services (paging with maxRecordCount, returnIdsOnly, resume on failure). Comfortable with object storage (S3 compatible), Docker, GitHub Actions, and writing tests. Fluent written English and at least four hours of daily overlap with US Mountain Time. You write clear documentation. The last milestone is judged on whether someone else can follow it. Nice to have: Parquet/GeoParquet and DuckDB, Prefect or Dagster, experience with county assessor or CAMA data. How hiring works Apply with answers to the four screening questions below. Generic or templated proposals are declined. Shortlisted candidates sign an NDA and receive a paid test (about 6 hours, $350 fixed) using a public county endpoint. The best test result gets the contract: five milestones, each with written acceptance criteria and its own payment. All work is work for hire and the IP belongs to us. Two-factor authentication on GitHub is required. We are open to freelancers worldwide. We are unable to work with contractors located in China, Russia, Iran, North Korea, Cuba, or Venezuela. Screening questions Describe a pipeline you built that re-ran safely against changing upstream data. How did you detect change and what happened when a source broke? A county ArcGIS Feature Service has 180,000 parcels, a maxRecordCount of 1,000, and features are being edited while you crawl. How do you get a consistent snapshot? A county replaces its assessment system and every parcel gets a new parcel number. How would you keep our internal parcel IDs stable? Link to a repository or write up of geospatial pipeline work you are allowed to share.
Отвори в Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Вход