← Oferty

Senior Data Acquisition & Document Intelligence Engineer

Budżet: - HOURLY / FULL_TIME ⭐ 0.00 (0) India

crawlers, database

Preferowane kwalifikacje

  • Doświadczenie: Średniozaawansowany
About the role We ingest very large volumes of publicly published documents and structured records from hundreds of external web sources every day. These sources are inconsistent, poorly standardised, frequently redesigned, and heavy on scanned PDFs and image-based documents. Your job is to make that mess arrive clean, complete and on time — every single day. This is not a one-off scraping project. You will own a production data acquisition system end to end: crawling, document extraction, OCR, parsing into structured schemas, and the ongoing quality assurance that keeps the dataset trustworthy as sources change underneath you. We are looking for someone who has already done this at scale and knows exactly where these pipelines break. What you will do 1. Crawling and acquisition • Design, build and operate large-scale crawlers and scrapers across hundreds of heterogeneous sources with differing structures, login flows and pagination patterns. • Handle sessions, cookies, tokens, dynamic JavaScript rendering, multi-step form navigation and file-download workflows. • Implement polite crawling: rate limiting, backoff, retry strategies, request budgeting and respect for source-side load. • Build incremental and change-detection crawls so we re-fetch only what actually changed, instead of re-scraping everything. • Monitor source drift — detect when a site’s structure, endpoint or format changes and fix the extractor before the data gap reaches downstream users. 2. Document extraction and OCR • Process high volumes of PDFs (native and scanned), images and office formats through a reliable extraction pipeline. • Build and tune OCR workflows for poor-quality inputs: low-DPI scans, skewed and rotated pages, stamps and seals, watermarks, handwritten annotations, faded prints and multi-column layouts. • Apply image pre-processing (deskew, denoise, binarisation, cropping, resolution upscaling) to lift OCR accuracy on difficult documents. • Extract tabular data reliably — including merged cells, multi-page tables, borderless tables and nested headers. • Work with multilingual documents, including Indian regional scripts alongside English. • Benchmark OCR engines against real samples and choose per document type rather than defaulting to one engine for everything. 3. Structuring and normalisation • Design the target schemas and translate raw, unstructured document text into consistent structured records. • Build extraction logic combining rules, patterns, layout awareness, NER and LLM-assisted parsing — and know when each approach is the right tool. • Normalise messy real-world values: dates in a dozen formats, amounts with mixed separators and units, inconsistent entity names, addresses, identifiers and codes. • Handle deduplication and entity resolution across sources where the same record appears in multiple forms. • Maintain and evolve taxonomies, classification logic and reference / master data. 4. Data quality — the part that never ends • Define and enforce quality dimensions: completeness, accuracy, freshness, consistency, uniqueness and validity. • Build automated validation, assertion and anomaly-detection layers that fail loudly before bad data reaches consumers. • Maintain a gold-standard sample set and run regular accuracy benchmarking against it, with published accuracy rates per source and per document type. • Set up field-level and record-level QA workflows, including a human-in-the-loop review path for low-confidence extractions. • Own reconciliation: prove that what we hold matches what the source actually published. • Maintain lineage and audit trails so any field in the final dataset can be traced back to the exact source document, page and extraction run. 5. Production operations • Orchestrate and schedule everything as monitored, alerting, self-healing production pipelines — not scripts run by hand. • Own uptime, throughput, latency and cost of the ingestion system. • Write runbooks, document extractor logic, and make the system operable by someone other than you. • Report on coverage and quality metrics to stakeholders on a fixed cadence. Must-have qualifications • 5+ years of hands-on experience building and operating production data extraction / scraping / document-processing pipelines. Not five years adjacent to them. • Strong Python (or equivalent) with production-grade engineering practice — testing, version control, code review, packaging. • Scraping at scale — Scrapy, Playwright, Selenium, Puppeteer, requests / httpx — with proxy and session management, and handling anti-bot friction lawfully. • Deep OCR experience — Tesseract, PaddleOCR, EasyOCR, or cloud OCR / IDP services (Google Document AI, AWS Textract, Azure Document Intelligence) — including tuning for accuracy, not just calling an API. • PDF and document parsing — PyMuPDF, pdfplumber, pdfminer, Camelot, Tabula or comparable, including layout-aware extraction. • Strong SQL and solid relational data modelling; experience with PostgreSQL or equivalent. • Workflow orchestration — Airflow, Prefect, Dagster or similar. • Cloud storage and compute — GCP (GCS, BigQuery) and/or AWS (S3, Lambda, Batch) — and processing datasets too large to fit in memory. • Data quality frameworks and testing — Great Expectations, Soda, Pandera, dbt tests, or a rigorous in-house equivalent. • Docker, CI/CD, logging, monitoring and alerting in a production setting. • A track record of maintaining a dataset over time — improving accuracy quarter on quarter, not just shipping v1. • Clear understanding of the legal and ethical boundaries of automated data collection — terms of service, robots.txt, personal data handling and applicable data protection obligations. To apply Send your CV along with a short note covering: • The largest extraction pipeline you have owned — sources, daily volume, document types. • Your measured OCR / extraction accuracy on difficult inputs, and how you measured it. • One time your pipeline broke silently, how you found out, and what you changed so it could not happen again.
Otwórz na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Zaloguj