Web Scraping, OCR & Data Pipeline Developer
Budget: $5.0
FIXED /
⭐ 4.99 (715)
United States
data-scraping, python, etl, selenium, data-extraction, data-mining, postgresql
Preferred qualifications
- Experience: Intermediate
We’re looking for an experienced developer to build a pipeline for processing authorized resume/CV documents and permitted public data.
Workflow:
Scraping → PDF Processing → OCR/Parsing → Data Extraction → Cleaning → Deduplication → PostgreSQL
Requirements:
- Python & Web Scraping
- OCR/PDF Document Parsing
- Resume/CV Data Extraction
- PostgreSQL & ETL Pipelines
- Selenium, Playwright, or Scrapy
- Tesseract, AWS Textract, or Google Document AI is a plus
All data sources and documents must be authorized for collection and processing. No bypassing CAPTCHAs, authentication, paywalls, or access controls.
Please share relevant examples of your previous scraping, OCR, document-processing, or data-pipeline work.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in