Playwright + Firebase Developer for Production Data Ingestion System
Бюджет: -
HOURLY / PART_TIME
⭐ 4.99 (25)
United States
typescript, google-firestore, node.js, firebase, javascript, browser-automation, data-extraction
Preferred qualifications
- Experience: Intermediate
We are looking for an experienced independent developer to build a production web-scraping and data-ingestion system that monitors approximately 30 public websites.
This is not a one-time scraping project.
The system will repeatedly monitor a known set of sources, remember previously observed state, detect meaningful changes, and generate structured proposals for human review before approved data reaches the production application.
A detailed technical architecture has already been designed. We are looking for someone who can implement it pragmatically, challenge unnecessary complexity, and favor simple, maintainable solutions over over-engineering.
WHAT YOU WILL OWN
+ Playwright scraper framework
+ Site-specific scraper adapters
+ TypeScript / Node.js ingestion logic
+ Listing-level and detail-page extraction
+ Data normalization
+ Source-state storage
+ Delta detection for new, changed, missing, and unchanged records
+ Source-health safeguards to prevent bad, partial, blocked, or incomplete scrapes from creating false changes
+ Partition-aware handling for cases where one section, city, page, or navigation branch fails while the rest of a source succeeds
+ Entity matching and deduplication across sources
+ Multi-source provenance
+ Structured proposal generation
+ Idempotency and duplicate protection
+ Rejection / reconciliation handling so the same unchanged observation does not repeatedly create proposals
+ Synthetic fixtures and regression tests
+ Scraper diagnostics, failure artifacts, and maintenance
+ Rapid site-specific adaptation once the real target sites begin launching
EXISTING APPLICATION DEVELOPER
The existing application developer owns:
+ Production Firebase / Firestore data model
+ Admin and human-review workflow
+ Production publishing pipeline
+ Client synchronization
+ Mobile and web application behavior
+ Canonical production entity schemas and matching rules
The ingestion developer will work against a clearly defined integration contract and apply the agreed matching rules against a stable interface to the production data.
The ingestion developer is not expected to redesign the production application or independently build the admin/publishing system.
REQUIRED EXPERIENCE
You should have strong experience with:
+ Playwright
+ Node.js / TypeScript
+ Firebase / Firestore
+ Recurring, stateful production scraping systems
+ Dynamic JavaScript websites
+ Data normalization and entity matching
+ Idempotent background jobs
+ Change / delta detection
+ Failure handling and retry logic
+ Testing and debugging scrapers as websites change
+ Observability and diagnostics for production jobs
Experience building a one-time scraper is not enough for this project.
We are especially interested in candidates who have built systems that run repeatedly over time and must safely distinguish between:
+ New records
+ Changed records
+ Missing records
+ Unchanged records
+ Partial or unhealthy scrape results
AI-ASSISTED DEVELOPMENT
We strongly value developers who use modern AI coding tools effectively, such as:
+ Claude Code
+ Codex
+ Cursor
+ Similar coding agents or AI-assisted development tools
We are not looking for someone who simply delegates the entire project to AI.
We are looking for a developer who can use these tools to accelerate repetitive implementation work while still exercising strong engineering judgment, reviewing generated code, writing tests, and validating production behavior.
PROJECT PHILOSOPHY
This should be a lean, maintainable v1.
We do not want enterprise infrastructure for its own sake.
The architecture describes required behavior and safeguards, but the developer is encouraged to simplify implementation when the same reliability and data integrity can be achieved with a simpler approach.
The core workflow is:
Source Websites
→ Extraction
→ State Comparison
→ Structured Proposal
→ Human Review
→ Production Publish
Human review is intentional.
The goal is to automate repetitive discovery and data entry, not eliminate human oversight.
IMPORTANT TIMING CONTEXT
Most target websites are seasonal and are not expected to expose their current production data for approximately 1–2 months.
We have screenshots and reference material from last year’s versions of the sites, which provide useful examples of likely layouts, navigation patterns, and data presentation.
Initial development should therefore focus on:
+ Reusable ingestion framework
+ Standard scraper-adapter contract
+ Synthetic fixtures
+ Source-state logic
+ Delta detection
+ Source-health and partition-health handling
+ Failure handling
+ Idempotency
+ Proposal generation
+ End-to-end test harness
Once the live target sites begin appearing, the developer will adapt the site-specific Playwright adapters to actual production behavior and convert important real-world cases into regression fixtures.
The project therefore has two major stages:
+ Build and validate the core system now using fixtures and reference material
+ Perform rapid live-source adaptation and hardening as the real sites launch
IMPLEMENTATION PHASES
1. Shared integration foundation
+ Define ingestion and proposal contracts
+ Establish source-state storage
+ Establish dependency and matching interfaces
+ Confirm Firebase integration boundaries
2. Playwright framework and test harness
+ Build reusable execution framework
+ Define site-adapter contract
+ Build synthetic fixtures
+ Implement health validation, retries, idempotency, and delta detection
3. Synthetic end-to-end vertical slice
+ Prove one complete workflow against deterministic fixture data
+ Validate extraction, state comparison, proposal generation, review, and non-production publication behavior
4. First live production source
+ Adapt the framework to the first representative live source
+ Validate actual navigation, DOM behavior, dynamic loading, access behavior, and real-world extraction
5. Representative source expansion
+ Expand to several structurally different sources
+ Use real failures and layouts to improve regression coverage
6. Remaining source rollout
+ Add adapters for the remaining target sources
+ Tune source-specific matching, health rules, cadence, and maintenance behavior
7. Stabilization and handoff
+ Final testing
+ Documentation
+ Regression coverage
+ Operational handoff
BUDGET
Fixed-price budget: $4,000–$6,000, divided into milestones.
PLEASE SEE SCREENING QUESTIONS
Please keep your proposal concise and specific to this project.
In your proposal, include your fixed-price estimate within the stated $4,000–$6,000 budget and an approximate timeline broken into:
+ Core framework and test harness
+ Synthetic end-to-end vertical slice
+ First complete live source
+ Expansion to approximately 5 representative sources
+ Remaining adapters to approximately 30 total sources
+ Final stabilization, documentation, and handoff
We are looking for an individual freelancer, not an agency. The detailed technical architecture brief will be provided to shortlisted applicants.
Generic scraping proposals, proposals that do not answer the screening questions, and applicants whose expected budget is substantially above the stated range will not be considered.
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти