Senior Python Engineer — production review + live smoke test of an LLM data-collection pipeline
Budget: $2000.0
FIXED /
⭐ 5.00 (27)
United States
api-integration, selenium-webdriver, selenium, python, data-mining, data-scraping, etl-pipelines
Preferred qualifications
- Experience: Expert
Summary
I run a research data-collection pipeline (open-source, Python) that discovers, scrapes, and uses LLMs (OpenAI Batch API) to extract structured records from government websites at scale. The code is mature — ~12k LOC, ~540 tests, 88% coverage, CI, packaged — but it has never had a live production run. Before I commit to one large, expensive, irreversible sweep, I want an experienced engineer to independently review it, run a small live end-to-end test, and harden the specific things that would waste money or corrupt output at full scale.
This is a review-and-harden engagement, not a greenfield build. I have a written scope with ranked, discrete deliverables.
What you'll do (ranked)
Run a small live end-to-end test (~100 institutions) against live search + OpenAI Batch APIs using throwaway, budget-capped keys I provide — and fix every real-world failure it surfaces.
Add a cost circuit-breaker so a run aborts if projected/actual spend crosses a ceiling.
Close a short list of known correctness gaps at the LLM-output boundaries (foreign-key integrity into the final dataset, date validation, schema-rule enforcement).
Validate the concurrency path under real parallelism (or tell me why it should stay single-threaded for the sweep).
Deliver a prioritized production-hardening findings memo (provider abstraction, observability for a multi-day run, error-path tests).
Must-have skills
Expert Python (typing, packaging, pytest); comfortable in a well-tested existing codebase.
Hands-on OpenAI Batch API experience (chunking, polling, cost, cached tokens).
Web scraping at scale (requests, Playwright/headless, robots.txt, rate limiting).
Pydantic / schema validation.
Track record hardening pipelines for cost-safe, resumable, idempotent long runs.
Nice to have
Research-data / reproducibility background.
Experience running LLM extraction against messy multilingual web content.
How I'll choose
Please answer in your proposal (see screening questions). I'll shortlist, share the public repo link + full written scope, and award a paid first milestone (a review memo) before the larger hardening milestones.
Screening questions (answer briefly)
Describe a data pipeline you made cost-safe or resumable for a long/expensive run. What specifically did you add?
You have an LLM emitting a structured record that includes an ID meant to key back to a master table. How do you guarantee the emitted ID is trustworthy before it lands in the final dataset?
How do you design a hard cost ceiling for an OpenAI Batch job that may run for days across many chunks?
How would you run a small, cheap, live end-to-end test to de-risk a large run without spending much?
Important
All research/methodology decisions stay with me; you flag, you don't change them.
You will not run the full sweep or receive production credentials or the full dataset — only a small sample + capped test keys.
Ouvrir sur Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Connexion