Adapt Existing Python Script + Build K-12 School Directory Dataset (Code Provided).
Budget: $400.0
FIXED /
⭐ 4.74 (122)
United States
data-extraction, python
Qualifiche preferite
- Esperienza: Intermedio
What this is:
I have working, tested Python code that extracts staff contact information from public school district staff directory pages. I do not need it built. I need someone to adapt it for district-level targeting, build the input dataset, run it, and hand me clean output.
This is a data engineering and data wrangling job, not a from-scratch development project. Please read the attached README before proposing — it documents the pipeline, the known limitations, and exactly what is already done.
Background:
I run a new national professional association for school counselors and college admissions professionals. I need to build a contact database of high school counselors, sourced from publicly published school and district staff directories.
What already exists (see attached):
extractor.py- Parses staff directory pages, identifies counselor-type roles by title, rejects non-target roles. Tested against table, card, and prose layouts.
fetcher.py: Polite fetching with robots.txt compliance, per-domain rate limiting, disk caching, resumable.
run_scrape.py: Pipeline driver, CSV output, progress tracking.
verify_emails.py: Syntax, duplicate, role-address and dead-domain filtering.
README.md: Full pipeline documentation.
Standard library only, no dependencies. Python 3.8+.
What I need you to do:
1. Build the target dataset. Using NCES Common Core of Data (ELSI table generator or the Public School Universe file), produce a list of US public school districts, ranked by Black student enrollment in grades 9–12. Include district name, state, NCES district ID, enrollment figures, and the district website URL. Sourcing the website URLs reliably is the hardest part of this job — please tell me in your proposal how you would approach it.
2. Adapt the script for district-level targeting. The current version targets individual schools. Many districts publish one central staff directory covering all schools, which is more efficient. Adjust the pipeline accordingly.
3. Handle JavaScript-rendered directories. The README documents this limitation — the current tool reads raw HTML only. Add a headless browser fallback (Playwright or Selenium) for pages that return no contacts, or propose a different approach.
4. Tune and test. Run against 50 districts. Share the output with me for review. Adjust title-matching patterns based on what you find before proceeding.
5. Full run and delivery. Run the top 500 districts. Deliver a verified CSV plus a short written handoff explaining how to rerun it.
Deliverables:
● districts.csv — ranked target list with working URLs
● Updated scripts with your changes documented in comments
● contacts_verified.csv — name, title, email, school, district, state, source URL
● A one-page written handoff covering how to rerun, and what to change to extend the list
● Brief notes on which districts failed and why
Scope boundaries — (please read):
This project collects data only from publicly published staff directory pages and public government datasets. It does not touch:
● LinkedIn or any social platform
● Purchased or third-party contact lists
The existing code respects robots.txt and rate-limits requests. Please preserve both. If you believe a target should be avoided, tell me — I would rather skip a source than create a problem.
Timeline:
Two weeks from award. Milestone 1 within the first four days.
Budget:
$400 fixed price, paid across three milestones. If you believe the scope warrants more, say so in your proposal with your reasoning — I would rather hear that upfront than mid-project.
To apply:
Answer the three screening questions. Proposals that don't answer them, or that propose rebuilding from scratch, will not be considered. Please don't send a generic template — I read every proposal and I can tell.
Apri su Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Accedi