Web Crawler/Script Developer — Patent & Citation Data Extraction
Budget: $8.0 - $25.0
HOURLY / PART_TIME
⭐ 0.00 (0)
Singapore
data-scraping, crawlers, scrapy-framework, data-extraction, python, microsoft-excel, data-mining
Gewenste kwalificaties
- Ervaring: Gevorderd
I need a web crawler (script/tool) built to extract patent data for a list of companies pairs (acquirer, target, deal close date) from a patent database. This is for a doctoral dissertation dataset — accuracy and reproducibility matter more than speed.
What the crawler needs to do, per company/deal:
Given a company name and a deal close date, identify that company's patent assignee record(s) in the target database — including resolving parent/subsidiary relationships where the deal-list company name isn't the literal patent assignee name (this is the hard part — a naive exact-name-match crawler will miss most subsidiary-filed patents).
Count patents filed (not granted) in two windows relative to the close date:
Pre: 2 years before close → close date
Post: close date → 3 years after close
For each patent in those windows, pull forward citation count and generality index (or the underlying citation data needed to compute generality — a Herfindahl-style measure of citation spread across technology classes), so I can compute Σ(citations × generality) per window.
Output one row per deal with: patents pre-count, patents post-count, Σ(citations × generality) pre, Σ(citations × generality) post, total patents used as the quality-calc denominator (pre and post), the source database/URL, and the extraction date.
Data source:
Needs a database that exposes citation and generality/technology-classification data per patent — Derwent Innovation, PatentSight, Orbit Intelligence, Lens.org, or PatentsView/Google Patents BigQuery if the crawler can derive generality from CPC/IPC class spread itself. Tell me in your proposal which source you'd use and whether you already have access/licensing, since some of these are paid.
Deliverables:
The crawler script/tool (Python preferred), documented and re-runnable
Output as a CSV/Excel file matching my column structure (I'll provide the exact template)
A short note on match/coverage rate and any companies it couldn't resolve confidently
Required:
Experience building web scrapers/crawlers against patent or IP databases specifically (general scraping experience alone isn't enough here — patent assignee data is messy)
Ability to handle corporate name variants (legal suffixes, exchange tickers, name changes, subsidiaries)
Provided by me: the list of 181 deals (4 already completed as a worked example/quality benchmark), and the exact output column template.
Openen op Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Inloggen