Extract & Classify PDF Product List Into Excel (Python + PDF Parsing)
Budżet: -
HOURLY / PART_TIME
⭐ 4.99 (94)
Ireland
data-scraping, data-extraction, data-mining, python, microsoft-excel, pdf-conversion
Title: Turn Large PDF Directory of Fertiliser Products into Clean Excel File
I have a large PDF that lists many companies and their agricultural products (fertilisers, soil amendments, bio‑inputs, etc.). I need this turned into a clean Excel/CSV file with some simple classification rules applied.
### What I need done
1. **Parse the PDF into structured data**
For each product entry, extract:
- Company name
- Country
- Website URL (if shown)
- Email address(es) (if shown)
- Product name
- Any category/description text that helps identify what kind of product it is
Where the PDF doesn’t clearly show website or email, I need you to:
- Look up the company on the relevant public directory/website I’ll point you to
- Capture main website URL and a sensible contact email (regulatory, sales, or general info)
2. **Classify each product**
For each product, create two extra fields:
- **Product type** – a short 1–5 word label, e.g.:
- “chicken manure pellets”
- “compost fertilizer”
- “biochar soil amendment”
- “fish fertilizer”
- “coconut coir substrate”
- “biopesticide – neem”
- “sanitizer / disinfectant”
- **Classification** – one of: **Yes / Maybe / No**, based on simple rules:
- **Yes** (directly relevant):
- Fertilisers and soil amendments, especially high organic‑matter products:
- Manure‑based products (e.g. chicken manure pellets)
- Compost / compost blends
- Biochar and biochar blends
- Humic/fulvic products
- Biofertilizers / microbial soil inoculants
- Growing media/substrates (coconut coir, peat/coir blends, potting soils, worm castings, etc.)
- **Maybe** (indirectly relevant / line‑extension potential):
- Herbicides and similar products with a soil or soil‑surface connection
- Biopesticides / botanical pesticides / mycoinsecticides (e.g. neem, spinosad, bacillus, viral biocontrols)
- Adjuvants, wetting agents, surfactants, spreaders used with fertilisers and soil amendments
- **No** (not relevant):
- Processing and sanitation products (disinfectants, detergents, cleaners, chlorine‑based systems, etc.)
- Food processing inputs only
- Livestock‑only products (feed additives, health products, bedding) where the company has no compost/manure/soil amendment products
3. **Add Region from Country**
- Parse **Country** from the address.
- Create a **Region** field with one of four values:
- **Americas** – North, Central, and South America
- **Europe** – all European countries, including UK & Ireland
- **Africa** – all African countries
- **Asia** – all Asian countries (including Middle East; we can agree edge cases)
4. **Final output format**
Deliver a single CSV/Excel file with **one row per product (and email, if distinct)**, with at least:
- Company
- Country
- Region (Americas / Europe / Africa / Asia)
- URL
- Email
- Product
- Product type (1–5 words)
- Classification (Yes / Maybe / No)
If a company has multiple products, repeat the company info on multiple rows. If a company has multiple clear contact emails, multiple rows are fine.
I would also like (if possible):
- The script or code you use, so this can be rerun on updated PDFs in future.
***
### Skills required
- Strong experience with:
- Python or similar scripting language for text/PDF parsing
- PDF text extraction (e.g. PyMuPDF, pdfplumber, or equivalent)
- String parsing and regular expressions
- CSV/Excel data handling
- Comfortable:
- Working with large, semi‑structured PDFs
- Applying and documenting simple business rules
- Doing light lookup on public web pages to fill missing details
***
### Deliverables
1. Clean CSV/Excel file with the columns listed above.
***
### Timeline and proposal
- ASAP
Experts only please and fluent English is essential. Project will be awarded immediately after a quick zoom call. Thanks for reading.
Otwórz na Upwork