← Missions

Extract & Classify PDF Product List Into Excel (Python + PDF Parsing)

Budget: - HOURLY / PART_TIME ⭐ 4.99 (94) Ireland

data-scraping, data-extraction, data-mining, python, microsoft-excel, pdf-conversion

Title: Turn Large PDF Directory of Fertiliser Products into Clean Excel File I have a large PDF that lists many companies and their agricultural products (fertilisers, soil amendments, bio‑inputs, etc.). I need this turned into a clean Excel/CSV file with some simple classification rules applied. ### What I need done 1. **Parse the PDF into structured data** For each product entry, extract: - Company name - Country - Website URL (if shown) - Email address(es) (if shown) - Product name - Any category/description text that helps identify what kind of product it is Where the PDF doesn’t clearly show website or email, I need you to: - Look up the company on the relevant public directory/website I’ll point you to - Capture main website URL and a sensible contact email (regulatory, sales, or general info) 2. **Classify each product** For each product, create two extra fields: - **Product type** – a short 1–5 word label, e.g.: - “chicken manure pellets” - “compost fertilizer” - “biochar soil amendment” - “fish fertilizer” - “coconut coir substrate” - “biopesticide – neem” - “sanitizer / disinfectant” - **Classification** – one of: **Yes / Maybe / No**, based on simple rules: - **Yes** (directly relevant): - Fertilisers and soil amendments, especially high organic‑matter products: - Manure‑based products (e.g. chicken manure pellets) - Compost / compost blends - Biochar and biochar blends - Humic/fulvic products - Biofertilizers / microbial soil inoculants - Growing media/substrates (coconut coir, peat/coir blends, potting soils, worm castings, etc.) - **Maybe** (indirectly relevant / line‑extension potential): - Herbicides and similar products with a soil or soil‑surface connection - Biopesticides / botanical pesticides / mycoinsecticides (e.g. neem, spinosad, bacillus, viral biocontrols) - Adjuvants, wetting agents, surfactants, spreaders used with fertilisers and soil amendments - **No** (not relevant): - Processing and sanitation products (disinfectants, detergents, cleaners, chlorine‑based systems, etc.) - Food processing inputs only - Livestock‑only products (feed additives, health products, bedding) where the company has no compost/manure/soil amendment products 3. **Add Region from Country** - Parse **Country** from the address. - Create a **Region** field with one of four values: - **Americas** – North, Central, and South America - **Europe** – all European countries, including UK & Ireland - **Africa** – all African countries - **Asia** – all Asian countries (including Middle East; we can agree edge cases) 4. **Final output format** Deliver a single CSV/Excel file with **one row per product (and email, if distinct)**, with at least: - Company - Country - Region (Americas / Europe / Africa / Asia) - URL - Email - Product - Product type (1–5 words) - Classification (Yes / Maybe / No) If a company has multiple products, repeat the company info on multiple rows. If a company has multiple clear contact emails, multiple rows are fine. I would also like (if possible): - The script or code you use, so this can be rerun on updated PDFs in future. *** ### Skills required - Strong experience with: - Python or similar scripting language for text/PDF parsing - PDF text extraction (e.g. PyMuPDF, pdfplumber, or equivalent) - String parsing and regular expressions - CSV/Excel data handling - Comfortable: - Working with large, semi‑structured PDFs - Applying and documenting simple business rules - Doing light lookup on public web pages to fill missing details *** ### Deliverables 1. Clean CSV/Excel file with the columns listed above. *** ### Timeline and proposal - ASAP Experts only please and fluent English is essential. Project will be awarded immediately after a quick zoom call. Thanks for reading.
Ouvrir sur Upwork