← Live-Feed

LLM Document Data Extraction with Schema Validation & OCR

Budget: - HOURLY / PART_TIME ⭐ 0.00 (0) Albania

data-extraction, computer-vision, ocr-algorithms, ocr-tesseract, opencv, artificial-intelligence, machine-learning, python, pdf-conversion, pdf, natural-language-processing, image-processing, data-science

Bevorzugte Qualifikationen

  • Erfahrung: Fortgeschritten
I need an engineer to build/harden a document data-extraction pipeline — not a chatbot, not a Q&A assistant. The system reads long, often scanned documents (PDFs with tables, multi-column layouts, numbered clauses) and must output a verified, structured checklist/dataset (JSON) that other systems depend on. Must have: Structured/schema-enforced LLM output (Pydantic, Instructor, or equivalent function-calling approach) — no free-text parsing Real experience with messy PDFs: scanned pages, tables, multi-column layouts, OCR when needed A concrete method to verify extracted fields against the source text (verbatim/fuzzy matching, confidence scoring) and flag/escalate low-confidence extractions to a human reviewer — zero silent misses is the requirement Experience measuring extraction quality (test/golden set, precision-recall, or similar) — not just "it works on my examples" Cost-aware production setup (batching, caching) for processing documents at volume
Auf Upwork öffnen

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Anmelden