LLM Document Data Extraction with Schema Validation & OCR
Presupuesto: -
HOURLY / PART_TIME
⭐ 0.00 (0)
Albania
data-extraction, computer-vision, ocr-algorithms, ocr-tesseract, opencv, artificial-intelligence, machine-learning, python, pdf-conversion, pdf, natural-language-processing, image-processing, data-science
Cualificaciones preferidas
- Experiencia: Intermedio
I need an engineer to build/harden a document data-extraction pipeline — not a chatbot, not a Q&A assistant. The system reads long, often scanned documents (PDFs with tables, multi-column layouts, numbered clauses) and must output a verified, structured checklist/dataset (JSON) that other systems depend on.
Must have:
Structured/schema-enforced LLM output (Pydantic, Instructor, or equivalent function-calling approach) — no free-text parsing
Real experience with messy PDFs: scanned pages, tables, multi-column layouts, OCR when needed
A concrete method to verify extracted fields against the source text (verbatim/fuzzy matching, confidence scoring) and flag/escalate low-confidence extractions to a human reviewer — zero silent misses is the requirement
Experience measuring extraction quality (test/golden set, precision-recall, or similar) — not just "it works on my examples"
Cost-aware production setup (batching, caching) for processing documents at volume
Abrir en Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Entrar