PDF Parsing & Document AI Expert Needed to Improve an Existing System
Rozpočet: $30.0 - $60.0
HOURLY / PART_TIME
⭐ 5.00 (1)
Germany
php, javascript, html5, java
Preferované kvalifikace
- Zkušenost: Expert
The system should accurately understand complex page structures, including multi-column layouts, mixed content, tables, captions, images, headers, footnotes, and other document regions.
Table and glyph extraction
Tables must be extracted while preserving:
1. Merged cells
2. Header relationships
3. Reading order
4. Special glyphs and symbols such as e.g ✓, ⊗, ●, and ○
4. Technical image and diagram extraction
The system must correctly detect and crop complete:
1. Technical drawings
2. Process-flow diagrams
3. Architecture diagrams
4. Schematics and other embedded figures
Images and diagrams must not be cropped in the middle or divided incorrectly across multiple regions.
Responsibilities
- Review our existing codebase and parsing pipeline
- Evaluate current extraction quality and identify failure patterns
- Improve the existing architecture build around and model-routing logic
- Determine which parser or extraction method should be used for each content type
- Implement validation, confidence scoring, and fallback mechanisms
- Improve consistency across different PDF layouts and versions
- Test improvements against a representative set of documents
- Define measurable accuracy and quality metrics
- Document the changes and technical decisions
We are open/looking for a hybrid or agentic approach that combines native PDF extraction, OCR, layout-detection models, multimodal models, and specialised parsing tools.
Please apply only if you have previously built or improved a similar production-level PDF parsing or Document AI system.
Otevřít na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Přihlásit