← Missions

Python/Document Intelligence AI Engineer (OCR, LLM Extraction, RAG with Citations)

Budget: $10.0 - $10.0 HOURLY / PART_TIME ⭐ 4.98 (135) United States

artificial-intelligence, data-extraction, postgresql, react-js, amazon-web-services, api-integration, next.js, docker, document-analysis

Qualifications préférées

  • Expérience : Intermédiaire
Looking for an experienced AI engineer to design and build a document intelligence pipeline for legal and business documents. The system takes scanned and digital PDFs, extracts structured data, validates it, routes uncertain cases to a human reviewer, and delivers approved data through an API. It also provides RAG based question answering across documents with page level citations. This is a production build, not a prototype or a chatbot wrapper. You should have shipped a document extraction system before and understand that OCR and layout handling come before the LLM, that extraction needs a schema, and that wrong data must never pass through silently. Pipeline Upload (PDF, scan, image, DOCX) → document classification → OCR and layout extraction → LLM structured extraction (JSON, Pydantic schema) → validation rules → confidence gating → human review queue → approval snapshot → export via REST API and CSV Phase 1 scope (MVP) • Ingestion service: accepts PDF, DOCX, JPG, PNG; detects scanned vs digital PDFs; splits multi document files; deduplicates; stores originals in S3 with hashing • OCR and layout layer: Azure Document Intelligence or AWS Textract (recommend and justify), with a fallback path (Tesseract or PaddleOCR) for offline or low volume use; preserve page and bounding box provenance for every extracted value • Classification: document type detection using an LLM or a lightweight classifier, with versioned schemas per type • Structured extraction: OpenAI or Anthropic API with JSON mode / tool use, Pydantic validation, confidence score per field, retry on validation failure, prompt caching to control cost • Validation rules: required fields present, date ordering, party name consistency across a document set, totals reconcile on invoices, unknown layouts fail safely instead of producing partial output • Review interface (React or Next.js): extracted value shown next to the source page image, field level correction, append only correction history, approval creates a locked final version • RAG question answering: chunking with page references, pgvector hybrid search (vector + BM25), reranking, streamed answers with document name and page citation, explicit "not found" when documents do not contain the answer • Async processing: Celery/Redis or ARQ, retries, dead letter handling, status tracking per document • Data layer: PostgreSQL + pgvector, tenant and matter level isolation with row level security, audit log of every extraction, correction, approval and export • Deployment: Docker Compose, AWS (EC2 or ECS, S3, RDS), GitHub Actions CI, environment based config, basic monitoring and error alerting • Evaluation: hand labelled golden set on day one, field level precision/recall report, regression check on every prompt or schema change Phase 2 (after MVP acceptance) Practice management integration (Clio or Filevine), deadline and obligation extraction to calendar, additional document types, Stripe billing, admin dashboard, cost per document optimisation, ongoing development. Long term engagement expected if Phase 1 goes well. Must have • At least one document extraction or IDP system shipped to production • Hands on OCR experience on real scanned documents: skew, noise, low resolution, tables, stamps, handwriting • Structured LLM output with schema validation and a clear plan for validation failures • RAG with source citations in production (pgvector, Pinecone or similar) • Full stack: Python backend plus a working React or Next.js review UI • Confidence gating, human in the loop review, audit logging • Multi tenant data isolation and secure handling of confidential documents • Cost awareness: able to quote and reduce cost per document Nice to have: legal or financial document domains, Clio/Filevine API, evaluation tooling (RAGAS, LangSmith) How to apply Start your proposal with DOC INTEL and answer: • One document extraction system you built for production: document types, OCR and LLM stack, how accuracy was measured • How you decide which fields go to human review and what the reviewer sees • How you prevent a wrong termination date from being exported silently • Approximate cost per document for a 20 page scanned contract with your stack, and how you would reduce it • Hourly rate and weekly availability
Ouvrir sur Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Connexion