Python Engineer to Extend a Document Processing Pipeline
Budżet: $2000.0
FIXED /
⭐ 5.00 (1)
United Kingdom
api-integration, data-extraction, data-source-integration, etl, python, etl-pipelines, data-science, python-script, automation
Preferowane kwalifikacje
- Typ talentu: Niezależny
- Lokalizacja: United Kingdom
- Doświadczenie: Ekspert
- Angielski: Biegły
- Job Success: 90%+
- Preferowany Rising Talent
- Min. zarobki: $1,000+
We are looking for an experienced Python engineer to extend an existing document processing application so that it can accept Word, HTML, Excel, and PowerPoint files in addition to PDFs.
We have an operational application that processes sets of publications relating to research studies. Our application currently:
Reads grouped PDF files from a SharePoint folder or local folder
Retrieves version-controlled extraction prompts from Supabase
Submits each study’s publication set to the Claude API
Stores raw LLM outputs for audit and traceability
Parses LLM outputs into structured fields
Populates an Excel-based data extraction table
This project is focused on extending its input capabilities while preserving the current downstream workflow, audit trail, and output structure.
The selected engineer will:
Review the existing Python codebase and workflow
Recommend an appropriate architecture for supporting multiple document formats
Implement reliable ingestion and parsing for Word, HTML, Excel and Powerpoint files
Integrate the new input formats into the existing study-level processing workflow
Ensure files from SharePoint and local folders are handled consistently
Preserve source metadata and raw Claude API responses for audit purposes
Plan for handling of unsupported, corrupted, encrypted, or partially readable files
Confirm that existing PDF processing continues to work without regression
Document the implementation, dependencies, limitations, and configuration requirements
Provide a clear handover of the completed work
We are open to the engineer recommending whether documents should be converted into a common intermediate representation or processed through format-specific adapters. The proposed approach must be maintainable and must preserve the content required for accurate extraction.
Deliverables
A short technical assessment of the existing architecture
An agreed implementation plan covering supported formats and parsing approach
Production-ready Word, HTML, Excel and PowerPoint support
Integration with the existing Claude API and Supabase workflow
Test results for a representative set of documents
Error handling and logging for failed or unsupported inputs
Updated technical documentation
Handover of the completed and tested code
The work will be considered complete when:
Supported files can enter the same study-level workflow currently used for PDFs
Relevant textual and tabular content is passed through the pipeline without material loss
The application retains sufficient source metadata for audit and troubleshooting
Raw Claude API outputs continue to be stored
Structured outputs continue to populate the existing Excel extraction table
Errors are logged clearly without causing unrelated studies to fail
Required experience
Strong Python development experience
Experience maintaining and extending an existing codebase
Practical experience with the Anthropic Claude API
Experience with Supabase, PostgreSQL, or a comparable backend platform
Experience parsing and normalising documents in Python
Experience working with structured and semi-structured content
Strong automated testing and debugging practices
Familiarity with Git-based development workflows
Ability to explain technical decisions and trade-offs clearly
Please include in your proposal:
A brief description of the most relevant pipeline you have built or maintained
Your experience with the Claude API or other LLM APIs
Your initial view on whether to use a common intermediate representation or format-specific adapters
Your availability and an indicative delivery timeline
Any access or technical information you would need before providing a detailed estimate
Please focus your proposal on directly relevant experience rather than sending a general development profile.
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj