AI Document Ingestion & Data Normalization System for Real Estate Platform
Buget: $6000.0
FIXED /
⭐ 5.00 (3)
Uruguay
python, react-js, artificial-intelligence, api-integration, data-extraction, ocr-tesseract, google-apis, amazon-web-services, machine-learning
Calificări preferate
- Experiență: Expert
PLANORBE is a real estate company with its own internal management platform (“Gestor”). We need a senior developer or small technical team to build a semi-automated document-ingestion and data-normalization system using modern AI tools.
IMPORTANT — PROVEN RELEVANT EXPERIENCE IS THE PRIMARY SELECTION CRITERION.
We want evidence that you or your team have actually built comparable systems in production. Proposals without concrete examples of relevant prior work will not be considered.
OBJECTIVE
Transform heterogeneous information received from real-estate developers into structured data compatible with our Gestor, with human review before publication.
TARGET DATA
Destination structure: Project → Typologies → Units.
Projects include general data, location, floors/units, completion, expenses, descriptions, financing, amenities, companies, renders, progress images, floorplans and downloads.
Typologies include bedrooms/bathrooms, attributes/equipment, surfaces, terraces, garage data and plans.
Units include typology, unit ID, floor, orientation/distribution, price, status and promotions.
Full schema and real project folders will be provided to shortlisted candidates.
SOURCE MATERIAL
Inputs may include PDF brochures/price lists, XLSX, Google Sheets/Docs, DOCX, JPG/PNG plans/renders, external Drive folders, partial updates and commercial information copied from email/WhatsApp. Formats vary by promoter and over time.
REAL CASES TO HANDLE
1. Conflicting sources, e.g. saved PDF vs live Google Sheet. Preserve source, date/version/validity and never silently overwrite conflicts.
2. Distinguish list, cash, financed and promotional prices and temporary discounts before mapping to production.
3. Repeated unit numbers across towers/blocks/stages. Unit 101 is not always unique; identity must include the relevant structural context.
4. Matrix-style lists where one row represents multiple floors/units (e.g. 101–1101) and must be expanded.
5. Partial availability lists: a missing unit must not automatically be considered sold.
6. Normalize commercial statuses such as Available/Reserved/Sold.
7. Correctly map interior, terrace, garden, rooftop, common, commercial and total areas.
8. Associate floorplans with typologies/units from brochure pages, images or filenames such as “1 bedroom 301–901”.
9. Drive links may point to a project or to a promoter root with multiple projects/folders; the operator may initially select the relevant source.
10. Retain creation/modification dates and validity periods for time-sensitive data.
EXPECTED WORKFLOW
1. Operator selects project and uploads files or selects/provides a Drive source.
2. System classifies source type/purpose/date/version.
3. AI plus deterministic parsing extracts facts and maps them to the Planorbe schema.
4. Each value retains provenance, source reference, confidence and relevant date/version.
5. Generate structured draft: Project → Typologies → Units.
6. Human Review UI allows edit/approve/reject and highlights low-confidence, missing or conflicting values with source evidence.
7. Approved data is sent to Gestor through an authenticated API.
8. Later updates generate a diff against current Gestor data (price/status/availability/new or missing units/promotions) for approval.
TECHNICAL APPROACH
We are open to AWS, Azure, OpenAI, Anthropic, Gemini or other suitable services. No custom model training is required. We prefer pragmatic, maintainable architecture, existing AI/document APIs where appropriate, deterministic code where reliability matters, low recurring costs, documented APIs and full source-code/repository ownership.
DEFINITION OF DONE
The system must process a representative set of our real heterogeneous sources, produce a reviewable structured draft with field-level provenance/confidence, handle the edge cases above, allow human correction/approval, push approved data through the Gestor API, generate update diffs, and be deployed with source code and technical documentation handed over.
PROPOSAL REQUIREMENTS
Provide 2–3 concrete examples of similar systems you personally or your team implemented. For each: inputs, structured output, technologies, your exact role, production status, and screenshots/demo/case study/repository excerpt/client reference where possible.
Also include proposed architecture, implementation time, fixed-price quotation for the COMPLETE solution with milestones, recurring infrastructure/API costs, main technical risks, explicit exclusions, and who will actually write/lead the code.
Shortlisted candidates will attend a technical video call with the person who will actually implement or lead the code, using one or two real Planorbe examples.
The budget is intentionally not disclosed. Please quote the complete solution based on the scope above.
Deschide pe Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Autentificare