Backend / Data Engineer Needed to Stabilize Existing Real Estate Scraper + Data Pipeline
Budżet: -
HOURLY / PART_TIME
⭐ 0.00 (0)
MEX
Preferowane kwalifikacje
- Doświadczenie: Ekspert
We are looking for an experienced backend/data engineer to take over and stabilize an existing real estate data pipeline project.
The system is already partially built and includes:
* Web scraper for a Mexico real estate source
* PostgreSQL database
* Data normalization
* Deduplication and update handling
* REST API
* Dockerized deployment
* Proxy integration framework
* Admin / monitoring tools
* Existing documentation and source code
The previous developer completed most of the initial foundation, but we now need someone to review the existing codebase, fix remaining issues, complete the proxy setup, and make the pipeline reliable enough for long-term use.
### Phase 1 – Recovery & Stabilization
The initial scope will include:
* Review and audit the existing codebase
* Identify and fix remaining scraper issues
* Validate and improve the rotating proxy integration
* Review retry, failover, cooldown, and rate-limit handling
* Improve scraper reliability and coverage
* Review data validation so incomplete or failed pages are not stored as valid listings
* Fix / clean existing bad records where necessary
* Validate normalization, deduplication, and update handling
* Review scraper run logs, statistics, and monitoring
* Confirm API and database outputs remain consistent
* Provide updated source code and documentation
We are **not starting from zero**. The goal is to inherit the existing system, understand it, and bring it to a stable production-ready state.
### Current Technology
The existing project is primarily based on:
* Python
* PostgreSQL
* FastAPI
* Docker
* Web scraping
* Rotating proxy infrastructure
Experience taking over and debugging another developer’s code is important.
### Potential Phase 2 – AI Real Estate Assistant
If Phase 1 goes well, there is a strong possibility of continuing with the same developer/team for a second phase.
The future AI assistant would use the real estate database created by this pipeline and may include:
* Natural-language property search
* AI-generated property recommendations
* Conversational real estate search
* Database-aware responses
* Lead / contact capture
* User search and conversation tracking
* Analytics on what buyers are looking for
* Potential agent recommendations / routing
* Standalone web application
* Embeddable chatbot/widget that can be added to our existing websites
The AI system should ultimately be able to work as its own tool while also being integrated into multiple websites.
### When Applying
Please include:
* Your experience with web scraping and data pipelines
* Experience taking over / debugging existing codebases
* Experience with rotating residential proxies and anti-blocking systems
* Experience with PostgreSQL, APIs, and Docker
* Any similar marketplace or real estate data projects
* Your approach to auditing and stabilizing an existing scraper
* Estimated time you would initially need to review the existing codebase before providing a final scope / estimate
Experience with AI / LLM systems is a plus because Phase 2 may continue with the same developer, but strong scraping and backend experience is the priority for Phase 1.
This is initially a defined recovery/stabilization project, with potential for significant ongoing work if the collaboration is successful.
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj