Build & deploy image + text lead-scoring system: XGBoost + CLIP image signal, CF Worker, GH Actions
Budget: -
HOURLY / PART_TIME
⭐ 4.95 (52)
United States
pandas, python, machine-learning, javascript
Budget
Open to proposals — please quote based on scope below. Competitive hourly rates or fixed bid.
## Overview
We're looking for a developer to build a lead-scoring pipeline for our platforms2. The system will score inbound leads using structured data (firmographic/CRM fields) combined with an image-based signal derived from photos submitted with consent during signup.
We have a specific architecture in mind and can provide technical direction, but implementation, integration, and deployment work is needed end-to-end.
## What we're looking to build
**Structured data scoring**
- A model trained on structured lead attributes (company info, funding stage, engagement data, and similar CRM-style fields) to output a numeric score predicting lead quality
- Training pipeline that can be re-run periodically as new outcome data comes in
- An API endpoint to submit a lead and receive a score
- An endpoint to log actual outcomes back into the system for future retraining
**Image-based signal**
- A pipeline that converts submitted profile photos into embeddings using an image model (CLIP or similar)
- Storage of embeddings in a vector-capable database
- A scoring mechanism comparing new photos against a reference set we'll provide, to produce a single presentation/formality score
- Support for multiple photos per lead (aggregated into one score)
- Longer-term: a trained classifier once we've accumulated a sufficient labeled image set, replacing or supplementing the reference-comparison approach
**Infrastructure preferences**
- Serverless/edge hosting (Cloudflare Workers) for API endpoints — no always-on servers if avoidable
- Postgres-based storage (Supabase) with vector search capability
- Scheduled retraining via CI/CD (GitHub Actions or similar) rather than a dedicated training server
- Cost-conscious throughout — this should run on pay-per-use infrastructure with minimal fixed monthly cost
## Scope boundaries (please read)
- Reference and training images will be provided by us; you're building the pipeline, not sourcing or curating the dataset.
- Specifics of our scoring criteria and business logic will be shared directly with the selected candidate, not detailed publicly here.
## Requirements
- Strong experience with Cloudflare Workers (or equivalent edge/serverless platform)
- Python + a standard ML library (scikit-learn, XGBoost, or similar) for structured-data modeling
- Experience with Postgres, ideally including pgvector or similar vector extensions
- Experience with CI/CD tooling (GitHub Actions or equivalent) for scheduled jobs
- Experience with image embedding APIs (CLIP via Replicate, Hugging Face, or similar)
## Nice to have
- Past work on lead-scoring, fraud-scoring, or similar predictive systems
- Experience with vector similarity search
## Deliverables
- Working scoring API for structured lead data, with a documented endpoint for logging outcomes
- Working image embedding + similarity scoring pipeline
- Automated retraining workflow for both components
- Documentation covering deployment, environment variables/secrets, and how to redeploy
## To apply
Please share relevant experience with (a) Cloudflare Workers or similar edge platforms, and (b) either structured ML scoring systems or image embedding/similarity work. A brief note on how you'd approach cost-efficient hosting for a low-to-moderate volume workload is welcome but not required.
Open job