← Jobs

Build & deploy image + text lead-scoring system: XGBoost + CLIP image signal, CF Worker, GH Actions

Budget: - HOURLY / PART_TIME ⭐ 4.95 (52) United States

pandas, python, machine-learning, javascript

Budget Open to proposals — please quote based on scope below. Competitive hourly rates or fixed bid. ## Overview We're looking for a developer to build a lead-scoring pipeline for our platforms2. The system will score inbound leads using structured data (firmographic/CRM fields) combined with an image-based signal derived from photos submitted with consent during signup. We have a specific architecture in mind and can provide technical direction, but implementation, integration, and deployment work is needed end-to-end. ## What we're looking to build **Structured data scoring** - A model trained on structured lead attributes (company info, funding stage, engagement data, and similar CRM-style fields) to output a numeric score predicting lead quality - Training pipeline that can be re-run periodically as new outcome data comes in - An API endpoint to submit a lead and receive a score - An endpoint to log actual outcomes back into the system for future retraining **Image-based signal** - A pipeline that converts submitted profile photos into embeddings using an image model (CLIP or similar) - Storage of embeddings in a vector-capable database - A scoring mechanism comparing new photos against a reference set we'll provide, to produce a single presentation/formality score - Support for multiple photos per lead (aggregated into one score) - Longer-term: a trained classifier once we've accumulated a sufficient labeled image set, replacing or supplementing the reference-comparison approach **Infrastructure preferences** - Serverless/edge hosting (Cloudflare Workers) for API endpoints — no always-on servers if avoidable - Postgres-based storage (Supabase) with vector search capability - Scheduled retraining via CI/CD (GitHub Actions or similar) rather than a dedicated training server - Cost-conscious throughout — this should run on pay-per-use infrastructure with minimal fixed monthly cost ## Scope boundaries (please read) - Reference and training images will be provided by us; you're building the pipeline, not sourcing or curating the dataset. - Specifics of our scoring criteria and business logic will be shared directly with the selected candidate, not detailed publicly here. ## Requirements - Strong experience with Cloudflare Workers (or equivalent edge/serverless platform) - Python + a standard ML library (scikit-learn, XGBoost, or similar) for structured-data modeling - Experience with Postgres, ideally including pgvector or similar vector extensions - Experience with CI/CD tooling (GitHub Actions or equivalent) for scheduled jobs - Experience with image embedding APIs (CLIP via Replicate, Hugging Face, or similar) ## Nice to have - Past work on lead-scoring, fraud-scoring, or similar predictive systems - Experience with vector similarity search ## Deliverables - Working scoring API for structured lead data, with a documented endpoint for logging outcomes - Working image embedding + similarity scoring pipeline - Automated retraining workflow for both components - Documentation covering deployment, environment variables/secrets, and how to redeploy ## To apply Please share relevant experience with (a) Cloudflare Workers or similar edge platforms, and (b) either structured ML scoring systems or image embedding/similarity work. A brief note on how you'd approach cost-efficient hosting for a low-to-moderate volume workload is welcome but not required.
Open job