Senior Machine Learning Engineer
Budget: $30.0 - $70.0
HOURLY / FULL_TIME
⭐ 0.00 (0)
USA
machine-learning, python, pytorch, gradient-boosting-technique, docker
Preferred qualifications
- Talent type: Independent
- Experience: Expert
- English: Fluent
- Job Success: 90%+
- Rising Talent preferred
We are an early-stage company building machine learning and geospatial software that helps utilities predict underground pipe failure risk.
We are looking for a senior Machine Learning Engineer to take ownership of an existing survival-modeling framework and evolve it into a reliable, reproducible platform. You will own model architecture, experimentation, training, validation, release, data-quality controls and the technical roadmap.
You will work with complex, varied datasets and use sound judgment to identify meaningful insights and opportunities that are not always immediately apparent.
The day-to-day work spans Python, PyTorch, gradient boosting, geospatial data, GitHub Actions, Docker and cloud GPU training.
This is a hands-on ownership role on a small, senior team.
What You Will Own
• End-to-end ownership of the modeling lifecycle, from development and validation through deployment and ongoing performance monitoring
• Development of meaningful model inputs from diverse infrastructure, environmental, and geospatial datasets
• Evaluation and improvement of the existing modeling framework, including its architecture, training pipeline, validation approach and production readiness
• Model evaluation using calibration, time-dependent concordance, cost-weighted analysis, decision-curve analysis, subgroup performance and uncertainty quantification
• Validation design that appropriately addresses leakage, temporal splits, spatial autocorrelation, censoring, class imbalance and data limitations
• Explainability for utility engineers, executive leadership, governing boards and other nontechnical stakeholders through SHAP, partial dependence and per-asset reason codes
• Repeatable pipelines with versioned data, features, models and reproducible training runs
• Traceable lineage from model predictions back to their source records
• Data-quality controls for schema, range, freshness, volume, referential integrity and distributional drift
• Containerized cloud training, checkpointing, cost management and reproducibility across GPU environments
• The technical roadmap for the modeling platform as new systems, schemas and datasets are onboarded
Your first priority will be to understand and reproduce the existing modeling framework, assess its architecture and training pipeline, build on the foundation already in place, identify opportunities for improvement and establish a roadmap for taking full technical ownership.
Core Qualifications
We will ask candidates to provide evidence of the following:
• Five or more years of experience building and shipping production machine learning models, including ownership of at least one model from raw data through deployment and monitoring
• Survival analysis, time-to-event, reliability or competing-risks modeling, including experience with methods such as:
o Cox proportional hazards or Cox-Time models
o Accelerated failure-time models
o DeepSurv or DeepHit
o Random survival forests
o Discrete-time hazard models
• Advanced Python and experience with modern machine learning tools, including:
o PyTorch or TensorFlow
o scikit-learn
o XGBoost or LightGBM
o pandas or Polars
o NumPy
• Experience working with incomplete, inconsistent and multi-source data, with sound judgment regarding imputation, proxies, exclusions and when the available data cannot support the question
• Strong model-validation experience, including temporal validation, leakage prevention, class imbalance, censoring, spatial autocorrelation and subgroup performance
• Production MLOps experience, including reproducible training, model versioning, testing, deployment, monitoring and release controls
• Experience authoring CI/CD workflows, preferably using GitHub Actions
• Experience implementing automated data-quality and model-health monitoring in production
• Cloud training experience, including containerized workloads, checkpointing, dependency management, cost control and GPU utilization
• Strong Git and GitHub practices, including short-lived branches, reviewable pull requests, protected branches, automated status checks and collaborative code review
• Clear communication of technical tradeoffs, uncertainty, limitations and recommendations to both technical and nontechnical stakeholders
• Ownership of outcomes, early surfacing of problems
• Documentation practices that allow teammates to understand and operate the system independently, including model cards, architecture decision records, runbooks and actionable READMEs
Engineering and MLOps Experience
Candidates should have practical experience with most of the following:
CI/CD
• GitHub Actions workflows that lint, type-check, test, build containers, run training smoke tests and gate merges
• Reusable workflows, matrix builds, caching, secrets handling and controlled release environments
Testing for Machine Learning
• Unit tests for transformations and feature logic
• Contract tests for schemas
• Regression tests for model metrics
• Reproducible fixtures that allow training and validation to run locally and in CI
DataOps Observability
• Automated testing for schema, range, freshness, volume, referential integrity and distributional drift
• Pipeline instrumentation, alerting, traceable lineage and clear failure handling
• Experience with Great Expectations, Soda, dbt tests or equivalent tools
Cloud and DevOps
• Docker and environment parity
• Infrastructure as code, with Terraform preferred
• Cloud deployment and monitoring
• Dependency, credential and secrets hygiene
• Ability to evaluate and independently stand-up training environments across unfamiliar cloud or GPU providers
Strongly Preferred
• Geospatial and spatiotemporal data, including:
o PostGIS
o GeoPandas
o rasterio
o Spatial joins
o Coordinate systems and projections
o ESRI and ArcGIS interoperability
• Multi-GPU or distributed training using DDP, FSDP, DeepSpeed, Ray or equivalent tools
• Resumable checkpointing on spot or preemptible capacity
• Memory profiling, mixed precision and training-throughput optimization
• Kubernetes with GPU scheduling, Ray clusters or Slurm
• MLflow, Weights & Biases, DVC, feature stores or model registries
• Azure, Databricks, ADLS or Delta Lake
• Multi-tenant modeling across organizations with different schemas and data standards
• Utilities, asset management, infrastructure, reliability engineering or predictive maintenance
How We Work
We are small, senior and collaborative. Decisions are documented, assumptions are tested and technical concerns are surfaced early.
Data quality is treated as a product feature, not as a pre-demo cleanup task. Public utilities may make long-term capital decisions using the outputs of this work, so model uncertainty and limitations must be identified and communicated plainly.
We value people who take ownership, work independently, ask thoughtful questions and bring forward ideas rather than waiting for detailed instructions.
Engagement
• Part-time and project-based to start, approximately 15–30 hours per week
• Long-term relationship intended
• Potential transition into a full-time technical leadership role within 6–12 months
• A future full-time position is not guaranteed and we will communicate candidly about where the opportunity stands
Links to relevant repositories, architecture diagrams, technical writing, project summaries or other evidence of your work are welcome.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in