Senior Machine Learning Engineer – ML Architecture, MLOps & Technical Leadership
Rozpočet: $20.0 - $60.0
HOURLY / FULL_TIME
⭐ 0.00 (0)
United Kingdom
sql, python, deep-learning, pytorch, apache-spark
Preferred qualifications
- Experience: Expert
We are looking for a Senior Machine Learning Engineer to lead the architecture and delivery of dependable machine-learning products from initial discovery through production operation.
The selected freelancer must combine strong hands-on modelling ability with software engineering, data engineering, MLOps, cloud architecture, security, GPU optimisation and technical leadership.
This role requires someone who can take ownership of complex ML systems, guide other engineers, manage technical and operational risks, communicate effectively with clients and ensure that solutions create measurable value without compromising reliability, security or data privacy.
Projects may include forecasting platforms, recommendation and ranking systems, advanced classification and anomaly detection, NLP, multimodal systems, deep-learning adaptation and reusable ML platforms.
Key Responsibilities:
- Own the end-to-end architecture of ML products, including data flows, training, evaluation, serving, monitoring, security and lifecycle management.
- Lead discovery sessions with clients and internal stakeholders.
- Convert ambiguous goals into measurable outcomes, technical requirements, constraints and phased delivery plans.
- Define credible baselines, modelling strategies and evaluation frameworks.
- Design temporal backtesting, ranking evaluation, calibration, uncertainty estimation and online experimentation.
- Review data and model designs for leakage, sampling bias, non-stationarity, target instability and operational risk.
- Lead feature engineering, data augmentation, transfer learning, fine-tuning and hyperparameter optimisation.
- Architect reliable batch, streaming and real-time inference systems.
- Define service-level expectations, observability, versioning, rollback and disaster-recovery approaches.
- Establish MLOps standards for experiment tracking, lineage, model registries, automated testing, CI/CD and approval gates.
- Optimise training and inference for throughput, latency, GPU memory usage and infrastructure cost.
- Guide infrastructure decisions across cloud, hybrid, local and on-premises environments.
- Design secure handling of client data, including tenant isolation, least-privilege access, encryption, key management and audit trails.
- Establish engineering standards for code quality, testing, documentation, reproducibility and architecture decisions.
- Conduct architecture, design and code reviews.
- Mentor machine-learning engineers and data scientists.
- Communicate technical trade-offs, risks and recommendations to clients, executives and engineering stakeholders.
- Challenge the use of ML where a simpler and more reliable solution would be more appropriate.
- Contribute to hiring, technical strategy, roadmap planning and build-versus-buy decisions.
The role has explicit responsibility for architecture, MLOps standards, infrastructure, security, mentoring and client-facing technical decisions.
Required Skills and Experience
Applicants should have:
- At least 6 years of relevant professional experience, including significant ownership of production ML systems.
- A track record of delivering and supporting multiple ML-enabled products in production.
- Strong hands-on Python and software-architecture skills.
- Advanced understanding of machine-learning model families, optimisation, generalisation, uncertainty and failure analysis.
- Strong experience with classical machine learning and deep learning.
- Practical experience with forecasting, recommendation systems or ranking models.
- Ability to review and guide technical work outside their main specialisation.
- Strong understanding of scalable data pipelines, data contracts, lineage and quality controls.
- Experience designing production-grade ML lifecycle systems.
- Strong cloud architecture experience on AWS, Azure or GCP.
- Experience with containerised or orchestrated workloads.
- Practical experience improving GPU utilisation, training efficiency, inference latency or infrastructure costs.
- Strong understanding of offline and online evaluation.
- Knowledge of model monitoring, drift detection, calibration and slice-based analysis.
- Strong understanding of security, privacy and governance for sensitive client data.
- Evidence of mentoring engineers, conducting reviews and improving engineering standards.
- Strong client-facing and stakeholder-management capability.
Relevant Technologies
Experience with several of the following is expected:
- Python and SQL
- FastAPI and typed APIs
scikit-learn
- XGBoost, LightGBM or CatBoost
- PyTorch or TensorFlow
- Hugging Face Transformers
- PEFT or LoRA
- Spark
- dbt
- Airflow, Prefect or Dagster
- Kafka
- Parquet, Delta Lake or Apache Iceberg
- MLflow or Weights & Biases
- Feature stores and model registries
- Docker and Kubernetes
- Terraform or similar infrastructure-as-code tools
- GitHub Actions, GitLab CI or Azure DevOps
- AWS, Azure or GCP
- Managed machine-learning services
- CUDA profiling
- Mixed precision
- Distributed training
- Quantisation
- ONNX or TensorRT
- Monitoring, logging and audit platforms
Additional Relevant Experience
The following would be advantageous:
- Computer vision, object detection, segmentation, - OCR or video analytics
- Multimodal or vision-language models
- LLM and RAG architecture
- LLM evaluation and guardrails
- Private or on-premises model deployment
- Vector databases
- Streaming inference
- Large-scale distributed training
- Causal inference
- Experimentation platforms
- Operations research or probabilistic programming
- Experience in healthcare, finance, government or other regulated sectors
Expected Deliverables
Depending on the project, deliverables may include:
- End-to-end ML architecture and technical design
- Data, training, evaluation and inference pipelines
- Forecasting, recommendation or ranking systems
- Production model APIs and services
- Batch, streaming or real-time inference architecture
- MLOps pipelines, model registries and approval workflows
- Automated testing and CI/CD configuration
- Cloud or on-premises deployment architecture
- Monitoring, drift detection and incident-response procedures
- GPU and infrastructure cost optimisation
- Security and privacy controls
- Architecture decision records
- Technical documentation and operational runbooks
- Design reviews, code reviews and mentoring support
Application Questions
Please answer the following five questions in your proposal:
1. Describe a production machine-learning system for which you owned or led the technical architecture. Explain the business objective, data flow, modelling approach, deployment architecture, monitoring and your specific contribution.
2. How would you architect a forecasting or recommendation platform that must remain reliable as data volumes, user traffic and business conditions change? Cover evaluation, scalability, drift, versioning and rollback.
3. Describe how you have established or improved MLOps practices. Include experiment tracking, model registries, testing, CI/CD, approval gates, monitoring and controlled promotion between environments.
4. Describe a difficult architectural, production or model-risk decision you led. What trade-offs did you consider, how did you communicate the decision, and what was the outcome?
5. Provide examples of relevant technical leadership and production ML work. This may include architecture diagrams, GitHub repositories, case studies, forecasting systems, recommendation systems, MLOps platforms, cloud deployments or mentoring responsibilities.
Otvoriť na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Prihlásiť