PyTorch / GPU ML Engineer — Deploy ESM-2 Protein Language Model
Presupuesto: $1750.0
FIXED /
⭐ 0.00 (0)
Belgium
pytorch, machine-learning, deep-learning, python, natural-language-processing, cuda
Cualificaciones preferidas
- Experiencia: Experto
We are building OSCAR, a protein R&D software platform that combines computational protein engineering, active learning, experimental data, and closed-loop scientific workflows.
The OSCAR backend and benchmark framework already exist.
For this contract, I am not looking for someone to redesign the platform or modify the benchmark methodology.
I need an experienced PyTorch and GPU inference engineer to deploy a real ESM-2 protein language model behind a small HTTP inference service, connect it to OSCAR's existing adapter contract, verify the exact model checkpoint being used, and run an existing frozen benchmark.
Scope
The worker should:
Run a real ESM-2 checkpoint on an NVIDIA GPU
Accept batches of protein sequences over HTTP
Return mean-pooled protein embeddings
Correctly handle padding and special tokens
Return exact model provenance
Integrate with OSCAR's existing ESM adapter
Pass the existing integration tests
Run the provided GB1 benchmark without changing its methodology
The basic workflow is:
OSCAR benchmark
to ESM-2 GPU inference worker
to protein embeddings and model provenance
to OSCAR benchmark result
Required API response
The service must return:
Protein embeddings
Model family: ESM-2
Exact model version, for example esm2_t33_650M_UR50D
Exact checkpoint identifier or weights SHA-256
Pooling method: mean
Layer used: -1
Embedding dimension: 1280
OSCAR will refuse to benchmark the model if the exact checkpoint cannot be identified.
Skills required
You should be comfortable with:
Python
PyTorch
NVIDIA CUDA
GPU inference
Transformer models
FAIR ESM or ESM-2
Hugging Face or equivalent model tooling
FastAPI or similar Python model-serving APIs
Docker
Batching variable-length sequences
Padding and attention masks
Mean pooling
GPU memory management
Model checkpoint and version tracking
Reproducible ML inference
Protein-engineering experience is helpful but not required.
Strong GPU and ML inference engineering is more important for this contract.
Milestone 1 — Working ESM-2 GPU service
Deliver:
Worker source code
Dockerfile or reproducible environment definition
Dependency and version information
Startup instructions
Embed endpoint
Health endpoint
GPU batching
Correct masking and mean pooling
Exact ESM-2 checkpoint identification
Checkpoint SHA-256 or equivalent exact identifier
Model version
Pooling method
Layer
Embedding dimension
Successful OSCAR adapter and integration test results
Example request and response
The repository already contains the adapter contract and tests.
Milestone 2 — Frozen benchmark run
Once Milestone 1 is accepted:
Connect OSCAR to the ESM-2 worker
Run the existing GB1 benchmark
Do not change the benchmark methodology
Do not change train or test splits
Do not change reference models
Do not change evaluation thresholds
Preserve the exact Git commit and model checkpoint used
Deliver benchmark_report.json
Deliver run logs
Report GPU model and runtime
Provide enough instructions for another engineer to reproduce the run
The current reference benchmark at the primary 104-measurement budget is approximately:
Position-only baseline — Spearman rho 0.626 plus or minus 0.015
OSCAR current GP — Spearman rho 0.153 plus or minus 0.112
ESM-2 plus ridge — result to be determined
The goal is not to force ESM-2 to beat the baseline.
A result below 0.626 is still a valid and useful result.
Scientific integrity and reproducibility are more important than a favorable score.
Important
Please do not:
Redesign OSCAR
Rewrite the benchmark
Change train or test splits after seeing results
Tune the benchmark specifically to improve the ESM-2 result
Substitute another model for ESM-2 without approval
Use mock or synthetic embeddings
Silently fall back to another checkpoint
Change scientific thresholds to make the result look better
If you believe there is a problem with the benchmark, report it before changing anything.
Deployment
I am open to deployment using:
RunPod
AWS
GCP
Lambda GPU Cloud
Another appropriate NVIDIA GPU platform
The deployment should be reproducible by another engineer.
Please answer these questions when applying
What PyTorch transformer models have you personally deployed on NVIDIA GPUs? Please describe the model, hardware, and how it was served.
Have you worked with ESM, ESM-2, ProtT5, or another protein language model? If not, describe relevant transformer-inference experience.
How would you create a mean-pooled protein embedding from padded ESM token representations while excluding padding and special tokens?
How would you prove that an API response came from the exact model checkpoint requested rather than another checkpoint with the same model name?
What factors determine the safe batch size for ESM-2 GPU inference?
If the configured ESM-2 model fails to load, what should the service do?
You receive an existing repository containing the adapter, tests, and benchmark. What would you inspect and run before modifying code?
How would you prevent cached embeddings from checkpoint A from being reused after the worker switches to checkpoint B?
Contract structure
I prefer a small fixed-price first milestone.
Milestone 1: Make the ESM-2 worker conform to OSCAR's existing API contract and pass the tests.
Milestone 2: Run the frozen benchmark and deliver the reproducible result.
There may be additional ML infrastructure and protein-model work afterward if the collaboration is successful.
Please focus your application on hands-on PyTorch, CUDA, and model-serving experience rather than general AI-agent or prompt-engineering experience.
Abrir en Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Entrar