← Jobb

PyTorch / GPU ML Engineer — Deploy ESM-2 Protein Language Model

Budget: $1750.0 FIXED / ⭐ 0.00 (0) Belgium

pytorch, machine-learning, deep-learning, python, natural-language-processing, cuda

Föredragna kvalifikationer

  • Erfarenhet: Expert
We are building OSCAR, a protein R&D software platform that combines computational protein engineering, active learning, experimental data, and closed-loop scientific workflows. The OSCAR backend and benchmark framework already exist. For this contract, I am not looking for someone to redesign the platform or modify the benchmark methodology. I need an experienced PyTorch and GPU inference engineer to deploy a real ESM-2 protein language model behind a small HTTP inference service, connect it to OSCAR's existing adapter contract, verify the exact model checkpoint being used, and run an existing frozen benchmark. Scope The worker should: Run a real ESM-2 checkpoint on an NVIDIA GPU Accept batches of protein sequences over HTTP Return mean-pooled protein embeddings Correctly handle padding and special tokens Return exact model provenance Integrate with OSCAR's existing ESM adapter Pass the existing integration tests Run the provided GB1 benchmark without changing its methodology The basic workflow is: OSCAR benchmark to ESM-2 GPU inference worker to protein embeddings and model provenance to OSCAR benchmark result Required API response The service must return: Protein embeddings Model family: ESM-2 Exact model version, for example esm2_t33_650M_UR50D Exact checkpoint identifier or weights SHA-256 Pooling method: mean Layer used: -1 Embedding dimension: 1280 OSCAR will refuse to benchmark the model if the exact checkpoint cannot be identified. Skills required You should be comfortable with: Python PyTorch NVIDIA CUDA GPU inference Transformer models FAIR ESM or ESM-2 Hugging Face or equivalent model tooling FastAPI or similar Python model-serving APIs Docker Batching variable-length sequences Padding and attention masks Mean pooling GPU memory management Model checkpoint and version tracking Reproducible ML inference Protein-engineering experience is helpful but not required. Strong GPU and ML inference engineering is more important for this contract. Milestone 1 — Working ESM-2 GPU service Deliver: Worker source code Dockerfile or reproducible environment definition Dependency and version information Startup instructions Embed endpoint Health endpoint GPU batching Correct masking and mean pooling Exact ESM-2 checkpoint identification Checkpoint SHA-256 or equivalent exact identifier Model version Pooling method Layer Embedding dimension Successful OSCAR adapter and integration test results Example request and response The repository already contains the adapter contract and tests. Milestone 2 — Frozen benchmark run Once Milestone 1 is accepted: Connect OSCAR to the ESM-2 worker Run the existing GB1 benchmark Do not change the benchmark methodology Do not change train or test splits Do not change reference models Do not change evaluation thresholds Preserve the exact Git commit and model checkpoint used Deliver benchmark_report.json Deliver run logs Report GPU model and runtime Provide enough instructions for another engineer to reproduce the run The current reference benchmark at the primary 104-measurement budget is approximately: Position-only baseline — Spearman rho 0.626 plus or minus 0.015 OSCAR current GP — Spearman rho 0.153 plus or minus 0.112 ESM-2 plus ridge — result to be determined The goal is not to force ESM-2 to beat the baseline. A result below 0.626 is still a valid and useful result. Scientific integrity and reproducibility are more important than a favorable score. Important Please do not: Redesign OSCAR Rewrite the benchmark Change train or test splits after seeing results Tune the benchmark specifically to improve the ESM-2 result Substitute another model for ESM-2 without approval Use mock or synthetic embeddings Silently fall back to another checkpoint Change scientific thresholds to make the result look better If you believe there is a problem with the benchmark, report it before changing anything. Deployment I am open to deployment using: RunPod AWS GCP Lambda GPU Cloud Another appropriate NVIDIA GPU platform The deployment should be reproducible by another engineer. Please answer these questions when applying What PyTorch transformer models have you personally deployed on NVIDIA GPUs? Please describe the model, hardware, and how it was served. Have you worked with ESM, ESM-2, ProtT5, or another protein language model? If not, describe relevant transformer-inference experience. How would you create a mean-pooled protein embedding from padded ESM token representations while excluding padding and special tokens? How would you prove that an API response came from the exact model checkpoint requested rather than another checkpoint with the same model name? What factors determine the safe batch size for ESM-2 GPU inference? If the configured ESM-2 model fails to load, what should the service do? You receive an existing repository containing the adapter, tests, and benchmark. What would you inspect and run before modifying code? How would you prevent cached embeddings from checkpoint A from being reused after the worker switches to checkpoint B? Contract structure I prefer a small fixed-price first milestone. Milestone 1: Make the ESM-2 worker conform to OSCAR's existing API contract and pass the tests. Milestone 2: Run the frozen benchmark and deliver the reproducible result. There may be additional ML infrastructure and protein-model work afterward if the collaboration is successful. Please focus your application on hands-on PyTorch, CUDA, and model-serving experience rather than general AI-agent or prompt-engineering experience.
Öppna på Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Logga in