← Live-Feed

Founding AI Platform Engineering Lead, Contract-to-Hire

Budget: $100.0 - $175.0 HOURLY / FULL_TIME ⭐ 4.81 (59) United States

python, kubernetes, artificial-intelligence

Bevorzugte Qualifikationen

  • Erfahrung: Experte
We are building a new multi-tenant AI platform that brings model access, policy controls, evaluation, and reusable agent systems into one dependable foundation. We need a founding engineering lead who can own the architecture, ship production code every week, and raise the level of a small cross-time-zone team. This is a contract-to-hire search for someone who wants a long-term, high-ownership role. The initial Upwork engagement gives both sides a practical way to work together before any longer-term arrangement is considered and formally approved. What you will own: - Define and evolve the common platform architecture across model access, inference, model lifecycle, evaluation, and agent runtime interfaces. - Turn ambiguous product and technical goals into executable boundaries, contracts, milestones, and acceptance tests. - Write and review production code, not only diagrams or strategy documents. - Establish secure multi-tenancy, identity propagation, observability, cost controls, release gates, and failure handling. - Make practical build-versus-buy and open-weight-versus-hosted decisions using quality, latency, cost, security, and portability evidence. - Lead design reviews, incident reviews, and technical hiring panels. - Create a clean handoff path across internal engineers, contractors, and product owners. You are likely a fit if you have: - Eight or more years in backend, platform, distributed-systems, or ML infrastructure engineering. - Staff, principal, founding, or early-engineer ownership of a production distributed system. - Personally built or operated LLM-era infrastructure under real traffic, such as an AI gateway, inference service, evaluation platform, model control plane, or agent platform. - Owned incidents, rollbacks, capacity decisions, observability, and measurable quality, latency, and cost tradeoffs. - Led a small team while remaining deeply hands-on in code and design. - Strong written communication and several hours of reliable overlap with US Eastern time. Experience with vLLM, TGI, TensorRT-LLM, KServe, Ray Serve, Kubernetes, OpenTelemetry, policy engines, or multi-provider model orchestration is useful, but named production ownership matters more than keyword coverage. To apply, answer every question below: 1. Name the most relevant production platform you personally owned. What traffic or scale did it serve, what code or architecture did you own, and what measurable outcome changed because of your work? 2. Describe one serious production incident or failed rollout. What did you diagnose, what did you change, and what permanent control did you add? 3. How would you separate policy eligibility from model or route selection in a multi-tenant AI platform? 4. Share one sanitized artifact you can walk through live: repository, code sample, design document, incident review, or system diagram. 5. Are you seeking a long-term full-time destination after an initial Upwork contract, and how many hours per week can you commit now?
Auf Upwork öffnen

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Anmelden