← Állások

LLM Infrastructure Specialist (Local / Self-Hosted Deployment)

Költségvetés: $35.0 - $45.0 HOURLY / PART_TIME ⭐ 5.00 (15) United States

Preferred qualifications

  • Location: Ukraine, Poland
  • Experience: Expert
About the role We're building an AI platform for commercial real estate workflows — document ingestion, AI-assisted extraction, and human-in-the-loop review — where security, data privacy, and auditability are first-class concerns. We're looking for a specialist who can deploy and operate large language models on our own infrastructure, so sensitive documents never have to leave our environment. You'll own the local LLM stack end to end: selecting models, standing up inference servers, optimizing them for our hardware, and making them reliable enough for production use. What you'll do Deploy open-weight LLMs (e.g. Llama, Mistral, Qwen, DeepSeek) on local / on-prem / private-cloud GPU infrastructure Set up and tune inference servers (vLLM, TGI, Ollama, llama.cpp, or similar) for throughput and latency Apply quantization and optimization techniques (GGUF, AWQ, GPTQ, etc.) to fit models to available hardware Expose models behind stable, OpenAI-compatible APIs for our application services Configure GPU environments (CUDA drivers, containerization, orchestration) Benchmark models for quality, speed, and cost; recommend the right model for each use case Establish monitoring, logging, and scaling for inference workloads Document the setup so the team can maintain and reproduce it What we're looking for Proven experience deploying LLMs locally or in a self-hosted/private environment (not just calling hosted APIs) Hands-on with at least one inference framework (vLLM, TGI, Ollama, llama.cpp, LM Studio, or equivalent) Solid understanding of GPU hardware, VRAM constraints, and model quantization trade-offs Comfortable with Linux, Docker/containers, and Python Able to reason about model selection: size vs. quality vs. speed vs. cost Nice to have Experience with fine-tuning, LoRA/PEFT, or RAG pipelines Kubernetes / GPU orchestration at scale Background in regulated or security-sensitive environments (data privacy, audit requirements) Familiarity with Azure / cloud GPU instances
Megnyitás Upworkön

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Bejelentkezés