LLM Infrastructure Specialist (Local / Self-Hosted Deployment)
Budżet: $35.0 - $45.0
HOURLY / PART_TIME
⭐ 5.00 (15)
United States
Preferowane kwalifikacje
- Lokalizacja: Ukraine, Poland
- Doświadczenie: Ekspert
About the role
We're building an AI platform for commercial real estate workflows — document ingestion, AI-assisted extraction, and human-in-the-loop review — where security, data privacy, and auditability are first-class concerns. We're looking for a specialist who can deploy and operate large language models on our own infrastructure, so sensitive documents never have to leave our environment.
You'll own the local LLM stack end to end: selecting models, standing up inference servers, optimizing them for our hardware, and making them reliable enough for production use.
What you'll do
Deploy open-weight LLMs (e.g. Llama, Mistral, Qwen, DeepSeek) on local / on-prem / private-cloud GPU infrastructure
Set up and tune inference servers (vLLM, TGI, Ollama, llama.cpp, or similar) for throughput and latency
Apply quantization and optimization techniques (GGUF, AWQ, GPTQ, etc.) to fit models to available hardware
Expose models behind stable, OpenAI-compatible APIs for our application services
Configure GPU environments (CUDA drivers, containerization, orchestration)
Benchmark models for quality, speed, and cost; recommend the right model for each use case
Establish monitoring, logging, and scaling for inference workloads
Document the setup so the team can maintain and reproduce it
What we're looking for
Proven experience deploying LLMs locally or in a self-hosted/private environment (not just calling hosted APIs)
Hands-on with at least one inference framework (vLLM, TGI, Ollama, llama.cpp, LM Studio, or equivalent)
Solid understanding of GPU hardware, VRAM constraints, and model quantization trade-offs
Comfortable with Linux, Docker/containers, and Python
Able to reason about model selection: size vs. quality vs. speed vs. cost
Nice to have
Experience with fine-tuning, LoRA/PEFT, or RAG pipelines
Kubernetes / GPU orchestration at scale
Background in regulated or security-sensitive environments (data privacy, audit requirements)
Familiarity with Azure / cloud GPU instances
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj