Senior AI/ML Infrastructure Engineer
Бюджет: $499.0
FIXED /
⭐ 5.00 (3)
United States
embedded-systems, microcontroller-programming, reverse-engineering, embedded-c
Preferred qualifications
- Experience: Expert
We are looking for a Senior AI/ML Infrastructure Engineer to design, deploy, and integrate a custom, privacy-first AI assistant ("Resi") directly into our urban planning and climate tech platform, Planner360.
Our infrastructure runs entirely on our own dedicated, bare-metal physical servers with zero cloud hosting dependencies. The goal is to build an air-gapped Retrieval-Augmented Generation (RAG) and local inference pipeline capable of processing complex, multi-format urban datasets—including unstructured files (PDFs, Word documents), structured data (Excel, SQL/NoSQL databases), and complex spatial/GIS layers (building footprints, infrastructure networks, map data).
The assistant must integrate cleanly via a custom backend API with Planner360, supporting structured inputs and output formatting that aligns with user requirements.
Infrastructure Profile
Your solution must be engineered and optimized specifically for our physical hardware cluster:
Server Count: 3x Physical Bare-Metal Servers
Processors: AMD Ryzen 5900X (per server)
Memory: 128GB RAM per server (Total system capacity: 384GB RAM)
Storage: 2x NVMe drives per server (ultra-fast storage for vector DB and caching)
Note: Because this is a CPU/RAM-optimized environment (rather than enterprise NVIDIA GPUs), the architecture must utilize efficient CPU runtimes (such as Ollama with GGUF/quantized weights) and intelligently distribute workloads across our 3 nodes.
Key Responsibilities
Local Inference Deployment: Configure, optimize, and run local model inference engines (Ollama / GGUF runtimes) across our 3-node physical server cluster using open-source weights (e.g., DeepSeek, Qwen).
Advanced RAG & Data Ingestion Pipeline: Build robust ingestion parsers for multi-format urban datasets, including text documents, tabular databases, and geospatial/GIS formats (GeoJSON, shapefiles, building footprints).
Vector Indexing & Storage: Set up and maintain a local vector database (Qdrant or Milvus) leveraging our fast NVMe drives for high-speed semantic search across city-scale data.
Planner360 API Integration: Develop a secure, high-throughput backend API (using FastAPI) to seamlessly connect the local AI engine with our existing Planner360 application interface.
Structured Output Mapping: Implement response-formatting logic so the LLM generates precise structured output blocks, code schemas, or JSON snippets that the Planner360 backend can programmatically compile back into user files (Excel, GIS layers, reports).
Tech Stack Experience Required
Inference Runtimes & Models: Ollama, GGUF/quantized model workflows, Hugging Face ecosystem.
Vector DBs & Orchestration: Qdrant, Milvus, LangChain, or LlamaIndex.
Backend & APIs: Python, FastAPI, Docker, RESTful APIs, WebSockets.
Data & Spatial Formats: Experience with Pandas, Excel processing, and geospatial/GIS data handling (GeoPandas, GeoJSON).
Infrastructure: Linux (Ubuntu Server), bare-metal CPU/RAM resource optimization, Docker Compose, containerized network isolation across multiple nodes.
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти