Build a Go LLM gateway in front of our Python RAG service (6 weeks, likely extension)
Budget: $20.0 - $40.0
HOURLY / FULL_TIME
⭐ 5.00 (16)
United States
golang, amazon-web-services, docker, postgresql, next.js
Preferred qualifications
- Experience: Expert
We're a SaaS running Go microservices on Kubernetes (EKS), PostgreSQL and Kafka. This spring we shipped an AI assistant built on Python/FastAPI and OpenAI. It works but it's fragile under load, we have no rate limiting per customer, and our OpenAI bill is out of control. We need a proper gateway in front of it.
Scope:
- Go service that sits between our app and the LLM providers: routing across OpenAI, Anthropic and Bedrock, streaming, retries and fallback, per-customer token budgets, response caching
- Hardening the existing FastAPI RAG workers: job queue, timeouts, structured logging. Not a rewrite
- Prompt and retrieval regression checks in our GitHub Actions pipeline so changes get tested before deploy
- Helm chart, Docker image, deployed to our EKS cluster, Grafana dashboard for latency and cost
- Handover doc for our two backend engineers
Must have: 5+ years Go in production, Docker and Kubernetes beyond local, PostgreSQL, one LLM feature shipped to real users, enough Python/FastAPI to work in the existing service. Nice to have: built an LLM proxy before or used LiteLLM/Portkey and know their limits, pgvector, Kafka.
After this there's a longer engagement on the table
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in