← Jobs

LLM Engineer to cut our AI costs (FastAPI + LangChain + RAG)

Budget: - HOURLY / FULL_TIME ⭐ 4.96 (16) India

artificial-intelligence, machine-learning, python, deep-learning

We run an AI chat assistant with some actions on a FastAPI backend. It uses LangChain, a RAG pipeline over a Milvus vector DB, and Claude models. It works, but our LLM cost per request is too high, mostly because some of our prompts are very large (one is around 12,000 input tokens per call). We need someone to bring that cost down without changing how the system actually behaves. That last part is the whole point — we don't want "shorter prompts that give worse answers." We want proof the answers stay the same. What we need done: 1. Go through our 4 main prompts and reduce token size where it's genuinely safe. We've already done an analysis that found real duplication (repeated examples, dead formatting instructions, etc.), so there's a starting point. 2. Build an automated test suite BEFORE making changes. Run our real queries through the current system, record what it does, then re-run after changes to prove decisions didn't change. English and Japanese both matter to us. 3. Evaluate 2–3 cheaper/alternative models against that same test suite, so we can see if switching models saves money while keeping quality. Deliver everything as reviewable pull requests with before/after numbers — token counts, cost, and test results. Our stack: Python, FastAPI, LangChain, Milvus, AWS Bedrock (Claude), some Node.js in front. A few things we care about: - You've worked on real production RAG/LLM systems, not just demos - You're comfortable using AI coding tools like Claude Code or Cursor — that's how we work - You can explain your reasoning to a non-technical founder in plain language When you apply, please skip the generic pitch. Just tell me: how would you prove a shorter prompt didn't hurt quality? That one answer tells me if you're the right person. This is a scoped project to start. If it goes well there's more ongoing work .
Openen op Upwork