LLM Engineer — Build Claude Pipeline That Classifies 50k Messages/Month (Python, Quo, Monday.com)
Бюджет: $5000.0
FIXED /
⭐ 5.00 (3)
United States
chatbot-development, artificial-intelligence, python
Бажана кваліфікація
- Досвід: Експерт
REQUIRED SCREENING QUESTION
Describe one system you built that processed thousands of real customer conversations in production. How did you measure its accuracy, and what was the number? Applications without a method and a figure will not be considered.
The role
We're a Los Angeles leasing company. Every lead reaches us by phone or SMS through Quo; our CRM is Monday.com. We have a working Python pipeline, built on the Claude API, that reads each conversation nightly, classifies the lead, and produces review pages for the owner. We're hiring a senior engineer to own it: raise its accuracy, make it reliable, and extend it from classification into follow-up and, later, a supervised SMS agent.
What exists today
SQLite mirror of the full Quo history (~500k messages, ~67k calls, ~36k call summaries, ~7k transcripts), refreshed nightly.
Monday.com schema layer: live ID verification, gated writes, per-write undo log.
Thread assembly per contact (texts + voicemails + transcripts + summaries), keyed by phone.
Nightly job: ingest → shape → Claude classification against a written rule set → source printout → attestation gate → triage and QC pages.
Rulings corpus (~370 owner decisions) with a regression test that can only tighten.
Templated follow-up sends, gated.
~20k lines of Python, 42 test files.
Volume
100 unique inbound numbers/day; ~54k messages and ~10k calls per 30 days; ~1,100 active CRM rows across 11 pipeline stages; multilingual callers; noisy speech-to-text.
The work
Evaluation & accuracy. Build the eval harness: labeled holdouts, per-class precision/recall, drift detection, regression on the rulings corpus. Current end-to-end agreement with human review is roughly 85%; the target is materially higher and measured continuously.
Reliability. The failure modes we've hit are the usual ones — silent empty inputs, entity resolution by name instead of ID, destructive partial updates, model paraphrase treated as fact. We want defensive engineering as a habit: verify dependencies, count sources, fail loudly, never write without a gate.
Rule application at scale. The classification rules are stated by the owner in plain language and change weekly. Design how those rules are captured, versioned, tested, and applied consistently across thousands of rows.
Follow-up automation. Status-driven cadences, suppression (opt-outs, quiet hours, DNC), approved templates, single outbound line.
Phase 2: supervised SMS agent. Answer questions, send listings, book tours — behind human approval until metrics justify autonomy.
Stack & integrations
Python, SQLite, Claude API (and Claude Code, which the pipeline shells out to), Quo REST API, Monday.com GraphQL. Prior Quo/Monday experience is not required; API/webhook fluency and the habit of reading vendor docs for sharp edges is.
Requirements
Shipped and operated a production LLM system on real customer conversations at volume.
Built evals for it and can talk about the numbers.
Strong Python; comfortable inheriting and refactoring an existing codebase rather than rewriting it.
Integration experience with third-party APIs and webhooks.
Nice to have: regulated messaging domains (housing, healthcare, finance), consent/compliance handling for outbound SMS.
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти