← Jobb

Lead Applied AI / LLM Systems Engineer - Secure Multi-Model Business Intelligence Platform

Budget: $1200.0 FIXED / ⭐ 4.88 (20) United Kingdom

docker, amazon-web-services, linux, artificial-intelligence, python, api, data-science, javascript, machine-learning

Welcome all, I'm seeking an experienced individual Lead Applied AI / LLM Systems Engineer to help review, harden and extend an existing early-stage prototype for a proprietary B2B artificial-intelligence platform. Note: This is not a basic chatbot, prompt-writing role, website-only project or unrestricted autonomous-agent product. The first stage is to create a secure, measurable and commercially viable internal pilot. Full product specifications and proprietary details will only be provided to the selected candidate after the required confidentiality, intellectual-property, security and contractor agreements have been signed. INITIAL RESPONSIBILITIES: The selected engineer will be expected to: 1. Review an existing founder-built/Replit prototype and assess its architecture, security, scalability and technical debt. 2. Design and implement a secure multi-tenant architecture so that one customer can never access another customer’s information. 3. Create strict company, user, role, department, room and source-level permissions. 4. Build a provider-neutral LLM gateway capable of securely using a selected choice of AI models. 5. Ensure all provider keys, OAuth credentials and secrets remain server-side and are stored securely. 6. Build structured, schema-validated AI outputs rather than relying on unstructured model responses. 7. Create an evidence system linking material AI outputs back to their exact authorised source passages. 8. Implement clear separation between: - confirmed facts; - inferred information; - recommendations; - proposed decisions; - approved decisions; and - actions. 9. Design context retrieval that filters information according to permissions before any information is sent to a model. 10. Implement cost-aware model routing so that inexpensive tasks use suitable lower-cost models and more capable models are used only where justified. 11. Implement hard execution limits including: - maximum input size; - maximum output size; - maximum model calls; - maximum tool calls; - maximum retries; - maximum execution time; and - maximum authorised task cost. 12. Build a prepaid internal usage and cost ledger. 13. Implement atomic cost reservation so a customer cannot begin a task unless sufficient prepaid usage capacity is available. 14. Prevent concurrent requests from spending the same allowance twice. 15. Build provider-usage reconciliation comparing estimated, reserved and actual cost. 16. Implement runtime monitoring for: - repeated model or tool calls; - semantic loops; - no measurable progress; - unsupported claims; - context growth; - budget overruns; - incorrect tenant or project access; and - attempted actions outside permission. 17. Build conditional verification so higher-risk outputs receive additional evidence checks, deterministic calculations, policy checks, second-model review or human approval where appropriate. 18. Create version-controlled prompt and model registries. 19. Build an evaluation framework capable of measuring: - decision extraction; - action extraction; - owner and deadline attribution; - evidence accuracy; - unsupported-claim rate; - contradiction detection; - correct abstention; - human correction rate; - latency; - model/tool cost; and - cost per successful task. 20. Create automated unit, integration, adversarial and end-to-end tests. 21. Produce architecture, security, deployment, testing, recovery and handover documentation. 22. Work continuously inside project-controlled GitHub, Replit/cloud and provider accounts. 23. Support future integrations with systems such as Slack, Microsoft Teams, email, calendars and document platforms, although these are not all required in the first milestone. REQUIRED EXPERIENCE: Applicants should have demonstrable experience with: - Production multi-tenant SaaS systems - Applied LLM systems - OpenAI and Anthropic APIs - TypeScript and/or Python - PostgreSQL - Background workers and queues - Retrieval and context-management systems - Structured LLM outputs and schema validation - Model routing and provider abstraction - LLM evaluation and regression testing - Usage metering and cost controls - OAuth and webhook security - Secrets management - Role-based access control - Audit logging and observability - GitHub branches, pull requests and CI/CD - Secure production deployment HIGHLY DESIRABLE EXPERIENCE: - Slack API integrations - Microsoft Graph / Teams integrations - Prompt caching and batch-processing optimisation - Agent-runtime or trajectory monitoring - Enterprise permission systems - Stripe subscriptions and prepaid usage systems - LLM security and prompt-injection defence - SOC 2-oriented engineering - Statistical evaluation of AI outputs FIRST PAID MILESTONE: The first milestone will be an architecture and foundation milestone rather than an open-ended build. Expected deliverables: 1. Written architecture assessment 2. Threat model 3. Data-model review 4. Multi-tenant security review 5. Provider-gateway design and baseline implementation 6. Usage-ledger and atomic-reservation design 7. Baseline evaluation runner 8. Cross-tenant access tests 9. Prompt/model version registry 10. Open-source and pre-existing-code disclosure 11. Prioritised defect list 12. Two-to-four-week implementation plan 13. Demonstration of the agreed baseline flow within the project-controlled repository OWNERSHIP, CONFIDENTIALITY AND ACCESS: Before receiving proprietary specifications or meaningful project access, the selected engineer must sign the required contractor development, confidentiality, intellectual-property assignment, data-protection, security and handover agreement. Non-negotiable conditions include: - All project-specific work product and assigned intellectual property belongs to the contracting company. - The engineer receives no ownership or equity. - All code must be committed to the project-controlled private GitHub repository. - No contractor-owned production repositories, domains, provider accounts, databases or cloud environments. - No subcontracting or agency substitution without prior written approval. - No provider secrets stored in source code. - No unapproved third-party confidential material. - Pre-existing code must be disclosed and approved before use. - Open-source dependencies must be documented and licence-compatible. - Complete source code, documentation, testing and handover are required. - No public portfolio, case study or disclosure without written permission. APPLICATION QUESTIONS: Please answer every question: 1. Describe a production LLM system you personally designed and state your exact contribution. 2. How would you prevent one customer’s data from appearing in another customer’s retrieval results or model context? 3. How would you implement atomic prepaid cost reservation before an LLM task begins? 4. How do you evaluate LLM output quality beyond subjective user ratings? 5. When would you use a second-model verifier, and when would that be unnecessary? 6. How would you detect semantic loops, repeated calls and lack of agent progress? 7. How would you calculate cost per successful business outcome? 8. How would you design model-provider fallback without disrupting the customer workflow? 9. How would you protect the platform from prompt injection contained inside uploaded documents or messages? 10. Which parts of this project would you simplify for the first controlled paid pilot? 11. Confirm that you will not subcontract any work without prior written approval. 12. Confirm that all project-specific work will be committed to the project-controlled repository and assigned under the signed agreement. SCREENING: Shortlisted candidates may be offered a paid, time-limited technical exercise using synthetic data. The exercise may include: - Reviewing a simplified meeting-analysis workflow - Identifying security and evaluation defects - Proposing a minimum production-safe architecture - Writing one cross-tenant security test - Writing one cost-reservation test - Explaining the major technical trade-offs ENGAGEMENT: - Individual applicants preferred - Remote - Initial paid milestone - Further work awarded through accepted milestones - Potential longer-term role for exceptional performance - No agencies or undisclosed subcontractors Fixed budget: $1,200 for completion of all deliverables listed in this job post, using and improving the existing Replit build rather than rebuilding from scratch. Applicants must confirm they accept the complete stated scope, budget and milestone acceptance criteria. No deliverable may be removed, substituted or treated as optional without my prior written approval. Milestones will be further discussed as potential shown can turn into full time.
Öppna på Upwork