← Zakázky

Senior AI Engineer, Autonomous Agents & LLM Evaluation

Rozpočet: $30.0 - $75.0 HOURLY / PART_TIME ⭐ 0.00 (0) United Kingdom

artificial-intelligence, python, bot-development, chatbot-development

You'll own the intelligence layer of our GenAI platform for e-commerce marketing. We run growth for a portfolio of in-house brands and for client accounts, and AI now sits at the center of how we produce, optimize, and scale that work: product content, ad copy, audience and campaign strategy, creative variation, and performance analysis. This is a senior, architecture-owning role. You'll design the tool-using autonomous agents that execute real marketing work, and you'll build the evaluation infrastructure that tells us, rigorously rather than anecdotally, whether those agents actually move the numbers. In performance marketing, "the copy sounds good" isn't the bar. Output has to be on-brand, conversion-oriented, safe to publish, and measurably tied to results across many brands with different voices and different audiences. You'll be the person who makes that true. You'll sit at the intersection of agent design and evaluation science, and you'll have the standing to say when something isn't ready to touch a live account. What You'll Own Agent architecture. Design and build tool-using autonomous agents (planning, tool selection, retrieval, multi-step reasoning, and graceful failure) that handle real marketing tasks: generating product descriptions and ad variants, analyzing campaign performance, researching audiences, and pulling from catalog, analytics, and ad-platform data. Decide where agents act autonomously and where a human approves before anything ships. Evaluation infrastructure. Architect LLM evaluation dashboards spanning quality, brand-safety, and robustness. Build eval harnesses (OpenAI Evals, TruLens, or custom) that run continuously so quality doesn't silently drift across dozens of brands and campaigns. Outcome correlation. Go beyond offline scoring. Design the instrumentation that connects eval metrics to real marketing outcomes such as CTR, conversion rate, ROAS, and engagement, so we know our quality scores actually predict performance rather than just reading well. Brand safety and robustness. Build testing for the failure modes that matter here: off-brand voice, false or unverifiable product claims, non-compliant advertising language, prompt injection, and behavior drift across model versions. Brand grounding. Ensure agents stay anchored to each brand's voice, guidelines, product catalog, and claims constraints, with traceability from output back to source, so a given brand's agent never sounds like another's. Technical leadership. Set the standard for how the team reasons about agent quality and evaluation. Document methodology so results are reproducible and repeatable across the portfolio. What We're Looking For 5+ years of production ML experience, including 2+ years hands-on with LLMs in production (not just prototypes) Demonstrated experience building tool-using or autonomous agent systems (planning, tool orchestration, multi-step reasoning) Deep experience designing LLM evaluation: quality, safety, and robustness metrics, and the harnesses to measure them (OpenAI Evals, TruLens, or custom) Strong ability to correlate model/eval metrics with downstream business outcomes Deep Python Solid AWS experience for deploying and scaling ML systems Clear written and verbal communication. You'll be defending evaluation methodology to technical and non-technical stakeholders, including marketing leads Strongly Preferred Experience building AI for marketing, advertising, e-commerce, or content-at-scale use cases Familiarity with RAG architectures and grounding LLM outputs in brand guidelines, catalogs, or authoritative sources Experience integrating with e-commerce or ad platforms and their APIs (catalog, analytics, ad managers) Experience with adversarial testing, red-teaming, or robustness evaluation of LLM systems Awareness of advertising-compliance and claims constraints across channels Why This Role This isn't a "wrap an API and generate some copy" job. You'll be building the agent and measurement infrastructure that lets us scale AI-driven marketing across many brands without quality, voice, or compliance falling apart. It's the layer that determines whether we can trust these systems on live client and in-house accounts. If you care about making AI systems provably good rather than plausibly good, this is that role.
Otevřít na Upwork