OpenClaw / Hermes Agent Engineer for Fleet Audit, Cost Reduction & Ongoing Ops
Бюджет: $35.0 - $80.0
HOURLY / FULL_TIME
⭐ 4.92 (9)
United States
Preferred qualifications
- Experience: Expert
About the work
View job posting online with proper formatting here: https://fountaincity.tech/jobs/openclaw-hermes-agent-engineer-for-fleet-audit-cost-reduction-ongoing-ops/
We are a technology studio that builds and operates autonomous AI agent systems. We run our own multi-agent platform on OpenClaw and Hermes, and we take on client fleets that have grown organically and now need real engineering underneath them.
We are bringing on a contractor to help deliver an engagement with a nonprofit client running a fleet of ~21 OpenClaw agents. The agents work and deliver real operational value, but the system was built conversationally rather than engineered, and it now shows it: high API spend, daily firefighting, memory and session issues, no staging environment, and no protection on an email pathway that agents read and act on.
The work starts with a two week assessment and then continues as ongoing operations and remediation. If the fit is good, this becomes recurring work across other client fleets.
We also have other areas you can extend into including, but not limited to:
Building and training open source LLMs
Building and managing backbone AI LLM platforms
Machine learning projects
Fleet deployment systems for LLMs at scale
Phase 1: Assessment (approximately 2 weeks)
Working from remote access to the client's staging machine, you will help produce a written findings and roadmap document covering:
Fleet inventory: full map of agents, roles, skills, crons, and dependencies
Cost forensics: where the monthly API spend actually goes, broken down by agent, model, and job
Deterministic vs LLM analysis: identifying work being done by live model calls that belongs in scripts, with savings estimates
State and memory review: tracing memory fill-up and session overload, defining the correct storage model
Agent and skill architecture review: consistency, modularity, and reusability across the fleet
Security assessment: prompt injection exposure through the inbox, credential handling, per-agent permission scope
Resilience review: staging and production separation, backup and recovery, monitoring
Immediate stabilization: safe quick fixes applied during the assessment where possible
Phase 2: Ongoing operations and remediation
Daily and weekly upkeep, monitoring, and support of the fleet
Moving deterministic work off the LLM into scripts and structured data
Re-architecting state storage away from Markdown and text files into a proper database
Hardening the email pathway against prompt injection
Standardizing agent identity structure and skill modularity across the fleet
Building out cost governance: per-agent attribution, spend caps, model-tier discipline
Staging and production separation on macOS hosts, plus backup, recovery, and observability
Scaling the fleet cleanly toward one assistant per employee plus department-level agents
Required experience
Hands-on production experience with OpenClaw and Hermes agent fleets, not just experimentation
Real software engineering background: Python and/or Node, version control, testing discipline, proper SDLC
Deep working knowledge of the LLM system primitives: skills, slash commands, hooks, subagents, and custom MCP servers. You should be able to build these from scratch and explain when each is the right tool
Agentic coding workflows (Claude Code, Codex or equivalent)
Certification: one or more completed AI engineering certifications. We give strong preference to the Anthropic Claude Certification Program (Claude Certified Developer or Claude Certified Architect, Foundations or Professional). Anthropic Academy course certificates and comparable credentials from other vendors also count
LLM cost optimization: model tiering, caching, context management, replacing model calls with deterministic code
Prompt injection and agent security: least-privilege permission scoping, credential handling, sandboxing
Database design for agent state, plus experience migrating file-based state into structured storage
Comfort with cron and launchd scheduling, Telegram, Slack, Discord integrations, and email pipelines
Required: hardware and local infrastructure
This fleet runs on physical Mac hardware, in addition to cloud instances. That is a deliberate choice and it means the person we hire has to be comfortable at the machine layer, not just the application layer.
macOS administration on Apple Silicon: Mac Mini and Mac Studio hosts, headless operation, launchd services, unattended reboot and auto-login behavior, keeping long-running agent processes alive
SSH and remote access: key management, hardened remote administration, tunneling, working confidently in a terminal on a machine you cannot physically touch
Local inference on Apple Silicon: Ollama, Qwen, LM Studio, MLX, or llama.cpp in production use. Model quantization tradeoffs, unified memory sizing, context window limits on local hardware, and knowing which workloads belong on a local model versus a hosted API call
Resource and thermal realities: memory pressure, disk capacity, throttling under sustained load, and sizing hardware for a growing fleet
Backup and recovery at the machine level: Time Machine or equivalent, snapshotting, and a tested path back from a dead host
Hybrid routing matters here. A meaningful part of the cost reduction work is deciding what runs locally on owned hardware, what runs deterministically in code, and what genuinely needs a frontier model API call.
Clear written English. A large part of this engagement is documentation that a non-technical client can read
Nice to have
Experience operating multi-agent systems at 20+ agents
Observability and monitoring tooling for agent fleets
Background working with nonprofits or handling sensitive personal data
To apply, please include
A specific example of an OpenClaw or Hermes fleet you have worked on: size, what broke, what you fixed
The largest LLM cost reduction you have delivered, with the before and after numbers and how you did it
How you would approach securing an agent that reads and acts on inbound email
Your certifications, with a verification link (Credly badge or equivalent) where you have one
A local inference setup you have run on Apple Silicon: hardware, models, quantization, and what you routed to it versus a hosted API
Your hourly rate and how many hours per week you can commit
Your time zone and overlap with US Pacific
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти