AI Agent Architect – Master Agent and Micro-Agents for Elastic Stack Operations
Budget: $200.0
FIXED /
⭐ 0.00 (0)
USA
javascript, elasticsearch, python
Bevorzugte Qualifikationen
- Erfahrung: Fortgeschritten
# AI Agent Developer – Master Agent and Multi-Agent Automation Using Elastic Stack Data
## Project Overview
We are looking for an experienced **AI Agent Developer / Multi-Agent Systems Architect** to design and build an agent-based automation framework within our existing application environment.
The solution should include a central **Master Agent** that receives requests, analyzes the task, delegates work to specialized micro-agents, combines the results, and provides clear recommendations or performs approved actions.
The micro-agents will use application telemetry and operational data stored in the **Elastic Stack**, including Elasticsearch, Kibana, application logs, infrastructure logs, metrics, alerts, traces, and related operational data.
The primary objective is to automate application monitoring, troubleshooting, health analysis, utility activities, and functional validation.
This is not only a chatbot project. We need a production-oriented agent system that can analyze Elastic data, correlate events, identify issues, recommend remediation, and safely execute authorized operational workflows.
## Proposed Agent Architecture
### Master Agent
The Master Agent should act as the main orchestration layer and be responsible for:
* Receiving requests from users, applications, monitoring systems, or APIs.
* Understanding the intent and required workflow.
* Selecting the appropriate micro-agent or group of agents.
* Passing relevant context and data to each agent.
* Coordinating multi-step investigations.
* Combining results from multiple agents.
* Handling failures, retries, timeouts, and incomplete responses.
* Maintaining task history and execution context.
* Requesting human approval before performing sensitive actions.
* Generating a final incident summary, recommendation, or operational report.
### Micro-Agents
The initial solution may include the following specialized agents.
#### 1. Application Health Agent
* Analyze application availability, latency, errors, and performance.
* Review Elasticsearch logs, metrics, traces, and alerts.
* Detect abnormal behavior or health degradation.
* Compare current behavior against historical baselines.
* Identify affected applications, services, hosts, clusters, or regions.
* Generate an application-health summary.
#### 2. Infrastructure and Utility Agent
* Monitor CPU, memory, disk usage, JVM performance, file systems, network connectivity, and host availability.
* Identify capacity risks and infrastructure bottlenecks.
* Review service status and supporting platform dependencies.
* Recommend or execute approved utility actions.
* Validate system health after remediation.
#### 3. Functional Validation Agent
* Validate whether critical application functions are operating correctly.
* Review API responses, application transactions, logs, and workflow completion.
* Detect failed or incomplete business transactions.
* Execute approved synthetic tests or validation scripts.
* Compare expected and actual application behavior.
* Provide pass/fail results with supporting evidence.
#### 4. Log Analysis and Root-Cause Agent
* Search and analyze Elasticsearch indices.
* Correlate errors across applications and infrastructure components.
* Identify recurring error patterns.
* Analyze timestamps, request IDs, transaction IDs, hostnames, service names, and dependency failures.
* Rank likely root causes based on available evidence.
* Separate confirmed findings from assumptions.
#### 5. Alert Triage Agent
* Review incoming alerts and identify duplicate or related alerts.
* Reduce alert noise.
* Assign severity and business impact.
* Correlate alerts with recent deployments, infrastructure changes, and application events.
* Recommend escalation, suppression, or investigation.
* Create a concise incident-triage summary.
#### 6. Remediation Agent
* Recommend remediation steps based on the detected issue.
* Execute only preapproved and low-risk actions.
* Require human approval for production-impacting actions.
* Capture the command, API call, script, output, and final status.
* Verify whether the remediation resolved the issue.
* Provide rollback guidance when applicable.
#### 7. Reporting and Knowledge Agent
* Create daily or on-demand application-health reports.
* Summarize incidents, root causes, resolutions, and outstanding risks.
* Store reusable troubleshooting knowledge.
* Retrieve relevant procedures or previous incidents.
* Generate operational recommendations based on historical patterns.
## Key Technical Requirements
The selected developer should be able to design and implement:
* Master-agent and micro-agent orchestration.
* Agent routing and task delegation.
* Context and state management.
* Tool calling and structured outputs.
* Elasticsearch query generation and execution.
* Integration with Elasticsearch REST APIs.
* Kibana, alerting, observability, and dashboard integration.
* Log, metric, trace, and application-event analysis.
* Retrieval-augmented generation using operational documentation.
* API, webhook, and automation-script integrations.
* Human-in-the-loop approval workflows.
* Agent execution logging and audit trails.
* Role-based access control.
* Secrets and credential management.
* Retry, timeout, fallback, and error-handling mechanisms.
* Evaluation and testing of agent accuracy.
* Protection against hallucinated actions or unsupported conclusions.
* Cost, token, and model-usage monitoring.
## Possible Technology Stack
We are open to recommendations, but relevant experience may include:
* Elasticsearch
* Kibana
* Elastic Observability
* Elastic APM
* Logstash
* Beats or Elastic Agent
* Python
* FastAPI
* REST APIs
* LangChain
* LangGraph
* LlamaIndex
* OpenAI-compatible models
* Microsoft AutoGen
* CrewAI
* MCP-based tools
* Docker
* Kubernetes
* Git
* CI/CD pipelines
* ServiceNow, Jira, Slack, Teams, PagerDuty, or similar integrations
The final framework should be selected based on maintainability, security, scalability, and suitability for our environment—not simply because it is currently popular.
## Primary Responsibilities
The selected professional will:
1. Review our existing Elastic Stack and application environment.
2. Identify available indices, mappings, dashboards, alerts, APIs, and data sources.
3. Define the master-agent and micro-agent architecture.
4. Recommend the appropriate agent framework and language-model strategy.
5. Develop a secure proof of concept.
6. Connect the agents to Elastic Stack data.
7. Implement routing, tool calling, context management, and structured outputs.
8. Create at least three initial micro-agents.
9. Add human-approval controls for sensitive actions.
10. Implement agent monitoring, logging, and auditability.
11. Test accuracy using real operational scenarios.
12. Document the architecture, configuration, deployment, and support procedures.
13. Train our technical team on extending and maintaining the solution.
## Initial Use Cases
The first implementation should support several use cases such as:
* “Check the health of Application X for the last two hours.”
* “Identify the reason for increased API response time.”
* “Analyze the latest production errors and group them by root cause.”
* “Determine whether the issue is related to the application, database, network, or infrastructure.”
* “Validate whether all critical application functions are working.”
* “Compare today’s errors with the previous seven-day baseline.”
* “Check whether a recent deployment caused the incident.”
* “Provide the top three likely root causes with supporting evidence.”
* “Recommend remediation steps and request approval before executing them.”
* “Generate a daily application-health and risk report.”
## Expected Deliverables
### Phase 1: Discovery and Architecture
* Review of the current Elastic and application environment.
* Use-case assessment and prioritization.
* Data-source and integration inventory.
* Agent architecture diagram.
* Security and access-control design.
* Recommended technology stack.
* Implementation roadmap.
### Phase 2: Proof of Concept
* Working Master Agent.
* At least three functional micro-agents.
* Elasticsearch integration.
* Natural-language and structured-query support.
* Agent routing and delegation.
* Incident analysis and health-report generation.
* Basic user interface, API, or chat integration.
* Execution logs and audit trail.
### Phase 3: Production Readiness
* Authentication and role-based access control.
* Human-approval workflows.
* Production error handling and fallback logic.
* Agent-quality evaluation framework.
* Deployment scripts and configuration.
* Monitoring and cost tracking.
* Technical documentation.
* Knowledge-transfer sessions.
## Required Experience
Applicants should have strong experience in several of the following areas:
* Building production AI agents or multi-agent systems.
* Developing master-agent and worker-agent architectures.
* LangGraph, LangChain, AutoGen, CrewAI, or equivalent frameworks.
* Elasticsearch queries, mappings, indices, APIs, and aggregations.
* Elastic Observability, APM, logs, metrics, and alerts.
* Application and infrastructure monitoring.
* Root-cause analysis and incident automation.
* Python and API development.
* RAG and operational knowledge-base integration.
* Secure tool execution and human-in-the-loop controls.
* Docker, Kubernetes, and enterprise application environments.
* AI evaluation, hallucination reduction, and production monitoring.
Generic chatbot development experience without Elastic Stack or operational automation experience will not be sufficient.
## Questions Applicants Must Answer
Please answer each question clearly:
1. What production AI-agent or multi-agent systems have you built?
2. Have you implemented a master-agent and worker-agent architecture?
3. What agent frameworks have you used, and which one would you recommend for this project?
4. Describe your experience with Elasticsearch and Elastic Observability.
5. How would you allow agents to query Elasticsearch safely?
6. How would you prevent an agent from executing an incorrect production action?
7. How would you preserve context across a multi-step incident investigation?
8. How would you evaluate whether the root-cause analysis is accurate?
9. How would you manage agent permissions, credentials, and audit logs?
10. What information would you need from our environment before starting?
11. What can you deliver during the first 30 days?
12. Are you applying as an individual or an agency?
13. Please provide examples of similar implementations, architecture diagrams, GitHub samples, or case studies.
Applications containing only generic AI, chatbot, or prompt-engineering descriptions will not be considered.
## Engagement Model
We prefer to begin with a paid discovery and proof-of-concept phase. The initial engagement should focus on architecture, Elastic Stack integration, and a small set of high-value agents.
Based on successful completion of the proof of concept, the engagement may expand into production implementation, additional agents, integrations, deployment, monitoring, and long-term support.
## Preferred Candidate
The ideal candidate is a senior AI agent engineer or solution architect who understands both:
* AI-agent orchestration and production LLM systems.
* Enterprise application monitoring and Elastic Stack operations.
The candidate should be capable of building a controlled, auditable operational assistant—not an uncontrolled autonomous system.
Auf Upwork öffnen
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Anmelden