AgentOps: Managing AI Agent Fleets in Production
Notice: This is an educational guide with illustrative code examples. It does not execute code, require credentials, or install dependencies. All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required to get started.
IBM coined "AgentOps" as a formal discipline in early 2026 for a reason that is now obvious: the bottleneck in enterprise AI shifted from building agents to keeping them alive. LangSmith Fleet launched in March 2026. Gartner projects 40% of agentic AI projects will fail by 2027 — not because the models are bad, but because nobody wrote the ops playbook. A Fortune 500 company leaked $400 million in cloud spend from unbudgeted agent compute last quarter. The EU AI Act's high-risk obligations become enforceable on August 2, 2026, and the Spanish DPA has already ruled that "greater technical autonomy does not reduce legal responsibility." This guide is the ops playbook. It covers the full agent lifecycle — provision, deploy, monitor, scale, retire — across fleets of tens to thousands of agents, all wired to the GreenHelix A2A Commerce Gateway's 128 tools across 15 services. You will build a Fleet Commander dashboard that consolidates provisioning, observability, cost control, SLA enforcement, scaling, and governance into a single operational view. Every chapter has working Python code, architecture diagrams, decision trees, and checklists. By the end, you will have the infrastructure to run agent fleets the way SREs run microservice fleets: with error budgets, automated escalation, cost guardrails, and an audit trail that holds up under regulatory scrutiny.
Getting started: All examples in this guide work with the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required.
What You'll Learn
- Chapter 1: The AgentOps Discipline
- Chapter 2: Fleet Provisioning and Identity at Scale
- Chapter 3: Observability for Agent Fleets
- Chapter 4: Cost-Aware Model Routing and FinOps
- Chapter 5: SLA Enforcement and Automated Escalation
- Chapter 6: Fleet Scaling Patterns
- Chapter 7: Governance, Audit, and EU AI Act Compliance
- Chapter 8: The Fleet Commander Dashboard
- What You Get
Full Guide
AgentOps: The Practitioner's Guide to Managing AI Agent Fleets in Production
IBM coined "AgentOps" as a formal discipline in early 2026 for a reason that is now obvious: the bottleneck in enterprise AI shifted from building agents to keeping them alive. LangSmith Fleet launched in March 2026. Gartner projects 40% of agentic AI projects will fail by 2027 — not because the models are bad, but because nobody wrote the ops playbook. A Fortune 500 company leaked $400 million in cloud spend from unbudgeted agent compute last quarter. The EU AI Act's high-risk obligations become enforceable on August 2, 2026, and the Spanish DPA has already ruled that "greater technical autonomy does not reduce legal responsibility." This guide is the ops playbook. It covers the full agent lifecycle — provision, deploy, monitor, scale, retire — across fleets of tens to thousands of agents, all wired to the GreenHelix A2A Commerce Gateway's 128 tools across 15 services. You will build a Fleet Commander dashboard that consolidates provisioning, observability, cost control, SLA enforcement, scaling, and governance into a single operational view. Every chapter has working Python code, architecture diagrams, decision trees, and checklists. By the end, you will have the infrastructure to run agent fleets the way SREs run microservice fleets: with error budgets, automated escalation, cost guardrails, and an audit trail that holds up under regulatory scrutiny.
Getting started: All examples in this guide work with the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required.
Table of Contents
- The AgentOps Discipline
- Fleet Provisioning and Identity at Scale
- Observability for Agent Fleets
- Cost-Aware Model Routing and FinOps
- SLA Enforcement and Automated Escalation
- Fleet Scaling Patterns
- Governance, Audit, and EU AI Act Compliance
- The Fleet Commander Dashboard
Chapter 1: The AgentOps Discipline
Why Agent Fleets Are Not Microservices
DevOps manages deterministic services. MLOps manages model training pipelines. Neither discipline handles the unique problem of autonomous software agents that make decisions, spend money, negotiate with counterparties, and degrade in ways that are stochastic rather than binary. A microservice either responds or it does not. An agent can respond with confident, plausible, expensive nonsense — and you will not know until the invoice arrives or the SLA breach notification lands.
The fundamental differences:
| Dimension | Microservices (DevOps) | Models (MLOps) | Agents (AgentOps) |
|---|---|---|---|
| Failure mode | Crash / timeout | Accuracy drift | Behavioral drift, cost explosion, SLA violation |
| State | Stateless or DB-backed | Model weights | Identity, wallet, reputation, active SLAs |
| Scaling unit | Container replicas | GPU hours | Capability + trust + budget |
| Rollback | Deploy previous image | Revert model version | Revoke access, freeze wallet, reassign tasks |
| Cost model | Compute per request | Training + inference | Token spend + tool calls + counterparty fees |
| Compliance | SOC2, PCI | Model cards | EU AI Act Article 12, audit trails, human oversight |
The Agent Lifecycle
Every agent in a fleet moves through five phases. GreenHelix tools map directly to each:
┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ PROVISION │────▶│ DEPLOY │────▶│ MONITOR │
│ │ │ │ │ │
│ register │ │ register │ │ submit │
│ _agent │ │ _service │ │ _metrics │
│ create │ │ create_sla │ │ get │
│ _wallet │ │ record │ │ _analytics │
│ build_claim │ │ _transaction│ │ check_sla │
│ _chain │ │ │ │ _compliance │
└─────────────┘ └─────────────┘ └─────────────┘
▲ │
│ ▼
┌─────────────┐ ┌─────────────┐
│ RETIRE │◀───────────────────────│ SCALE │
│ │ │ │
│ freeze │ │ search │
│ wallet │ │ _services │
│ revoke keys │ │ best_match │
│ archive │ │ estimate │
│ audit trail │ │ _cost │
└─────────────┘ └─────────────┘
The GreenHelix Tool Surface
All operations go through a single endpoint:
import requests
base_url = "https://api.greenhelix.net/v1"
api_key = "your-api-key"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {api_key}"
# Every tool call follows this pattern
resp = session.post(f"{base_url}/v1", json={
"tool": "tool_name",
"input": {"param": "value"}
})
result = resp.json()
This uniformity is what makes fleet management tractable. You are not integrating 15 different APIs with 15 different auth mechanisms and 15 different error formats. You are calling one endpoint, one auth header, one JSON envelope, 128 tools. The fleet controller you build in Chapter 8 exploits this to treat every GreenHelix operation as a uniform unit of work.
The AgentOps Maturity Model
Score your organization:
| Level | Name | Characteristics |
|---|---|---|
| 0 | Ad Hoc | Agents run on developer laptops, no monitoring, manual key management |
| 1 | Provisioned | Agents registered with identities and wallets, basic logging |
| 2 | Observable | Per-agent dashboards, SLOs defined, cost tracking active |
| 3 | Governed | SLA enforcement automated, audit trails, escalation pipelines |
| 4 | Autonomous | Self-scaling fleets, budget-aware routing, compliance continuous |
This guide takes you from Level 0 to Level 4. Each chapter corresponds to a maturity jump.
The Cost of Getting AgentOps Wrong
The numbers are unforgiving. The $400M Fortune 500 cloud leak was not a single dramatic failure — it was thousands of agents, each overspending by small amounts, compounding over months with no aggregation layer to surface the trend. A mid-size fintech reported a $2.3M surprise bill from a fleet of 80 research agents that were supposed to cost $12K/month. The root cause: each agent retried failed tool calls indefinitely because nobody configured a retry ceiling, and a downstream API had a 72-hour partial outage. The agents kept hammering the endpoint, accumulating token charges on every attempt. In both cases, the technology worked exactly as instructed. The agents completed their tasks. They just did it at 50x the expected cost because nobody was watching.
These are not engineering failures. They are operational failures — the same class of failure that drove the DevOps and SRE movements in the 2010s. The difference is that DevOps had a decade to develop its playbook. AgentOps does not. The EU AI Act deadline is in months. Gartner's 40% failure projection is for next year. The teams that build their ops layer now will be the ones that survive the Gartner trough.
Who This Guide Is For
This guide is written for platform engineers, SREs, and infrastructure leads who have been handed a fleet of AI agents and told to "make them production-ready." You know how to run Kubernetes clusters, set up Prometheus dashboards, and configure PagerDuty escalation policies. What you do not know — yet — is how to apply those instincts to software that makes its own decisions, spends its own money, and degrades in ways that look like success until the SLA report comes in.
You do not need to be an ML engineer. You do not need to understand transformer architectures or prompt engineering. You need to understand operational discipline — and this guide translates that discipline into the agent domain.
Chapter 2: Fleet Provisioning and Identity at Scale
The Provisioning Pipeline
Provisioning a single agent is trivial. Provisioning 200 agents with correct identities, funded wallets, appropriate tier access, reputation seeds, and SLA contracts — without a single credential collision or orphaned wallet — requires a pipeline.
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ MANIFEST │───▶│ REGISTER │───▶│ WALLET │───▶│ TRUST │───▶│ VERIFY │
│ │ │ │ │ │ │ │ │ │
│ YAML spec│ │ identity │ │ create + │ │ claim │ │ health │
│ per agent│ │ + keys │ │ fund │ │ chains │ │ check │
└──────────┘ └──────────┘ └──────────┘ └──────────┘ └──────────┘
Agent Manifest Format
Define each agent in a declarative manifest before touching any API:
# fleet-manifest.yaml
fleet:
name: "customer-support-fleet"
environment: "production"
agents:
- id_prefix: "cs-triage"
count: 10
tier: "pro"
initial_balance: 50.00
capabilities: ["text-classification", "sentiment-analysis"]
trust_claims: ["response-time-p99-under-2s", "accuracy-above-95"]
sla_template: "standard-support"
- id_prefix: "cs-escalation"
count: 3
tier: "pro"
initial_balance: 200.00
capabilities: ["complex-reasoning", "multi-turn-dialogue"]
trust_claims: ["human-supervised", "pii-compliant"]
sla_template: "premium-support"
Programmatic Registration
import requests
import uuid
import time
base_url = "https://api.greenhelix.net/v1"
api_key = "your-fleet-admin-key"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {api_key}"
def provision_agent(id_prefix, index, tier, initial_balance, capabilities, trust_claims):
"""Provision a single agent: register, wallet, trust chain."""
agent_id = f"{id_prefix}-{index:04d}-{uuid.uuid4().hex[:8]}"
# Step 1: Register the agent identity
resp = session.post(f"{base_url}/v1", json={
"tool": "register_agent",
"input": {
"agent_id": agent_id,
"tier": tier,
"capabilities": capabilities,
"metadata": {
"fleet": id_prefix,
"provisioned_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
"environment": "production"
}
}
})
if resp.status_code != 200:
raise RuntimeError(f"Registration failed for {agent_id}: {resp.text}")
registration = resp.json()
# Step 2: Create and fund wallet
resp = session.post(f"{base_url}/v1", json={
"tool": "create_wallet",
"input": {
"agent_id": agent_id,
"initial_balance": str(initial_balance)
}
})
if resp.status_code != 200:
raise RuntimeError(f"Wallet creation failed for {agent_id}: {resp.text}")
wallet = resp.json()
# Step 3: Bootstrap trust via claim chain
resp = session.post(f"{base_url}/v1", json={
"tool": "build_claim_chain",
"input": {
"agent_id": agent_id,
"claims": trust_claims
}
})
if resp.status_code != 200:
raise RuntimeError(f"Claim chain failed for {agent_id}: {resp.text}")
return {
"agent_id": agent_id,
"wallet_id": wallet.get("wallet_id"),
"balance": initial_balance,
"claims": trust_claims
}
def provision_fleet(manifest):
"""Provision an entire fleet from a manifest."""
results = {"succeeded": [], "failed": []}
for agent_spec in manifest["fleet"]["agents"]:
for i in range(agent_spec["count"]):
try:
agent = provision_agent(
id_prefix=agent_spec["id_prefix"],
index=i,
tier=agent_spec["tier"],
initial_balance=agent_spec["initial_balance"],
capabilities=agent_spec["capabilities"],
trust_claims=agent_spec["trust_claims"]
)
results["succeeded"].append(agent)
except RuntimeError as e:
results["failed"].append({
"index": i,
"prefix": agent_spec["id_prefix"],
"error": str(e)
})
return results
API Key Management at Scale
Never share a single API key across agents. Each agent gets its own key scoped to its tier and capabilities. Store keys in a secrets manager (Vault, AWS Secrets Manager, GCP Secret Manager) and inject at deploy time.
Key rotation pattern:
def rotate_agent_key(agent_id, old_key):
"""Rotate API key with zero-downtime overlap window."""
# Generate new key (via your identity provider)
# Configure agent to accept both keys during overlap
# Verify new key works
resp = session.post(f"{base_url}/v1", json={
"tool": "register_agent",
"input": {
"agent_id": agent_id,
"rotate_key": True
}
})
new_key = resp.json().get("api_key")
# Update secrets manager
# Wait for propagation (60s overlap window)
# Revoke old key
return new_key
Reputation Seeding
New agents start with zero reputation, which creates a cold-start problem — no one wants to transact with an unproven agent. Reputation seeds solve this:
def seed_reputation(agent_id, seed_metrics):
"""Submit initial metrics to bootstrap reputation."""
resp = session.post(f"{base_url}/v1", json={
"tool": "submit_metrics",
"input": {
"agent_id": agent_id,
"metrics": seed_metrics
}
})
return resp.json()
# Seed a triage agent with baseline metrics
seed_reputation("cs-triage-0001-a1b2c3d4", {
"response_time_ms": 450,
"accuracy_rate": 0.96,
"uptime_rate": 0.999,
"tasks_completed": 0,
"error_rate": 0.0
})
Fleet Provisioning Checklist
- All agents have unique IDs with fleet prefix for grouping
- Each agent has its own API key stored in secrets manager
- Wallets are funded with initial balance matching expected burn rate
- Trust claims are registered and verifiable
- Tier access matches required tool permissions
- Agent metadata includes fleet name, environment, and provisioning timestamp
- Key rotation schedule is configured (every 90 days minimum)
- Provisioning script is idempotent (safe to re-run)
- Failed provisioning steps trigger cleanup (no orphaned wallets)
- Fleet manifest is version-controlled
Chapter 3: Observability for Agent Fleets
The Three Pillars, Adapted for Agents
Traditional observability has three pillars: metrics, logs, traces. Agent observability adds two more: behavioral signals (what the agent decided to do and why) and economic signals (what it cost and who paid). Without these two additions, you can tell that your agent is up and fast but not that it is spending $14 per task on a job budgeted at $2.
Per-Agent Metrics Collection
import requests
import time
base_url = "https://api.greenhelix.net/v1"
api_key = "your-api-key"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {api_key}"
def collect_agent_metrics(agent_id):
"""Collect comprehensive metrics for a single agent."""
metrics = {}
# Performance metrics
resp = session.post(f"{base_url}/v1", json={
"tool": "get_analytics",
"input": {
"agent_id": agent_id,
"time_range": "1h",
"metrics": ["latency_p50", "latency_p99", "error_rate",
"tool_call_volume", "task_completion_rate"]
}
})
if resp.status_code == 200:
metrics["performance"] = resp.json()
# Reputation and trust
resp = session.post(f"{base_url}/v1", json={
"tool": "get_agent_reputation",
"input": {"agent_id": agent_id}
})
if resp.status_code == 200:
metrics["reputation"] = resp.json()
# Financial position
resp = session.post(f"{base_url}/v1", json={
"tool": "get_balance",
"input": {"agent_id": agent_id}
})
if resp.status_code == 200:
metrics["balance"] = resp.json()
# SLA compliance
resp = session.post(f"{base_url}/v1", json={
"tool": "check_sla_compliance",
"input": {"agent_id": agent_id}
})
if resp.status_code == 200:
metrics["sla"] = resp.json()
return metrics
Fleet Health Aggregation
Individual agent metrics are necessary but insufficient. You need fleet-level aggregation to spot systemic issues — a fleet where 30% of agents have degraded latency is a different problem than one agent having a bad day.
┌─────────────────────────────────────────────────────────┐
│ FLEET HEALTH DASHBOARD │
├─────────────┬─────────────┬─────────────┬───────────────┤
│ Agents │ Healthy │ Degraded │ Critical │
│ Total: 200 │ ███ 174 │ ██ 21 │ █ 5 │
├─────────────┼─────────────┼─────────────┼───────────────┤
│ P99 Lat. │ Fleet Err │ SLA Comp. │ Burn Rate │
│ 1,240 ms │ Rate: 2.1% │ 97.3% │ $42.10/hr │
├─────────────┴─────────────┴─────────────┴───────────────┤
│ ▁▂▃▄▅▆▇█▇▆▅▄▃▂▁ Tool call volume (24h) │
│ ▕████████████████▏ Peak: 14:00 UTC │
├─────────────────────────────────────────────────────────┤
│ ALERTS │
│ ⚠ cs-triage-0042: latency P99 > 3000ms (15 min) │
│ ⚠ cs-escalation-0002: balance below $20 threshold │
│ ✗ cs-triage-0187: SLA breach — response time │
└─────────────────────────────────────────────────────────┘
def aggregate_fleet_health(agent_ids):
"""Aggregate metrics across a fleet of agents."""
fleet_metrics = {
"total": len(agent_ids),
"healthy": 0,
"degraded": 0,
"critical": 0,
"total_burn_rate": 0.0,
"sla_compliant": 0,
"latencies_p99": [],
"error_rates": [],
"alerts": []
}
for agent_id in agent_ids:
metrics = collect_agent_metrics(agent_id)
perf = metrics.get("performance", {})
balance = metrics.get("balance", {})
sla = metrics.get("sla", {})
error_rate = perf.get("error_rate", 0)
latency_p99 = perf.get("latency_p99", 0)
fleet_metrics["latencies_p99"].append(latency_p99)
fleet_metrics["error_rates"].append(error_rate)
# Classify agent health
if error_rate > 0.10 or latency_p99 > 5000:
fleet_metrics["critical"] += 1
fleet_metrics["alerts"].append({
"agent_id": agent_id,
"severity": "critical",
"reason": f"error_rate={error_rate}, p99={latency_p99}ms"
})
elif error_rate > 0.05 or latency_p99 > 3000:
fleet_metrics["degraded"] += 1
fleet_metrics["alerts"].append({
"agent_id": agent_id,
"severity": "warning",
"reason": f"error_rate={error_rate}, p99={latency_p99}ms"
})
else:
fleet_metrics["healthy"] += 1
# SLA tracking
if sla.get("compliant", False):
fleet_metrics["sla_compliant"] += 1
# Cost tracking
fleet_metrics["total_burn_rate"] += float(balance.get("burn_rate_per_hour", 0))
fleet_metrics["sla_compliance_pct"] = (
fleet_metrics["sla_compliant"] / fleet_metrics["total"] * 100
if fleet_metrics["total"] > 0 else 0
)
return fleet_metrics
Alerting Pipeline
Define alert rules that escalate from notification to automated remediation:
| Severity | Condition | Action |
|---|---|---|
| Info | P99 latency > 2x baseline for 5 min | Log, notify Slack |
| Warning | Error rate > 5% for 10 min | Notify on-call, reduce traffic share |
| Critical | SLA breach or balance < $5 | Page on-call, auto-pause agent, escalate |
| Emergency | > 20% of fleet critical | Page fleet commander, freeze provisioning |
def evaluate_alerts(fleet_metrics, thresholds):
"""Evaluate fleet metrics against alerting thresholds."""
alerts = []
critical_pct = fleet_metrics["critical"] / fleet_metrics["total"] * 100
if critical_pct > thresholds.get("emergency_critical_pct", 20):
alerts.append({
"severity": "emergency",
"message": f"{critical_pct:.1f}% of fleet in critical state",
"action": "freeze_provisioning"
})
if fleet_metrics["sla_compliance_pct"] < thresholds.get("min_sla_pct", 95):
alerts.append({
"severity": "critical",
"message": f"Fleet SLA compliance at {fleet_metrics['sla_compliance_pct']:.1f}%",
"action": "page_fleet_commander"
})
if fleet_metrics["total_burn_rate"] > thresholds.get("max_burn_rate_per_hour", 100):
alerts.append({
"severity": "warning",
"message": f"Fleet burn rate ${fleet_metrics['total_burn_rate']:.2f}/hr exceeds budget",
"action": "notify_finops"
})
return alerts
Behavioral Observability: The Fifth Pillar
Traditional metrics tell you the agent is slow. Behavioral observability tells you why. An agent that suddenly starts calling search_services 40 times per task instead of 3 might have a perfectly healthy error rate and latency — but it is thrashing through the marketplace because its ranking heuristic has drifted. Without behavioral signals, you would never catch this until the cost spike hits.
Track these behavioral metrics per agent:
| Metric | What It Reveals | Alert Threshold |
|---|---|---|
| Tool calls per task | Efficiency / thrashing | > 2x baseline |
| Unique tools per task | Task complexity drift | > 3x baseline |
| Retry ratio | Downstream health | > 20% of calls are retries |
| Escalation rate | Agent capability limits | > 15% of tasks escalated |
| Decision latency | Reasoning bottlenecks | P99 > 5x P50 |
| Counterparty diversity | Market exploration | < 2 unique counterparties/day |
def collect_behavioral_metrics(agent_id, task_log):
"""Extract behavioral signals from a task execution log."""
metrics = {
"tool_calls_per_task": len(task_log.get("tool_calls", [])),
"unique_tools": len(set(tc["tool"] for tc in task_log.get("tool_calls", []))),
"retries": sum(1 for tc in task_log.get("tool_calls", []) if tc.get("is_retry")),
"escalated": task_log.get("escalated", False),
"decision_latency_ms": task_log.get("decision_latency_ms", 0)
}
# Submit behavioral metrics alongside performance metrics
resp = session.post(f"{base_url}/v1", json={
"tool": "submit_metrics",
"input": {
"agent_id": agent_id,
"metrics": metrics
}
})
return resp.json()
Distributed Tracing for Multi-Agent Workflows
When Agent A delegates a subtask to Agent B, which calls Agent C for data enrichment, a single task spans three agents and potentially dozens of tool calls. Without distributed tracing, debugging a failed task means manually correlating logs across agents — the same problem microservices solved with Jaeger and Zipkin.
Implement trace propagation by passing a trace context through every tool call:
import uuid
def create_trace_context(parent_trace_id=None):
"""Create or extend a trace context for cross-agent tracing."""
trace_id = parent_trace_id or str(uuid.uuid4())
span_id = str(uuid.uuid4())[:16]
return {"trace_id": trace_id, "span_id": span_id}
def traced_tool_call(agent_id, tool_name, tool_input, trace_ctx):
"""Execute a tool call with trace context propagation."""
resp = session.post(f"{base_url}/v1", json={
"tool": tool_name,
"input": {
**tool_input,
"_trace": trace_ctx
}
})
result = resp.json()
# Record the span in the audit trail for correlation
session.post(f"{base_url}/v1", json={
"tool": "create_audit_trail",
"input": {
"agent_id": agent_id,
"action": f"tool_call:{tool_name}",
"details": {
"trace_id": trace_ctx["trace_id"],
"span_id": trace_ctx["span_id"],
"status_code": resp.status_code,
"latency_ms": resp.elapsed.microseconds // 1000
}
}
})
return result
Observability Checklist
- Per-agent metrics collection runs every 60 seconds
- Fleet aggregation computes health classification (healthy/degraded/critical)
- Alerting thresholds configured for latency, error rate, SLA, and cost
- Escalation pipeline tested end-to-end (can you verify the page fires?)
- Dashboard shows real-time fleet health with drill-down to individual agents
- Metrics retained for 90 days minimum (compliance requirement)
- Anomaly detection enabled for cost spikes (>200% of hourly baseline)
- Tool-call-level tracing enabled for debugging cascading failures
- Behavioral metrics tracked per agent (tool calls per task, retry ratio, escalation rate)
- Distributed trace context propagated across multi-agent workflows
Chapter 4: Cost-Aware Model Routing and FinOps
The Agent FinOps Problem
A microservice has a predictable cost profile: X requests per second times Y compute per request. An agent has a stochastic cost profile: it might resolve a task in 2 tool calls or 47, it might choose a cheap model or an expensive one, and it might retry failed calls indefinitely if you do not stop it. The $400M Fortune 500 cloud leak happened because nobody modeled the cost distribution of agent behavior under production load.
Token-Level Cost Tracking
import requests
base_url = "https://api.greenhelix.net/v1"
api_key = "your-api-key"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {api_key}"
def track_task_cost(agent_id, task_id, tool_calls):
"""Track per-task cost across all tool calls."""
total_cost = 0.0
for call in tool_calls:
# Estimate cost before executing
resp = session.post(f"{base_url}/v1", json={
"tool": "estimate_cost",
"input": {
"tool_name": call["tool"],
"input_size": call.get("estimated_tokens", 100)
}
})
estimate = resp.json()
estimated_cost = float(estimate.get("estimated_cost", "0"))
# Check budget before executing
if total_cost + estimated_cost > call.get("budget_remaining", float("inf")):
return {
"status": "budget_exceeded",
"total_cost": str(total_cost),
"blocked_call": call["tool"]
}
# Execute the tool call
resp = session.post(f"{base_url}/v1", json={
"tool": call["tool"],
"input": call["input"]
})
# Record the transaction
resp = session.post(f"{base_url}/v1", json={
"tool": "record_transaction",
"input": {
"agent_id": agent_id,
"task_id": task_id,
"tool": call["tool"],
"cost": str(estimated_cost),
"timestamp": call.get("timestamp")
}
})
total_cost += estimated_cost
return {"status": "completed", "total_cost": str(total_cost)}
The Model Routing Decision Tree
Not every task needs GPT-4 or Claude Opus. Most fleets waste 60-70% of their token budget routing simple classification tasks to frontier models.
┌──────────────────────┐
│ Incoming Task │
└──────────┬───────────┘
│
┌──────────▼───────────┐
│ Complexity estimate? │
└──────────┬───────────┘
┌───────┴───────┐
▼ ▼
┌─────────┐ ┌──────────┐
│ LOW │ │ MED/HIGH │
│(<100tok)│ │ │
└────┬────┘ └────┬─────┘
│ │
▼ ▼
┌──────────────┐ ┌─────────────────┐
│ Small model │ │ Budget remaining │
│ (Haiku/mini) │ │ > threshold? │
│ $0.001/task │ └────────┬────────┘
└──────────────┘ ┌─────┴─────┐
▼ ▼
┌──────────┐ ┌──────────┐
│ YES │ │ NO │
│ Frontier │ │ Mid-tier │
│ model │ │ model │
│$0.03/task│ │$0.008/tsk│
└──────────┘ └──────────┘
Budget Guardrails
class BudgetGuardrail:
"""Enforce per-agent, per-fleet, and per-task budget limits."""
def __init__(self, session, base_url):
self.session = session
self.base_url = base_url
self.limits = {
"per_task_max": 5.00,
"per_agent_hourly_max": 50.00,
"fleet_daily_max": 2000.00
}
def check_balance(self, agent_id):
"""Check if agent has sufficient funds to continue."""
resp = self.session.post(f"{self.base_url}/v1", json={
"tool": "get_balance",
"input": {"agent_id": agent_id}
})
balance = resp.json()
current = float(balance.get("balance", "0"))
return current > self.limits["per_task_max"]
def get_volume_discount(self, agent_id):
"""Check if agent qualifies for volume pricing."""
resp = self.session.post(f"{self.base_url}/v1", json={
"tool": "get_volume_discount",
"input": {"agent_id": agent_id}
})
return resp.json()
def convert_cost_estimate(self, amount, from_currency, to_currency):
"""Convert cost estimates across currencies for global fleets."""
resp = self.session.post(f"{self.base_url}/v1", json={
"tool": "convert_currency",
"input": {
"amount": str(amount),
"from_currency": from_currency,
"to_currency": to_currency
}
})
return resp.json()
def enforce(self, agent_id, task_cost):
"""Enforce budget guardrails. Returns (allowed, reason)."""
if task_cost > self.limits["per_task_max"]:
return False, f"Task cost ${task_cost:.2f} exceeds per-task limit ${self.limits['per_task_max']:.2f}"
if not self.check_balance(agent_id):
return False, f"Agent {agent_id} balance below minimum threshold"
return True, "within_budget"
Fleet Economics Model
Build a spreadsheet model before deploying. Here is the math:
Fleet Monthly Cost = Σ(agent_i) [
(tasks_per_day × avg_tool_calls_per_task × avg_cost_per_call × 30)
+ (provisioning_cost)
+ (monitoring_overhead)
]
ROI = (revenue_from_agent_tasks - fleet_monthly_cost) / fleet_monthly_cost
Break-even point = fleet_monthly_cost / revenue_per_task
Example: A 50-agent triage fleet handling 10,000 tasks/day at $0.003/task average cost:
- Monthly tool cost: 10,000 x 3.2 avg calls x $0.003 x 30 = $2,880
- Monitoring overhead (5%): $144
- Total: $3,024/month
- If each task saves $0.50 in human labor: Revenue = 10,000 x $0.50 x 30 = $150,000/month
- ROI: 4,860%
Cost Attribution: Who Pays for Shared Work?
In multi-agent workflows, a single task may involve several agents. The triage agent classifies it, the specialist agent processes it, the quality agent validates it. Who pays for each step? Without clear cost attribution, fleet economics modeling is fiction.
Three attribution models:
Caller-pays: The agent that initiates a tool call or delegates a subtask pays the full cost. Simple to implement. Penalizes orchestrator agents that coordinate work but add the least value per call.
Proportional split: Each agent in the workflow pays a share proportional to the token volume or tool calls they contributed. Fair but complex to implement — requires full trace context to compute shares.
Task-budget pool: A budget is allocated to the task, not the agent. All agents draw from the same pool. The fleet operator monitors the pool, not individual agents. Best for workflows where agent boundaries are fluid.
def attribute_task_cost(task_trace, model="proportional"):
"""Attribute costs across agents in a multi-agent workflow."""
agent_costs = {}
for span in task_trace.get("spans", []):
agent_id = span["agent_id"]
cost = float(span.get("cost", "0"))
if agent_id not in agent_costs:
agent_costs[agent_id] = 0.0
agent_costs[agent_id] += cost
total = sum(agent_costs.values())
if model == "caller_pays":
initiator = task_trace["initiating_agent"]
return {initiator: str(total)}
elif model == "proportional":
return {
agent_id: str(cost)
for agent_id, cost in agent_costs.items()
}
elif model == "task_pool":
return {"task_pool": str(total), "breakdown": agent_costs}
return agent_costs
The Weekly FinOps Review
Schedule a weekly review with this agenda:
- Budget vs. actual — Compare each fleet segment's actual spend to its budget. Flag any segment exceeding 110% of plan.
- Cost per task trend — Is cost per task stable, rising, or falling? Rising cost at stable volume means agent efficiency is degrading.
- Model mix audit — What percentage of tasks went to frontier models vs. mid-tier vs. small? Is the routing decision tree working?
- Top 10 most expensive tasks — Review the outliers. Were they legitimately complex or did an agent thrash?
- Volume discount capture — Are you hitting volume tiers? If not, should you consolidate traffic to fewer agents to hit the next tier?
FinOps Checklist
- Per-task budget limits configured and enforced
- Per-agent hourly spending caps in place
- Fleet daily budget ceiling with automated pause on breach
- Cost estimates run before every tool call (estimate_cost)
- Volume discount eligibility checked at provisioning
- Model routing routes <100 token tasks to cheapest model
- Weekly cost review comparing actual vs. budgeted spend
- Anomaly detection for cost spikes (agent spending >3x normal)
- Currency conversion configured for global fleet operations
- Break-even analysis documented per fleet segment
- Cost attribution model chosen and implemented for multi-agent workflows
- Top-10 expensive task review runs weekly
Chapter 5: SLA Enforcement and Automated Escalation
Defining Agent-Level SLAs
An SLA for an agent is not the same as an SLA for a web service. A web service SLA covers availability and latency. An agent SLA must also cover accuracy, cost-per-task, and behavioral bounds — things that are harder to measure and harder to enforce.
SLA dimensions for agents:
| Dimension | Metric | Example Target |
|---|---|---|
| Response time | P99 latency | < 2,000 ms |
| Availability | Uptime percentage | > 99.5% |
| Accuracy | Task success rate | > 95% |
| Cost | Average cost per task | < $0.05 |
| Behavioral | Escalation rate | < 10% of tasks |
Creating SLAs Programmatically
import requests
base_url = "https://api.greenhelix.net/v1"
api_key = "your-api-key"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {api_key}"
def create_agent_sla(agent_id, sla_template):
"""Create an SLA contract for an agent."""
sla_definitions = {
"standard-support": {
"response_time_p99_ms": 2000,
"uptime_pct": 99.5,
"accuracy_pct": 95.0,
"max_cost_per_task": "0.05",
"review_period_days": 30
},
"premium-support": {
"response_time_p99_ms": 500,
"uptime_pct": 99.9,
"accuracy_pct": 98.0,
"max_cost_per_task": "0.10",
"review_period_days": 7
}
}
sla_config = sla_definitions.get(sla_template)
if not sla_config:
raise ValueError(f"Unknown SLA template: {sla_template}")
resp = session.post(f"{base_url}/v1", json={
"tool": "create_sla",
"input": {
"agent_id": agent_id,
"terms": sla_config
}
})
return resp.json()
Continuous Compliance Monitoring
def monitor_sla_compliance(agent_ids, interval_minutes=5):
"""Continuously monitor SLA compliance across fleet."""
violations = []
for agent_id in agent_ids:
resp =
…(truncated)