Agent Observability Stack: Distributed Tracing, Metrics, and Alerting for Multi-Agent Systems
Notice: This is an educational guide with illustrative code examples. It does not execute code, require credentials, or install dependencies. All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required to get started.
When Agent A calls Agent B which calls Agent C, and the transaction takes 12 seconds instead of 200ms, where is the bottleneck? When your agent fleet processes 10,000 transactions per day and revenue drops 15%, which agent is underperforming? When an escrow settlement fails silently at 3am, how quickly do you find out? Traditional monitoring tools -- ping checks, CPU graphs, uptime dashboards -- cannot answer these questions for distributed agent systems. They were built for monoliths and simple request-response services. Agent commerce is fundamentally different: transactions span multiple autonomous agents, each with its own state, pricing, and failure modes. A single customer-facing operation might traverse five agents, two escrow contracts, and three separate billing events before completing. The failure surface is combinatorial, not linear. You need observability, not just monitoring. Monitoring tells you something is broken. Observability tells you why, where, and how to fix it. For agent commerce, this means distributed tracing to follow transactions across agent boundaries, custom metrics to measure business outcomes (not just infrastructure health), and intelligent alerting to catch problems before your users do. The difference between a 15-minute outage and a 4-hour revenue leak is whether your observability stack understands agent-to-agent transaction flows. This guide builds a complete observability stack from scratch. We use OpenTelemetry standards for interoperability, integrate directly with GreenHelix's metrics and event tools for agent-specific telemetry, and build anomaly detection that understands the patterns unique to agent commerce. Every component is production Python code you can deploy today. By the end, you will have distributed tracing across agent calls, custom business metrics, anomaly detection, dashboards, alerting with escalation policies, and SLA monitoring with cost attribution.
What You'll Learn
- Chapter 1: The Three Pillars of Agent Observability
- Chapter 2: AgentTracer Class
- Chapter 3: Distributed Tracing Across Agent Calls
- Chapter 4: MetricsCollector Class
- Chapter 5: Anomaly Detection
- Chapter 6: Dashboard Design
- Chapter 7: AlertManager Class
- Chapter 8: SLA Monitoring and Cost Attribution
- What's Next
Full Guide
Agent Observability Stack: Distributed Tracing, Metrics, and Alerting for Multi-Agent Systems
When Agent A calls Agent B which calls Agent C, and the transaction takes 12 seconds instead of 200ms, where is the bottleneck? When your agent fleet processes 10,000 transactions per day and revenue drops 15%, which agent is underperforming? When an escrow settlement fails silently at 3am, how quickly do you find out? Traditional monitoring tools -- ping checks, CPU graphs, uptime dashboards -- cannot answer these questions for distributed agent systems. They were built for monoliths and simple request-response services. Agent commerce is fundamentally different: transactions span multiple autonomous agents, each with its own state, pricing, and failure modes. A single customer-facing operation might traverse five agents, two escrow contracts, and three separate billing events before completing. The failure surface is combinatorial, not linear.
You need observability, not just monitoring. Monitoring tells you something is broken. Observability tells you why, where, and how to fix it. For agent commerce, this means distributed tracing to follow transactions across agent boundaries, custom metrics to measure business outcomes (not just infrastructure health), and intelligent alerting to catch problems before your users do. The difference between a 15-minute outage and a 4-hour revenue leak is whether your observability stack understands agent-to-agent transaction flows.
This guide builds a complete observability stack from scratch. We use OpenTelemetry standards for interoperability, integrate directly with GreenHelix's metrics and event tools for agent-specific telemetry, and build anomaly detection that understands the patterns unique to agent commerce. Every component is production Python code you can deploy today. By the end, you will have distributed tracing across agent calls, custom business metrics, anomaly detection, dashboards, alerting with escalation policies, and SLA monitoring with cost attribution.
Getting started: All examples in this guide work with the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required.
Chapter 1: The Three Pillars of Agent Observability
Observability for distributed systems rests on three pillars: traces, metrics, and logs. Each answers a different class of question, and none is sufficient alone. For agent commerce, each pillar has specific requirements that differ from traditional web service observability.
Traces: Following a Transaction Across Agent Boundaries
A trace is the complete record of a single transaction as it flows through your system. In agent commerce, a trace might begin when a customer agent requests a service, pass through a gateway agent, invoke a specialist agent, trigger an escrow creation, wait for fulfillment, and end with settlement and payment distribution. Each step in that journey is a span, and the collection of spans forms a trace.
The critical difference from traditional distributed tracing is that agent boundaries are not just service boundaries -- they are trust boundaries. When Agent A calls Agent B, there is an economic transaction embedded in that call. The trace must capture not just latency and status codes, but also billing events, escrow state transitions, and the economic context of each span. A span that shows "200 OK in 150ms" is incomplete if it does not also show "billed 0.003 credits, escrow ID esw_abc123 created."
Traces answer questions like: Why did this specific transaction fail? Where did the latency come from? Which agent in the chain was the bottleneck? Did the billing event match the actual work performed?
Metrics: Measuring Agent Health, Performance, and Business Outcomes
Metrics are aggregated numerical measurements over time. Where traces give you detail about individual transactions, metrics give you the big picture. For agent commerce, metrics fall into three categories.
Infrastructure metrics measure the health of your agents as software systems: CPU usage, memory consumption, request queue depth, connection pool utilization. These are table stakes and most monitoring tools handle them well.
Performance metrics measure how your agents behave under load: request latency (p50, p95, p99), throughput (requests per second), error rates, and timeout rates. These require histograms and percentile calculations, not just averages.
Business metrics are where agent commerce diverges from traditional services. You need to track revenue per agent, cost per transaction, escrow settlement rates, dispute rates, SLA compliance percentages, and customer satisfaction proxies. These metrics directly measure whether your agent fleet is making money or losing it.
GreenHelix's submit_metrics tool accepts custom metric submissions, making it the natural sink for all three categories. The tool accepts metric name, value, dimensions, and timestamp, allowing you to build rich, queryable metric series.
Logs: Structured Event Logging for Debugging
Logs are discrete events with context. In agent commerce, structured logs are essential because you need to correlate log events across agent boundaries. A plain text log line like "Error processing request" is useless when you have 50 agents each producing thousands of log lines per hour.
Structured logs include the trace ID, span ID, agent ID, transaction ID, and any relevant business context as machine-parseable fields. When something goes wrong, you filter logs by trace ID and see every event from every agent involved in that transaction, in chronological order.
GreenHelix's get_events tool retrieves event streams that serve as a structured log source. Events include agent interactions, billing events, escrow state changes, and webhook deliveries. By correlating your application logs with GreenHelix events, you get a complete picture of what happened and why.
Why All Three Matter
Metrics tell you something is wrong: "Error rate spiked to 5% at 14:32." Traces tell you where and why: "Transaction txn_789 failed at Agent C because the escrow creation timed out after 30 seconds." Logs give you the details: "Agent C's connection pool was exhausted because Agent D was not releasing connections after failed settlements."
Without metrics, you do not know there is a problem until users complain. Without traces, you cannot pinpoint which agent or which step is responsible. Without logs, you cannot understand the root cause well enough to fix it. Agent commerce multiplies this dependency because the interactions are more complex, the failure modes are more subtle, and the economic consequences of missed problems are direct and measurable.
GreenHelix Tools for Observability
Three GreenHelix tools form the foundation of our observability stack:
submit_metrics accepts custom metric data points with dimensions. We use this to push agent performance metrics, business metrics, and custom counters into GreenHelix's metric store.
get_events retrieves event streams filtered by agent, time range, and event type. We use this to pull structured event data for trace correlation and log enrichment.
register_webhook creates webhook subscriptions for specific event types. We use this to trigger real-time alerts when critical events occur, rather than polling for problems.
from greenhelix import GreenHelixClient
client = GreenHelixClient(api_key="your-api-key")
# Submit a custom metric
client.execute_tool("submit_metrics", {
"agent_id": "agent-payment-processor",
"metrics": [
{
"name": "transaction.latency_ms",
"value": 245.3,
"dimensions": {
"agent": "agent-payment-processor",
"operation": "process_payment",
"status": "success"
},
"timestamp": "2026-04-07T14:30:00Z"
}
]
})
# Retrieve events for trace correlation
events = client.execute_tool("get_events", {
"agent_id": "agent-payment-processor",
"event_types": ["escrow.created", "escrow.settled", "billing.charged"],
"start_time": "2026-04-07T14:00:00Z",
"end_time": "2026-04-07T15:00:00Z"
})
# Register a webhook for real-time alerting
client.execute_tool("register_webhook", {
"url": "https://alerts.myfleet.com/webhook",
"event_types": ["escrow.failed", "billing.error"],
"agent_id": "agent-payment-processor"
})
These three tools, combined with OpenTelemetry's tracing and metrics APIs, give us everything we need to build a production observability stack.
Chapter 2: AgentTracer Class
The core of our observability stack is a tracer that understands agent commerce. Standard OpenTelemetry tracers create spans for HTTP requests and database calls. Our AgentTracer wraps OpenTelemetry to add agent-specific context: billing events, escrow state, economic metadata, and cross-agent trace propagation.
Design Principles
The tracer must be lightweight. Adding observability should not measurably increase transaction latency. We target less than 1ms overhead per span, which means in-memory buffering with asynchronous export.
The tracer must be OpenTelemetry-compatible. This means traces can be exported to Jaeger, Zipkin, Grafana Tempo, or any other OTel-compatible backend. We do not lock you into a proprietary format.
The tracer must understand agent commerce primitives. Spans should automatically capture billing amounts, escrow IDs, agent identities, and transaction types without requiring manual instrumentation at every call site.
The AgentTracer Implementation
import time
import uuid
import threading
from dataclasses import dataclass, field
from typing import Optional, Dict, Any, List
from contextlib import contextmanager
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import (
BatchSpanProcessor,
SpanExporter,
SpanExportResult,
)
from opentelemetry.trace import StatusCode, SpanKind
from opentelemetry.context import attach, detach, set_value, get_value
@dataclass
class AgentSpanAttributes:
"""Standard attributes for agent commerce spans."""
AGENT_ID = "agent.id"
AGENT_ROLE = "agent.role"
TRANSACTION_ID = "agent.transaction.id"
ESCROW_ID = "agent.escrow.id"
BILLING_AMOUNT = "agent.billing.amount"
BILLING_CURRENCY = "agent.billing.currency"
OPERATION_TYPE = "agent.operation.type"
PEER_AGENT_ID = "agent.peer.id"
SLA_TIER = "agent.sla.tier"
COST_CENTER = "agent.cost_center"
class GreenHelixSpanExporter(SpanExporter):
"""Exports spans as metrics to GreenHelix submit_metrics."""
def __init__(self, client, agent_id: str, batch_size: int = 50):
self._client = client
self._agent_id = agent_id
self._batch_size = batch_size
def export(self, spans) -> SpanExportResult:
metrics_batch = []
for span in spans:
duration_ms = (span.end_time - span.start_time) / 1_000_000
attrs = dict(span.attributes) if span.attributes else {}
metrics_batch.append({
"name": "agent.span.duration_ms",
"value": duration_ms,
"dimensions": {
"agent": self._agent_id,
"operation": span.name,
"status": "ok" if span.status.status_code == StatusCode.OK
else "error",
"span_kind": span.kind.name if span.kind else "INTERNAL",
"peer_agent": attrs.get(AgentSpanAttributes.PEER_AGENT_ID, ""),
},
"timestamp": _ns_to_iso(span.end_time),
})
if AgentSpanAttributes.BILLING_AMOUNT in attrs:
metrics_batch.append({
"name": "agent.billing.amount",
"value": float(attrs[AgentSpanAttributes.BILLING_AMOUNT]),
"dimensions": {
"agent": self._agent_id,
"operation": span.name,
"currency": attrs.get(
AgentSpanAttributes.BILLING_CURRENCY, "credits"
),
},
"timestamp": _ns_to_iso(span.end_time),
})
if len(metrics_batch) >= self._batch_size:
self._flush(metrics_batch)
metrics_batch = []
if metrics_batch:
self._flush(metrics_batch)
return SpanExportResult.SUCCESS
def _flush(self, metrics):
try:
self._client.execute_tool("submit_metrics", {
"agent_id": self._agent_id,
"metrics": metrics,
})
except Exception:
return SpanExportResult.FAILURE
def shutdown(self):
pass
def _ns_to_iso(ns_timestamp: int) -> str:
"""Convert nanosecond timestamp to ISO 8601."""
import datetime
dt = datetime.datetime.fromtimestamp(
ns_timestamp / 1e9, tz=datetime.timezone.utc
)
return dt.isoformat()
class AgentTracer:
"""OpenTelemetry-compatible tracer for agent commerce."""
def __init__(
self,
agent_id: str,
client=None,
service_name: str = None,
exporters: List[SpanExporter] = None,
):
self.agent_id = agent_id
self._client = client
self._service_name = service_name or f"agent-{agent_id}"
provider = TracerProvider()
if client:
gh_exporter = GreenHelixSpanExporter(client, agent_id)
provider.add_span_processor(BatchSpanProcessor(gh_exporter))
if exporters:
for exporter in exporters:
provider.add_span_processor(BatchSpanProcessor(exporter))
self._provider = provider
self._tracer = provider.get_tracer(
self._service_name, schema_url="https://greenhelix.net/schemas/agent/1.0"
)
@contextmanager
def start_span(
self,
name: str,
kind: SpanKind = SpanKind.INTERNAL,
attributes: Dict[str, Any] = None,
peer_agent_id: str = None,
):
"""Start a new span with agent commerce context."""
attrs = {
AgentSpanAttributes.AGENT_ID: self.agent_id,
}
if peer_agent_id:
attrs[AgentSpanAttributes.PEER_AGENT_ID] = peer_agent_id
if attributes:
attrs.update(attributes)
with self._tracer.start_as_current_span(
name, kind=kind, attributes=attrs
) as span:
yield span
@contextmanager
def trace_agent_call(
self,
target_agent_id: str,
operation: str,
transaction_id: str = None,
):
"""Trace a call from this agent to another agent."""
txn_id = transaction_id or str(uuid.uuid4())
attrs = {
AgentSpanAttributes.OPERATION_TYPE: operation,
AgentSpanAttributes.TRANSACTION_ID: txn_id,
}
with self.start_span(
f"call.{target_agent_id}.{operation}",
kind=SpanKind.CLIENT,
attributes=attrs,
peer_agent_id=target_agent_id,
) as span:
yield span
@contextmanager
def trace_escrow(self, escrow_id: str, operation: str):
"""Trace an escrow lifecycle operation."""
attrs = {
AgentSpanAttributes.ESCROW_ID: escrow_id,
AgentSpanAttributes.OPERATION_TYPE: f"escrow.{operation}",
}
with self.start_span(
f"escrow.{operation}",
kind=SpanKind.INTERNAL,
attributes=attrs,
) as span:
yield span
@contextmanager
def trace_billing(self, amount: float, currency: str = "credits"):
"""Trace a billing event."""
attrs = {
AgentSpanAttributes.BILLING_AMOUNT: amount,
AgentSpanAttributes.BILLING_CURRENCY: currency,
AgentSpanAttributes.OPERATION_TYPE: "billing.charge",
}
with self.start_span(
"billing.charge",
kind=SpanKind.INTERNAL,
attributes=attrs,
) as span:
yield span
def get_trace_context(self) -> Dict[str, str]:
"""Extract current trace context for propagation to other agents."""
span = trace.get_current_span()
ctx = span.get_span_context()
if not ctx or not ctx.is_valid:
return {}
return {
"traceparent": f"00-{format(ctx.trace_id, '032x')}-"
f"{format(ctx.span_id, '016x')}-"
f"{'01' if ctx.trace_flags & 1 else '00'}",
"x-agent-id": self.agent_id,
}
def inject_context(self, headers: Dict[str, str]) -> Dict[str, str]:
"""Inject trace context into outgoing request headers."""
headers.update(self.get_trace_context())
return headers
def shutdown(self):
"""Flush pending spans and shut down the tracer."""
self._provider.shutdown()
Using the AgentTracer
The tracer integrates naturally with GreenHelix client calls. Each API interaction gets a span, and the spans automatically capture timing, status, and agent commerce attributes.
from greenhelix import GreenHelixClient
client = GreenHelixClient(api_key="your-api-key")
tracer = AgentTracer(agent_id="agent-marketplace", client=client)
# Trace a multi-step transaction
with tracer.trace_agent_call("agent-data-provider", "fetch_dataset") as span:
result = client.execute_tool("call_agent", {
"target": "agent-data-provider",
"operation": "fetch_dataset",
"params": {"dataset_id": "ds_12345"},
"headers": tracer.get_trace_context(),
})
span.set_attribute("dataset.size_bytes", result.get("size", 0))
if result.get("status") == "error":
span.set_status(StatusCode.ERROR, result.get("message", ""))
else:
span.set_status(StatusCode.OK)
# Trace an escrow lifecycle
with tracer.trace_escrow("esw_abc123", "create") as span:
escrow = client.execute_tool("create_escrow", {
"amount": 5.00,
"payer": "agent-buyer",
"payee": "agent-seller",
})
span.set_attribute("escrow.amount", 5.00)
span.set_status(StatusCode.OK)
The key design choice is that AgentTracer wraps OpenTelemetry rather than replacing it. You can add any standard OTel exporter (Jaeger, OTLP, console) alongside the GreenHelix exporter. This means your agent traces show up in your existing observability infrastructure without any migration.
The GreenHelixSpanExporter converts spans to metrics via submit_metrics. This dual-purpose export means every traced operation automatically generates latency and billing metrics, eliminating the need to instrument metrics separately for operations you are already tracing.
Chapter 3: Distributed Tracing Across Agent Calls
Single-agent tracing is straightforward. The real challenge is tracing a transaction as it flows through multiple independent agents, each potentially running on different infrastructure, operated by different teams, and communicating through the GreenHelix gateway.
The Problem: Trace Context Propagation
When Agent A calls Agent B, Agent B has no inherent knowledge of Agent A's trace. Without explicit context propagation, Agent B starts a new, disconnected trace. You end up with five isolated traces instead of one unified trace showing the complete transaction flow.
The W3C Trace Context standard solves this with two HTTP headers: traceparent (containing trace ID, span ID, and sampling flags) and tracestate (containing vendor-specific data). Our AgentTracer already generates these headers via get_trace_context(). The challenge is ensuring every agent in the chain extracts, uses, and propagates these headers.
Trace Context in Agent-to-Agent Calls
Here is the pattern for traced agent-to-agent calls. The calling agent injects context into outgoing headers. The receiving agent extracts context and creates child spans.
# === Calling Agent (Agent A) ===
class TracedAgentClient:
"""HTTP client that automatically propagates trace context."""
def __init__(self, tracer: AgentTracer, client):
self._tracer = tracer
self._client = client
def call_agent(
self,
target_agent: str,
operation: str,
params: Dict[str, Any],
transaction_id: str = None,
) -> Dict[str, Any]:
"""Call another agent with automatic trace propagation."""
with self._tracer.trace_agent_call(
target_agent, operation, transaction_id
) as span:
headers = self._tracer.get_trace_context()
headers["x-transaction-id"] = transaction_id or ""
result = self._client.execute_tool("call_agent", {
"target": target_agent,
"operation": operation,
"params": params,
"headers": headers,
})
span.set_attribute("response.status", result.get("status", "unknown"))
if result.get("error"):
span.set_status(StatusCode.ERROR, result["error"])
else:
span.set_status(StatusCode.OK)
return result
# === Receiving Agent (Agent B) ===
from opentelemetry.trace.propagation import TraceContextTextMapPropagator
from opentelemetry.context import Context
class TracedAgentHandler:
"""Request handler that extracts and continues trace context."""
def __init__(self, tracer: AgentTracer):
self._tracer = tracer
self._propagator = TraceContextTextMapPropagator()
def handle_request(
self,
operation: str,
params: Dict[str, Any],
headers: Dict[str, str],
) -> Dict[str, Any]:
"""Handle an incoming agent request, continuing the trace."""
# Extract trace context from incoming headers
ctx = self._propagator.extract(carrier=headers)
# Create a child span under the extracted context
token = attach(ctx)
try:
with self._tracer.start_span(
f"handle.{operation}",
kind=SpanKind.SERVER,
attributes={
"caller.agent_id": headers.get("x-agent-id", "unknown"),
AgentSpanAttributes.OPERATION_TYPE: operation,
},
) as span:
result = self._dispatch(operation, params, span)
return result
finally:
detach(token)
def _dispatch(
self, operation: str, params: Dict[str, Any], span
) -> Dict[str, Any]:
"""Route to the appropriate operation handler."""
handler = getattr(self, f"op_{operation}", None)
if not handler:
span.set_status(StatusCode.ERROR, f"Unknown operation: {operation}")
return {"error": f"Unknown operation: {operation}"}
return handler(params, span)
Tracing a Complete Transaction Through Five Agents
Consider an agent commerce transaction where a customer agent orders a data analysis report. The transaction flows through five agents:
- Customer Agent -- initiates the request
- Marketplace Agent -- matches request to provider
- Data Provider Agent -- supplies raw data
- Analysis Agent -- processes and analyzes data
- Delivery Agent -- formats and delivers the report
def traced_execute(
client: TracedAgentClient,
tracer: AgentTracer,
request: Dict[str, Any],
) -> Dict[str, Any]:
"""Execute a full traced transaction across multiple agents."""
transaction_id = str(uuid.uuid4())
with tracer.start_span(
"transaction.data_analysis_report",
kind=SpanKind.INTERNAL,
attributes={
AgentSpanAttributes.TRANSACTION_ID: transaction_id,
AgentSpanAttributes.OPERATION_TYPE: "data_analysis_report",
},
) as root_span:
# Step 1: Find a provider through the marketplace
with tracer.start_span("step.find_provider") as step_span:
match = client.call_agent(
"agent-marketplace",
"find_provider",
{"capability": "data_analysis", "budget": request["budget"]},
transaction_id=transaction_id,
)
step_span.set_attribute("provider.matched", match.get("provider_id", ""))
provider_id = match["provider_id"]
# Step 2: Create escrow for the transaction
with tracer.trace_escrow(f"esw_{transaction_id[:8]}", "create") as esc_span:
escrow = client.call_agent(
"agent-marketplace",
"create_escrow",
{
"amount": match["price"],
"payer": request["customer_agent"],
"payee": provider_id,
},
transaction_id=transaction_id,
)
esc_span.set_attribute("escrow.id", escrow.get("escrow_id", ""))
# Step 3: Fetch raw data
with tracer.start_span("step.fetch_data") as step_span:
data = client.call_agent(
match["data_source"],
"fetch_dataset",
{"dataset_id": request["dataset_id"]},
transaction_id=transaction_id,
)
step_span.set_attribute("data.records", data.get("record_count", 0))
# Step 4: Run analysis
with tracer.start_span("step.analyze") as step_span:
analysis = client.call_agent(
provider_id,
"analyze_data",
{"data": data["data"], "analysis_type": request["analysis_type"]},
transaction_id=transaction_id,
)
step_span.set_attribute("analysis.confidence", analysis.get("confidence", 0))
# Step 5: Deliver report
with tracer.start_span("step.deliver") as step_span:
delivery = client.call_agent(
"agent-delivery",
"deliver_report",
{
"report": analysis["report"],
"recipient": request["customer_agent"],
"format": request.get("format", "pdf"),
},
transaction_id=transaction_id,
)
step_span.set_attribute("delivery.channel", delivery.get("channel", ""))
# Step 6: Settle escrow
with tracer.trace_escrow(escrow["escrow_id"], "settle") as esc_span:
settlement = client.call_agent(
"agent-marketplace",
"settle_escrow",
{"escrow_id": escrow["escrow_id"]},
transaction_id=transaction_id,
)
esc_span.set_attribute("settlement.status", settlement.get("status", ""))
root_span.set_status(StatusCode.OK)
return {
"transaction_id": transaction_id,
"report_url": delivery.get("url"),
"total_cost": match["price"],
}
Trace Correlation with GreenHelix Events
GreenHelix events have their own event IDs. To get a complete picture, you need to correlate your OpenTelemetry trace IDs with GreenHelix event IDs. The approach is to embed the trace ID in your GreenHelix API calls and then join on it during analysis.
def trace_escrow_lifecycle(
tracer: AgentTracer,
client,
escrow_id: str,
) -> Dict[str, Any]:
"""Trace and correlate an escrow's full lifecycle with GreenHelix events."""
trace_ctx = tracer.get_trace_context()
trace_id = trace_ctx.get("traceparent", "").split("-")[1] if trace_ctx else ""
with tracer.trace_escrow(escrow_id, "lifecycle") as span:
# Fetch GreenHelix events for this escrow
events = client.execute_tool("get_events", {
"filters": {"escrow_id": escrow_id},
"start_time": "2026-04-07T00:00:00Z",
"end_time": "2026-04-07T23:59:59Z",
})
lifecycle = {
"escrow_id": escrow_id,
"trace_id": trace_id,
"events": [],
}
for event in events.get("events", []):
event_type = event.get("type", "")
lifecycle["events"].append({
"type": event_type,
"timestamp": event.get("timestamp"),
"greenhelix_event_id": event.get("event_id"),
"trace_id": trace_id,
})
# Create a child span for each GreenHelix event
with tracer.start_span(
f"greenhelix.event.{event_type}",
attributes={
"greenhelix.event_id": event.get("event_id", ""),
"greenhelix.event_type": event_type,
AgentSpanAttributes.ESCROW_ID: escrow_id,
},
) as event_span:
event_span.set_status(StatusCode.OK)
span.set_attribute("lifecycle.event_count", len(lifecycle["events"]))
return lifecycle
Parent-Child Span Relationships
OpenTelemetry automatically manages parent-child relationships through Python context managers. When you nest start_span calls, each inner span becomes a child of the outer span. This creates the tree structure visible in trace visualization tools like Jaeger.
For agent commerce, the hierarchy typically looks like:
transaction.data_analysis_report (root)
|-- step.find_provider
| |-- call.agent-marketplace.find_provider (CLIENT)
|-- escrow.create
| |-- call.agent-marketplace.create_escrow (CLIENT)
|-- step.fetch_data
| |-- call.agent-data-source.fetch_dataset (CLIENT)
|-- step.analyze
| |-- call.agent-analysis.analyze_data (CLIENT)
|-- step.deliver
| |-- call.agent-delivery.deliver_report (CLIENT)
|-- escrow.settle
|-- call.agent-marketplace.settle_escrow (CLIENT)
Each CLIENT span on the calling side has a corresponding SERVER span on the receiving side. The trace ID is the same across all agents, so a trace visualization tool shows the entire transaction as a single, unified trace spanning all five agents.
Chapter 4: MetricsCollector Class
Tracing gives you per-transaction detail. Metrics give you the aggregate view: how is the fleet performing right now, how does today compare to yesterday, and are we meeting our business targets? The MetricsCollector class provides a clean API for collecting, buffering, and exporting metrics from your agent fleet.
Metric Types
There are three fundamental metric types for agent commerce:
Counters are monotonically increasing values. Use them for total requests, total errors, total revenue, and total transactions. Counters answer "how many" questions.
Gauges are point-in-time values that can go up or down. Use them for active connections, queue depth, current balance, and active escrows. Gauges answer "how much right now" questions.
Histograms track the distribution of values. Use them for latency, transaction amounts, and response sizes. Histograms answer "what is the distribution" questions, giving you percentiles (p50, p95, p99) rather than just averages.
The MetricsCollector Implementation
import time
import math
import threading
from collections import defaultdict
from dataclasses import dataclass, field
from typing import Dict, List, Optional, Callable, Any
@dataclass
class MetricPoint:
"""A single metric data point."""
name: str
value: float
dimensions: Dict[str, str]
timestamp: float
metric_type: str # "counter", "gauge", "histogram"
class HistogramBuckets:
"""Tracks value distribution for histogram metrics."""
def __init__(self, boundaries: List[float] = None):
self.boundaries = boundaries or [
5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000, 10000
]
self.buckets = [0] * (len(self.boundaries) + 1)
self.sum = 0.0
self.count = 0
self.min = float("inf")
self.max = float("-inf")
self._values: List[float] = []
def observe(self, value: float):
self.sum += value
self.count += 1
self.min = min(self.min, value)
self.max = max(self.max, value)
self._values.append(value)
for i, boundary in enumerate(self.boundaries):
if value <= boundary:
self.buckets[i] += 1
return
self.buckets[-1] += 1
def percentile(self, p: float) -> float:
if not self._values:
return 0.0
sorted_vals = sorted(self._values)
idx = int(math.ceil(p / 100.0 * len(sorted_vals))) - 1
return sorted_vals[max(0, idx)]
def reset(self):
self.buckets = [0] * (len(self.boundaries) + 1)
self.sum = 0.0
self.count = 0
self.min = float("inf")
self.max = float("-inf")
self._values = []
class MetricsCollector:
"""Collects, buffers, and exports agent commerce metrics."""
def __init__(
self,
agent_id: str,
client=None,
flush_interval_seconds: float = 60.0,
buffer_size: int = 1000,
):
self.agent_id = agent_id
self._client = client
self._flush_interval = flush_interval_seconds
self._buffer_size = buffer_size
self._counters: Dict[str, float] = defaultdict(float)
self._gauges: Dict[str, float] = {}
self._histograms: Dict[str, HistogramBuckets] = {}
self._buffer: List[MetricPoint] = []
self._lock = threading.Lock()
self._dimensions_registry: Dict[str, Dict[str, str]] = {}
self._flush_callbacks: List[Callable] = []
self._running = False
self._flush_thread: Optional[threading.Thread] = None
def start(self):
"""Start the background flush thread."""
self._running = True
self._flush_thread = threading.Thread(
target=self._flush_loop, daemon=True
)
self._flush_thread.start()
def stop(self):
"""Stop the background flush thread and flush remaining metrics."""
self._running = False
if self._flush_thread:
self._flush_thread.join(timeout=5.0)
self.flush()
# --- Counter Operations ---
def increment(
self,
name: str,
value: float = 1.0,
dimensions: Dict[str, str] = None,
):
"""Increment a counter metric."""
key = self._make_key(name, dimensions)
with self._lock:
self._counters[key] += value
self._dimensions_registry[key] = dimensions or {}
self._buffer_point(name, self._counters[key], dimensions, "counter")
# --- Gauge Operations ---
def gauge_set(
self,
name: str,
value: float,
dimensions: Dict[str, str] = None,
):
"""Set a gauge metric to an absolute value."""
key = self._make_key(name, dimensions)
with self._lock:
self._gauges[key] = value
self._dimensions_registry[key] = dimensions or {}
self._buffer_point(name, value, dimensions, "gauge")
def gauge_increment(
self,
name: str,
value: float = 1.0,
dimensions: Dict[str, str] = None,
):
"""Increment a gauge metric."""
key = self._make_key(name, dimensions)
with self._lock:
self._gauges[key] = self._gauges.get(key, 0.0) + value
self._dimensions_registry[key] = dimensions or {}
self._buffer_point(name, self._gauges[key], dimensions, "gauge")
# --- Histogram Operations ---
def observe(
self,
name: str,
value: float,
dimensions: Dict[str, str] = None,
boundaries: List[float] = None,
):
"""Record an observation in a histogram metric."""
key = self._make_key(name, dimensions)
with self._lock:
if key not in self._histograms:
self._histograms[key] = HistogramBuckets(boundaries)
self._histograms[key].observe(value)
self._dimensions_registry[key] = dimensions or {}
self._buffer_point(name, value, dimensions, "histogram")
# --- Business Metric Helpers ---
def record_transaction(
self,
operation: str,
duration_ms: float,
amount: float,
status: str,
peer_agent: str = "",
):
"""Record a complete transaction with all standard metrics."""
dims = {
"operation": operation,
"status": status,
"peer_agent": peer_agent,
}
self.increment("transactions.total", 1.0, dims)
self.observe("transactions.duration_ms", duration_ms, dims)
self.observe("transactions.amount", amount, dims)
if status == "error":
self.increment("transactions.errors", 1.0, dims)
def record_revenue(self, amount: float, source:
…(truncated)