# Greenhelix Agent Observability Stack

> Agent Observability Stack: Distributed Tracing, Metrics, and Alerting for Multi-Agent Systems. Build a complete observability stack for agent commerce: OpenTelemetry integration, distributed tracing across agent calls, custom metrics, anomaly detection, dashboard design, SLA monitoring, and cost attribution. Includes detailed Python code examples with full API integration.

- Skill: `lord1egypt/greenhelix-agent-observability-stack` (Agent Skill)
- Install (CLI): `npx skillmds@latest add lord1egypt/greenhelix-agent-observability-stack`
- Raw SKILL.md: https://api.skillmd.com/api/skills/lord1egypt/greenhelix-agent-observability-stack/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: AI & ML
- License: MIT
- Author: Lord1Egypt (https://skillmd.com/u/lord1egypt)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/lord1egypt/greenhelix-agent-observability-stack

---

# Agent Observability Stack: Distributed Tracing, Metrics, and Alerting for Multi-Agent Systems

> **Notice**: This is an educational guide with illustrative code examples.
> It does not execute code, require credentials, or install dependencies.
> All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which
> provides 500 free credits — no API key required to get started.


When Agent A calls Agent B which calls Agent C, and the transaction takes 12 seconds instead of 200ms, where is the bottleneck? When your agent fleet processes 10,000 transactions per day and revenue drops 15%, which agent is underperforming? When an escrow settlement fails silently at 3am, how quickly do you find out? Traditional monitoring tools -- ping checks, CPU graphs, uptime dashboards -- cannot answer these questions for distributed agent systems. They were built for monoliths and simple request-response services. Agent commerce is fundamentally different: transactions span multiple autonomous agents, each with its own state, pricing, and failure modes. A single customer-facing operation might traverse five agents, two escrow contracts, and three separate billing events before completing. The failure surface is combinatorial, not linear.
You need observability, not just monitoring. Monitoring tells you something is broken. Observability tells you why, where, and how to fix it. For agent commerce, this means distributed tracing to follow transactions across agent boundaries, custom metrics to measure business outcomes (not just infrastructure health), and intelligent alerting to catch problems before your users do. The difference between a 15-minute outage and a 4-hour revenue leak is whether your observability stack understands agent-to-agent transaction flows.
This guide builds a complete observability stack from scratch. We use OpenTelemetry standards for interoperability, integrate directly with GreenHelix's metrics and event tools for agent-specific telemetry, and build anomaly detection that understands the patterns unique to agent commerce. Every component is production Python code you can deploy today. By the end, you will have distributed tracing across agent calls, custom business metrics, anomaly detection, dashboards, alerting with escalation policies, and SLA monitoring with cost attribution.

## What You'll Learn
- Chapter 1: The Three Pillars of Agent Observability
- Chapter 2: AgentTracer Class
- Chapter 3: Distributed Tracing Across Agent Calls
- Chapter 4: MetricsCollector Class
- Chapter 5: Anomaly Detection
- Chapter 6: Dashboard Design
- Chapter 7: AlertManager Class
- Chapter 8: SLA Monitoring and Cost Attribution
- What's Next

## Full Guide

# Agent Observability Stack: Distributed Tracing, Metrics, and Alerting for Multi-Agent Systems

When Agent A calls Agent B which calls Agent C, and the transaction takes 12 seconds instead of 200ms, where is the bottleneck? When your agent fleet processes 10,000 transactions per day and revenue drops 15%, which agent is underperforming? When an escrow settlement fails silently at 3am, how quickly do you find out? Traditional monitoring tools -- ping checks, CPU graphs, uptime dashboards -- cannot answer these questions for distributed agent systems. They were built for monoliths and simple request-response services. Agent commerce is fundamentally different: transactions span multiple autonomous agents, each with its own state, pricing, and failure modes. A single customer-facing operation might traverse five agents, two escrow contracts, and three separate billing events before completing. The failure surface is combinatorial, not linear.

You need observability, not just monitoring. Monitoring tells you something is broken. Observability tells you why, where, and how to fix it. For agent commerce, this means distributed tracing to follow transactions across agent boundaries, custom metrics to measure business outcomes (not just infrastructure health), and intelligent alerting to catch problems before your users do. The difference between a 15-minute outage and a 4-hour revenue leak is whether your observability stack understands agent-to-agent transaction flows.

This guide builds a complete observability stack from scratch. We use OpenTelemetry standards for interoperability, integrate directly with GreenHelix's metrics and event tools for agent-specific telemetry, and build anomaly detection that understands the patterns unique to agent commerce. Every component is production Python code you can deploy today. By the end, you will have distributed tracing across agent calls, custom business metrics, anomaly detection, dashboards, alerting with escalation policies, and SLA monitoring with cost attribution.

---


> **Getting started**: All examples in this guide work with the GreenHelix sandbox
> (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required.

## Chapter 1: The Three Pillars of Agent Observability

Observability for distributed systems rests on three pillars: traces, metrics, and logs. Each answers a different class of question, and none is sufficient alone. For agent commerce, each pillar has specific requirements that differ from traditional web service observability.

### Traces: Following a Transaction Across Agent Boundaries

A trace is the complete record of a single transaction as it flows through your system. In agent commerce, a trace might begin when a customer agent requests a service, pass through a gateway agent, invoke a specialist agent, trigger an escrow creation, wait for fulfillment, and end with settlement and payment distribution. Each step in that journey is a span, and the collection of spans forms a trace.

The critical difference from traditional distributed tracing is that agent boundaries are not just service boundaries -- they are trust boundaries. When Agent A calls Agent B, there is an economic transaction embedded in that call. The trace must capture not just latency and status codes, but also billing events, escrow state transitions, and the economic context of each span. A span that shows "200 OK in 150ms" is incomplete if it does not also show "billed 0.003 credits, escrow ID esw_abc123 created."

Traces answer questions like: Why did this specific transaction fail? Where did the latency come from? Which agent in the chain was the bottleneck? Did the billing event match the actual work performed?

### Metrics: Measuring Agent Health, Performance, and Business Outcomes

Metrics are aggregated numerical measurements over time. Where traces give you detail about individual transactions, metrics give you the big picture. For agent commerce, metrics fall into three categories.

Infrastructure metrics measure the health of your agents as software systems: CPU usage, memory consumption, request queue depth, connection pool utilization. These are table stakes and most monitoring tools handle them well.

Performance metrics measure how your agents behave under load: request latency (p50, p95, p99), throughput (requests per second), error rates, and timeout rates. These require histograms and percentile calculations, not just averages.

Business metrics are where agent commerce diverges from traditional services. You need to track revenue per agent, cost per transaction, escrow settlement rates, dispute rates, SLA compliance percentages, and customer satisfaction proxies. These metrics directly measure whether your agent fleet is making money or losing it.

GreenHelix's `submit_metrics` tool accepts custom metric submissions, making it the natural sink for all three categories. The tool accepts metric name, value, dimensions, and timestamp, allowing you to build rich, queryable metric series.

### Logs: Structured Event Logging for Debugging

Logs are discrete events with context. In agent commerce, structured logs are essential because you need to correlate log events across agent boundaries. A plain text log line like "Error processing request" is useless when you have 50 agents each producing thousands of log lines per hour.

Structured logs include the trace ID, span ID, agent ID, transaction ID, and any relevant business context as machine-parseable fields. When something goes wrong, you filter logs by trace ID and see every event from every agent involved in that transaction, in chronological order.

GreenHelix's `get_events` tool retrieves event streams that serve as a structured log source. Events include agent interactions, billing events, escrow state changes, and webhook deliveries. By correlating your application logs with GreenHelix events, you get a complete picture of what happened and why.

### Why All Three Matter

Metrics tell you something is wrong: "Error rate spiked to 5% at 14:32." Traces tell you where and why: "Transaction txn_789 failed at Agent C because the escrow creation timed out after 30 seconds." Logs give you the details: "Agent C's connection pool was exhausted because Agent D was not releasing connections after failed settlements."

Without metrics, you do not know there is a problem until users complain. Without traces, you cannot pinpoint which agent or which step is responsible. Without logs, you cannot understand the root cause well enough to fix it. Agent commerce multiplies this dependency because the interactions are more complex, the failure modes are more subtle, and the economic consequences of missed problems are direct and measurable.

### GreenHelix Tools for Observability

Three GreenHelix tools form the foundation of our observability stack:

**submit_metrics** accepts custom metric data points with dimensions. We use this to push agent performance metrics, business metrics, and custom counters into GreenHelix's metric store.

**get_events** retrieves event streams filtered by agent, time range, and event type. We use this to pull structured event data for trace correlation and log enrichment.

**register_webhook** creates webhook subscriptions for specific event types. We use this to trigger real-time alerts when critical events occur, rather than polling for problems.

```python
from greenhelix import GreenHelixClient

client = GreenHelixClient(api_key="your-api-key")

# Submit a custom metric
client.execute_tool("submit_metrics", {
    "agent_id": "agent-payment-processor",
    "metrics": [
        {
            "name": "transaction.latency_ms",
            "value": 245.3,
            "dimensions": {
                "agent": "agent-payment-processor",
                "operation": "process_payment",
                "status": "success"
            },
            "timestamp": "2026-04-07T14:30:00Z"
        }
    ]
})

# Retrieve events for trace correlation
events = client.execute_tool("get_events", {
    "agent_id": "agent-payment-processor",
    "event_types": ["escrow.created", "escrow.settled", "billing.charged"],
    "start_time": "2026-04-07T14:00:00Z",
    "end_time": "2026-04-07T15:00:00Z"
})

# Register a webhook for real-time alerting
client.execute_tool("register_webhook", {
    "url": "https://alerts.myfleet.com/webhook",
    "event_types": ["escrow.failed", "billing.error"],
    "agent_id": "agent-payment-processor"
})
```

These three tools, combined with OpenTelemetry's tracing and metrics APIs, give us everything we need to build a production observability stack.

---

## Chapter 2: AgentTracer Class

The core of our observability stack is a tracer that understands agent commerce. Standard OpenTelemetry tracers create spans for HTTP requests and database calls. Our `AgentTracer` wraps OpenTelemetry to add agent-specific context: billing events, escrow state, economic metadata, and cross-agent trace propagation.

### Design Principles

The tracer must be lightweight. Adding observability should not measurably increase transaction latency. We target less than 1ms overhead per span, which means in-memory buffering with asynchronous export.

The tracer must be OpenTelemetry-compatible. This means traces can be exported to Jaeger, Zipkin, Grafana Tempo, or any other OTel-compatible backend. We do not lock you into a proprietary format.

The tracer must understand agent commerce primitives. Spans should automatically capture billing amounts, escrow IDs, agent identities, and transaction types without requiring manual instrumentation at every call site.

### The AgentTracer Implementation

```python
import time
import uuid
import threading
from dataclasses import dataclass, field
from typing import Optional, Dict, Any, List
from contextlib import contextmanager
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import (
    BatchSpanProcessor,
    SpanExporter,
    SpanExportResult,
)
from opentelemetry.trace import StatusCode, SpanKind
from opentelemetry.context import attach, detach, set_value, get_value


@dataclass
class AgentSpanAttributes:
    """Standard attributes for agent commerce spans."""
    AGENT_ID = "agent.id"
    AGENT_ROLE = "agent.role"
    TRANSACTION_ID = "agent.transaction.id"
    ESCROW_ID = "agent.escrow.id"
    BILLING_AMOUNT = "agent.billing.amount"
    BILLING_CURRENCY = "agent.billing.currency"
    OPERATION_TYPE = "agent.operation.type"
    PEER_AGENT_ID = "agent.peer.id"
    SLA_TIER = "agent.sla.tier"
    COST_CENTER = "agent.cost_center"


class GreenHelixSpanExporter(SpanExporter):
    """Exports spans as metrics to GreenHelix submit_metrics."""

    def __init__(self, client, agent_id: str, batch_size: int = 50):
        self._client = client
        self._agent_id = agent_id
        self._batch_size = batch_size

    def export(self, spans) -> SpanExportResult:
        metrics_batch = []
        for span in spans:
            duration_ms = (span.end_time - span.start_time) / 1_000_000
            attrs = dict(span.attributes) if span.attributes else {}

            metrics_batch.append({
                "name": "agent.span.duration_ms",
                "value": duration_ms,
                "dimensions": {
                    "agent": self._agent_id,
                    "operation": span.name,
                    "status": "ok" if span.status.status_code == StatusCode.OK
                             else "error",
                    "span_kind": span.kind.name if span.kind else "INTERNAL",
                    "peer_agent": attrs.get(AgentSpanAttributes.PEER_AGENT_ID, ""),
                },
                "timestamp": _ns_to_iso(span.end_time),
            })

            if AgentSpanAttributes.BILLING_AMOUNT in attrs:
                metrics_batch.append({
                    "name": "agent.billing.amount",
                    "value": float(attrs[AgentSpanAttributes.BILLING_AMOUNT]),
                    "dimensions": {
                        "agent": self._agent_id,
                        "operation": span.name,
                        "currency": attrs.get(
                            AgentSpanAttributes.BILLING_CURRENCY, "credits"
                        ),
                    },
                    "timestamp": _ns_to_iso(span.end_time),
                })

            if len(metrics_batch) >= self._batch_size:
                self._flush(metrics_batch)
                metrics_batch = []

        if metrics_batch:
            self._flush(metrics_batch)

        return SpanExportResult.SUCCESS

    def _flush(self, metrics):
        try:
            self._client.execute_tool("submit_metrics", {
                "agent_id": self._agent_id,
                "metrics": metrics,
            })
        except Exception:
            return SpanExportResult.FAILURE

    def shutdown(self):
        pass


def _ns_to_iso(ns_timestamp: int) -> str:
    """Convert nanosecond timestamp to ISO 8601."""
    import datetime
    dt = datetime.datetime.fromtimestamp(
        ns_timestamp / 1e9, tz=datetime.timezone.utc
    )
    return dt.isoformat()


class AgentTracer:
    """OpenTelemetry-compatible tracer for agent commerce."""

    def __init__(
        self,
        agent_id: str,
        client=None,
        service_name: str = None,
        exporters: List[SpanExporter] = None,
    ):
        self.agent_id = agent_id
        self._client = client
        self._service_name = service_name or f"agent-{agent_id}"

        provider = TracerProvider()

        if client:
            gh_exporter = GreenHelixSpanExporter(client, agent_id)
            provider.add_span_processor(BatchSpanProcessor(gh_exporter))

        if exporters:
            for exporter in exporters:
                provider.add_span_processor(BatchSpanProcessor(exporter))

        self._provider = provider
        self._tracer = provider.get_tracer(
            self._service_name, schema_url="https://greenhelix.net/schemas/agent/1.0"
        )

    @contextmanager
    def start_span(
        self,
        name: str,
        kind: SpanKind = SpanKind.INTERNAL,
        attributes: Dict[str, Any] = None,
        peer_agent_id: str = None,
    ):
        """Start a new span with agent commerce context."""
        attrs = {
            AgentSpanAttributes.AGENT_ID: self.agent_id,
        }
        if peer_agent_id:
            attrs[AgentSpanAttributes.PEER_AGENT_ID] = peer_agent_id
        if attributes:
            attrs.update(attributes)

        with self._tracer.start_as_current_span(
            name, kind=kind, attributes=attrs
        ) as span:
            yield span

    @contextmanager
    def trace_agent_call(
        self,
        target_agent_id: str,
        operation: str,
        transaction_id: str = None,
    ):
        """Trace a call from this agent to another agent."""
        txn_id = transaction_id or str(uuid.uuid4())
        attrs = {
            AgentSpanAttributes.OPERATION_TYPE: operation,
            AgentSpanAttributes.TRANSACTION_ID: txn_id,
        }

        with self.start_span(
            f"call.{target_agent_id}.{operation}",
            kind=SpanKind.CLIENT,
            attributes=attrs,
            peer_agent_id=target_agent_id,
        ) as span:
            yield span

    @contextmanager
    def trace_escrow(self, escrow_id: str, operation: str):
        """Trace an escrow lifecycle operation."""
        attrs = {
            AgentSpanAttributes.ESCROW_ID: escrow_id,
            AgentSpanAttributes.OPERATION_TYPE: f"escrow.{operation}",
        }

        with self.start_span(
            f"escrow.{operation}",
            kind=SpanKind.INTERNAL,
            attributes=attrs,
        ) as span:
            yield span

    @contextmanager
    def trace_billing(self, amount: float, currency: str = "credits"):
        """Trace a billing event."""
        attrs = {
            AgentSpanAttributes.BILLING_AMOUNT: amount,
            AgentSpanAttributes.BILLING_CURRENCY: currency,
            AgentSpanAttributes.OPERATION_TYPE: "billing.charge",
        }

        with self.start_span(
            "billing.charge",
            kind=SpanKind.INTERNAL,
            attributes=attrs,
        ) as span:
            yield span

    def get_trace_context(self) -> Dict[str, str]:
        """Extract current trace context for propagation to other agents."""
        span = trace.get_current_span()
        ctx = span.get_span_context()
        if not ctx or not ctx.is_valid:
            return {}

        return {
            "traceparent": f"00-{format(ctx.trace_id, '032x')}-"
                           f"{format(ctx.span_id, '016x')}-"
                           f"{'01' if ctx.trace_flags & 1 else '00'}",
            "x-agent-id": self.agent_id,
        }

    def inject_context(self, headers: Dict[str, str]) -> Dict[str, str]:
        """Inject trace context into outgoing request headers."""
        headers.update(self.get_trace_context())
        return headers

    def shutdown(self):
        """Flush pending spans and shut down the tracer."""
        self._provider.shutdown()
```

### Using the AgentTracer

The tracer integrates naturally with GreenHelix client calls. Each API interaction gets a span, and the spans automatically capture timing, status, and agent commerce attributes.

```python
from greenhelix import GreenHelixClient

client = GreenHelixClient(api_key="your-api-key")
tracer = AgentTracer(agent_id="agent-marketplace", client=client)

# Trace a multi-step transaction
with tracer.trace_agent_call("agent-data-provider", "fetch_dataset") as span:
    result = client.execute_tool("call_agent", {
        "target": "agent-data-provider",
        "operation": "fetch_dataset",
        "params": {"dataset_id": "ds_12345"},
        "headers": tracer.get_trace_context(),
    })
    span.set_attribute("dataset.size_bytes", result.get("size", 0))

    if result.get("status") == "error":
        span.set_status(StatusCode.ERROR, result.get("message", ""))
    else:
        span.set_status(StatusCode.OK)

# Trace an escrow lifecycle
with tracer.trace_escrow("esw_abc123", "create") as span:
    escrow = client.execute_tool("create_escrow", {
        "amount": 5.00,
        "payer": "agent-buyer",
        "payee": "agent-seller",
    })
    span.set_attribute("escrow.amount", 5.00)
    span.set_status(StatusCode.OK)
```

The key design choice is that `AgentTracer` wraps OpenTelemetry rather than replacing it. You can add any standard OTel exporter (Jaeger, OTLP, console) alongside the GreenHelix exporter. This means your agent traces show up in your existing observability infrastructure without any migration.

The `GreenHelixSpanExporter` converts spans to metrics via `submit_metrics`. This dual-purpose export means every traced operation automatically generates latency and billing metrics, eliminating the need to instrument metrics separately for operations you are already tracing.

---

## Chapter 3: Distributed Tracing Across Agent Calls

Single-agent tracing is straightforward. The real challenge is tracing a transaction as it flows through multiple independent agents, each potentially running on different infrastructure, operated by different teams, and communicating through the GreenHelix gateway.

### The Problem: Trace Context Propagation

When Agent A calls Agent B, Agent B has no inherent knowledge of Agent A's trace. Without explicit context propagation, Agent B starts a new, disconnected trace. You end up with five isolated traces instead of one unified trace showing the complete transaction flow.

The W3C Trace Context standard solves this with two HTTP headers: `traceparent` (containing trace ID, span ID, and sampling flags) and `tracestate` (containing vendor-specific data). Our `AgentTracer` already generates these headers via `get_trace_context()`. The challenge is ensuring every agent in the chain extracts, uses, and propagates these headers.

### Trace Context in Agent-to-Agent Calls

Here is the pattern for traced agent-to-agent calls. The calling agent injects context into outgoing headers. The receiving agent extracts context and creates child spans.

```python
# === Calling Agent (Agent A) ===

class TracedAgentClient:
    """HTTP client that automatically propagates trace context."""

    def __init__(self, tracer: AgentTracer, client):
        self._tracer = tracer
        self._client = client

    def call_agent(
        self,
        target_agent: str,
        operation: str,
        params: Dict[str, Any],
        transaction_id: str = None,
    ) -> Dict[str, Any]:
        """Call another agent with automatic trace propagation."""
        with self._tracer.trace_agent_call(
            target_agent, operation, transaction_id
        ) as span:
            headers = self._tracer.get_trace_context()
            headers["x-transaction-id"] = transaction_id or ""

            result = self._client.execute_tool("call_agent", {
                "target": target_agent,
                "operation": operation,
                "params": params,
                "headers": headers,
            })

            span.set_attribute("response.status", result.get("status", "unknown"))

            if result.get("error"):
                span.set_status(StatusCode.ERROR, result["error"])
            else:
                span.set_status(StatusCode.OK)

            return result


# === Receiving Agent (Agent B) ===

from opentelemetry.trace.propagation import TraceContextTextMapPropagator
from opentelemetry.context import Context

class TracedAgentHandler:
    """Request handler that extracts and continues trace context."""

    def __init__(self, tracer: AgentTracer):
        self._tracer = tracer
        self._propagator = TraceContextTextMapPropagator()

    def handle_request(
        self,
        operation: str,
        params: Dict[str, Any],
        headers: Dict[str, str],
    ) -> Dict[str, Any]:
        """Handle an incoming agent request, continuing the trace."""
        # Extract trace context from incoming headers
        ctx = self._propagator.extract(carrier=headers)

        # Create a child span under the extracted context
        token = attach(ctx)
        try:
            with self._tracer.start_span(
                f"handle.{operation}",
                kind=SpanKind.SERVER,
                attributes={
                    "caller.agent_id": headers.get("x-agent-id", "unknown"),
                    AgentSpanAttributes.OPERATION_TYPE: operation,
                },
            ) as span:
                result = self._dispatch(operation, params, span)
                return result
        finally:
            detach(token)

    def _dispatch(
        self, operation: str, params: Dict[str, Any], span
    ) -> Dict[str, Any]:
        """Route to the appropriate operation handler."""
        handler = getattr(self, f"op_{operation}", None)
        if not handler:
            span.set_status(StatusCode.ERROR, f"Unknown operation: {operation}")
            return {"error": f"Unknown operation: {operation}"}
        return handler(params, span)
```

### Tracing a Complete Transaction Through Five Agents

Consider an agent commerce transaction where a customer agent orders a data analysis report. The transaction flows through five agents:

1. **Customer Agent** -- initiates the request
2. **Marketplace Agent** -- matches request to provider
3. **Data Provider Agent** -- supplies raw data
4. **Analysis Agent** -- processes and analyzes data
5. **Delivery Agent** -- formats and delivers the report

```python
def traced_execute(
    client: TracedAgentClient,
    tracer: AgentTracer,
    request: Dict[str, Any],
) -> Dict[str, Any]:
    """Execute a full traced transaction across multiple agents."""
    transaction_id = str(uuid.uuid4())

    with tracer.start_span(
        "transaction.data_analysis_report",
        kind=SpanKind.INTERNAL,
        attributes={
            AgentSpanAttributes.TRANSACTION_ID: transaction_id,
            AgentSpanAttributes.OPERATION_TYPE: "data_analysis_report",
        },
    ) as root_span:

        # Step 1: Find a provider through the marketplace
        with tracer.start_span("step.find_provider") as step_span:
            match = client.call_agent(
                "agent-marketplace",
                "find_provider",
                {"capability": "data_analysis", "budget": request["budget"]},
                transaction_id=transaction_id,
            )
            step_span.set_attribute("provider.matched", match.get("provider_id", ""))

        provider_id = match["provider_id"]

        # Step 2: Create escrow for the transaction
        with tracer.trace_escrow(f"esw_{transaction_id[:8]}", "create") as esc_span:
            escrow = client.call_agent(
                "agent-marketplace",
                "create_escrow",
                {
                    "amount": match["price"],
                    "payer": request["customer_agent"],
                    "payee": provider_id,
                },
                transaction_id=transaction_id,
            )
            esc_span.set_attribute("escrow.id", escrow.get("escrow_id", ""))

        # Step 3: Fetch raw data
        with tracer.start_span("step.fetch_data") as step_span:
            data = client.call_agent(
                match["data_source"],
                "fetch_dataset",
                {"dataset_id": request["dataset_id"]},
                transaction_id=transaction_id,
            )
            step_span.set_attribute("data.records", data.get("record_count", 0))

        # Step 4: Run analysis
        with tracer.start_span("step.analyze") as step_span:
            analysis = client.call_agent(
                provider_id,
                "analyze_data",
                {"data": data["data"], "analysis_type": request["analysis_type"]},
                transaction_id=transaction_id,
            )
            step_span.set_attribute("analysis.confidence", analysis.get("confidence", 0))

        # Step 5: Deliver report
        with tracer.start_span("step.deliver") as step_span:
            delivery = client.call_agent(
                "agent-delivery",
                "deliver_report",
                {
                    "report": analysis["report"],
                    "recipient": request["customer_agent"],
                    "format": request.get("format", "pdf"),
                },
                transaction_id=transaction_id,
            )
            step_span.set_attribute("delivery.channel", delivery.get("channel", ""))

        # Step 6: Settle escrow
        with tracer.trace_escrow(escrow["escrow_id"], "settle") as esc_span:
            settlement = client.call_agent(
                "agent-marketplace",
                "settle_escrow",
                {"escrow_id": escrow["escrow_id"]},
                transaction_id=transaction_id,
            )
            esc_span.set_attribute("settlement.status", settlement.get("status", ""))

        root_span.set_status(StatusCode.OK)
        return {
            "transaction_id": transaction_id,
            "report_url": delivery.get("url"),
            "total_cost": match["price"],
        }
```

### Trace Correlation with GreenHelix Events

GreenHelix events have their own event IDs. To get a complete picture, you need to correlate your OpenTelemetry trace IDs with GreenHelix event IDs. The approach is to embed the trace ID in your GreenHelix API calls and then join on it during analysis.

```python
def trace_escrow_lifecycle(
    tracer: AgentTracer,
    client,
    escrow_id: str,
) -> Dict[str, Any]:
    """Trace and correlate an escrow's full lifecycle with GreenHelix events."""
    trace_ctx = tracer.get_trace_context()
    trace_id = trace_ctx.get("traceparent", "").split("-")[1] if trace_ctx else ""

    with tracer.trace_escrow(escrow_id, "lifecycle") as span:
        # Fetch GreenHelix events for this escrow
        events = client.execute_tool("get_events", {
            "filters": {"escrow_id": escrow_id},
            "start_time": "2026-04-07T00:00:00Z",
            "end_time": "2026-04-07T23:59:59Z",
        })

        lifecycle = {
            "escrow_id": escrow_id,
            "trace_id": trace_id,
            "events": [],
        }

        for event in events.get("events", []):
            event_type = event.get("type", "")
            lifecycle["events"].append({
                "type": event_type,
                "timestamp": event.get("timestamp"),
                "greenhelix_event_id": event.get("event_id"),
                "trace_id": trace_id,
            })

            # Create a child span for each GreenHelix event
            with tracer.start_span(
                f"greenhelix.event.{event_type}",
                attributes={
                    "greenhelix.event_id": event.get("event_id", ""),
                    "greenhelix.event_type": event_type,
                    AgentSpanAttributes.ESCROW_ID: escrow_id,
                },
            ) as event_span:
                event_span.set_status(StatusCode.OK)

        span.set_attribute("lifecycle.event_count", len(lifecycle["events"]))
        return lifecycle
```

### Parent-Child Span Relationships

OpenTelemetry automatically manages parent-child relationships through Python context managers. When you nest `start_span` calls, each inner span becomes a child of the outer span. This creates the tree structure visible in trace visualization tools like Jaeger.

For agent commerce, the hierarchy typically looks like:

```
transaction.data_analysis_report (root)
  |-- step.find_provider
  |     |-- call.agent-marketplace.find_provider (CLIENT)
  |-- escrow.create
  |     |-- call.agent-marketplace.create_escrow (CLIENT)
  |-- step.fetch_data
  |     |-- call.agent-data-source.fetch_dataset (CLIENT)
  |-- step.analyze
  |     |-- call.agent-analysis.analyze_data (CLIENT)
  |-- step.deliver
  |     |-- call.agent-delivery.deliver_report (CLIENT)
  |-- escrow.settle
        |-- call.agent-marketplace.settle_escrow (CLIENT)
```

Each CLIENT span on the calling side has a corresponding SERVER span on the receiving side. The trace ID is the same across all agents, so a trace visualization tool shows the entire transaction as a single, unified trace spanning all five agents.

---

## Chapter 4: MetricsCollector Class

Tracing gives you per-transaction detail. Metrics give you the aggregate view: how is the fleet performing right now, how does today compare to yesterday, and are we meeting our business targets? The `MetricsCollector` class provides a clean API for collecting, buffering, and exporting metrics from your agent fleet.

### Metric Types

There are three fundamental metric types for agent commerce:

**Counters** are monotonically increasing values. Use them for total requests, total errors, total revenue, and total transactions. Counters answer "how many" questions.

**Gauges** are point-in-time values that can go up or down. Use them for active connections, queue depth, current balance, and active escrows. Gauges answer "how much right now" questions.

**Histograms** track the distribution of values. Use them for latency, transaction amounts, and response sizes. Histograms answer "what is the distribution" questions, giving you percentiles (p50, p95, p99) rather than just averages.

### The MetricsCollector Implementation

```python
import time
import math
import threading
from collections import defaultdict
from dataclasses import dataclass, field
from typing import Dict, List, Optional, Callable, Any


@dataclass
class MetricPoint:
    """A single metric data point."""
    name: str
    value: float
    dimensions: Dict[str, str]
    timestamp: float
    metric_type: str  # "counter", "gauge", "histogram"


class HistogramBuckets:
    """Tracks value distribution for histogram metrics."""

    def __init__(self, boundaries: List[float] = None):
        self.boundaries = boundaries or [
            5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000, 10000
        ]
        self.buckets = [0] * (len(self.boundaries) + 1)
        self.sum = 0.0
        self.count = 0
        self.min = float("inf")
        self.max = float("-inf")
        self._values: List[float] = []

    def observe(self, value: float):
        self.sum += value
        self.count += 1
        self.min = min(self.min, value)
        self.max = max(self.max, value)
        self._values.append(value)

        for i, boundary in enumerate(self.boundaries):
            if value <= boundary:
                self.buckets[i] += 1
                return
        self.buckets[-1] += 1

    def percentile(self, p: float) -> float:
        if not self._values:
            return 0.0
        sorted_vals = sorted(self._values)
        idx = int(math.ceil(p / 100.0 * len(sorted_vals))) - 1
        return sorted_vals[max(0, idx)]

    def reset(self):
        self.buckets = [0] * (len(self.boundaries) + 1)
        self.sum = 0.0
        self.count = 0
        self.min = float("inf")
        self.max = float("-inf")
        self._values = []


class MetricsCollector:
    """Collects, buffers, and exports agent commerce metrics."""

    def __init__(
        self,
        agent_id: str,
        client=None,
        flush_interval_seconds: float = 60.0,
        buffer_size: int = 1000,
    ):
        self.agent_id = agent_id
        self._client = client
        self._flush_interval = flush_interval_seconds
        self._buffer_size = buffer_size

        self._counters: Dict[str, float] = defaultdict(float)
        self._gauges: Dict[str, float] = {}
        self._histograms: Dict[str, HistogramBuckets] = {}
        self._buffer: List[MetricPoint] = []
        self._lock = threading.Lock()

        self._dimensions_registry: Dict[str, Dict[str, str]] = {}
        self._flush_callbacks: List[Callable] = []

        self._running = False
        self._flush_thread: Optional[threading.Thread] = None

    def start(self):
        """Start the background flush thread."""
        self._running = True
        self._flush_thread = threading.Thread(
            target=self._flush_loop, daemon=True
        )
        self._flush_thread.start()

    def stop(self):
        """Stop the background flush thread and flush remaining metrics."""
        self._running = False
        if self._flush_thread:
            self._flush_thread.join(timeout=5.0)
        self.flush()

    # --- Counter Operations ---

    def increment(
        self,
        name: str,
        value: float = 1.0,
        dimensions: Dict[str, str] = None,
    ):
        """Increment a counter metric."""
        key = self._make_key(name, dimensions)
        with self._lock:
            self._counters[key] += value
            self._dimensions_registry[key] = dimensions or {}
            self._buffer_point(name, self._counters[key], dimensions, "counter")

    # --- Gauge Operations ---

    def gauge_set(
        self,
        name: str,
        value: float,
        dimensions: Dict[str, str] = None,
    ):
        """Set a gauge metric to an absolute value."""
        key = self._make_key(name, dimensions)
        with self._lock:
            self._gauges[key] = value
            self._dimensions_registry[key] = dimensions or {}
            self._buffer_point(name, value, dimensions, "gauge")

    def gauge_increment(
        self,
        name: str,
        value: float = 1.0,
        dimensions: Dict[str, str] = None,
    ):
        """Increment a gauge metric."""
        key = self._make_key(name, dimensions)
        with self._lock:
            self._gauges[key] = self._gauges.get(key, 0.0) + value
            self._dimensions_registry[key] = dimensions or {}
            self._buffer_point(name, self._gauges[key], dimensions, "gauge")

    # --- Histogram Operations ---

    def observe(
        self,
        name: str,
        value: float,
        dimensions: Dict[str, str] = None,
        boundaries: List[float] = None,
    ):
        """Record an observation in a histogram metric."""
        key = self._make_key(name, dimensions)
        with self._lock:
            if key not in self._histograms:
                self._histograms[key] = HistogramBuckets(boundaries)
            self._histograms[key].observe(value)
            self._dimensions_registry[key] = dimensions or {}
            self._buffer_point(name, value, dimensions, "histogram")

    # --- Business Metric Helpers ---

    def record_transaction(
        self,
        operation: str,
        duration_ms: float,
        amount: float,
        status: str,
        peer_agent: str = "",
    ):
        """Record a complete transaction with all standard metrics."""
        dims = {
            "operation": operation,
            "status": status,
            "peer_agent": peer_agent,
        }

        self.increment("transactions.total", 1.0, dims)
        self.observe("transactions.duration_ms", duration_ms, dims)
        self.observe("transactions.amount", amount, dims)

        if status == "error":
            self.increment("transactions.errors", 1.0, dims)

    def record_revenue(self, amount: float, source: 

…(truncated)
