# Greenhelix Agent Testing Observability

> The Agent Testing & Observability Cookbook: Ship Reliable Agent Commerce Systems. Practitioner cookbook for testing and monitoring agent commerce: tool contract tests, workflow saga tests, chaos injection, OpenTelemetry tracing, health checks, alerting, and CI/CD pipelines.

- Skill: `lord1egypt/greenhelix-agent-testing-observability` (Agent Skill)
- Install (CLI): `npx skillmds@latest add lord1egypt/greenhelix-agent-testing-observability`
- Raw SKILL.md: https://api.skillmd.com/api/skills/lord1egypt/greenhelix-agent-testing-observability/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: DevOps & Infra
- License: MIT
- Author: Lord1Egypt (https://skillmd.com/u/lord1egypt)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/lord1egypt/greenhelix-agent-testing-observability

---

# The Agent Testing & Observability Cookbook: Ship Reliable Agent Commerce Systems

> **Notice**: This is an educational guide with illustrative code examples.
> It does not execute code or install dependencies.
> All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which
> provides 500 free credits — no API key required to get started.
>
> **Referenced credentials** (you supply these in your own environment):
> - `GREENHELIX_API_KEY`: API authentication for GreenHelix gateway (read/write access to purchased API tools only)
> - `AGENT_SIGNING_KEY`: Cryptographic signing key for agent identity (Ed25519 key pair for request signing)


Your agent commerce system works on your laptop. It passes a smoke test against the GreenHelix sandbox. You deploy to production on a Friday afternoon and go home. By Saturday morning, a retry loop has created 47 duplicate escrows, a performance escrow released funds against stale metrics, and a settlement webhook silently failed for six hours because the endpoint returned 503 and nobody was watching. The system was never tested for these failures because the traditional testing pyramid -- unit tests at the bottom, integration tests in the middle, end-to-end tests at the top -- was not designed for autonomous agents that make financial decisions across unreliable networks against counterparties that may themselves be failing. This guide rebuilds the testing pyramid for agent commerce, then layers production observability, chaos testing, alerting, and CI/CD on top. Every pattern is backed by working Python code, grounded in the 260-test suite that ships with the GreenHelix gateway, and designed to be copied directly into your project.
1. [The Testing Pyramid for Agent Systems](#chapter-1-the-testing-pyramid-for-agent-systems)
2. [Tool-Level Testing Patterns](#chapter-2-tool-level-testing-patterns)

## What You'll Learn
- Chapter 1: The Testing Pyramid for Agent Systems
- Chapter 2: Tool-Level Testing Patterns
- Chapter 3: Workflow & Integration Testing
- Chapter 4: Chaos Testing for Agent Commerce
- Chapter 5: Production Observability
- Chapter 6: Alerting & Incident Response
- Chapter 7: CI/CD for Agent Systems
- Chapter 8: What to Do Next

## Full Guide

# The Agent Testing & Observability Cookbook: Ship Reliable Agent Commerce Systems

Your agent commerce system works on your laptop. It passes a smoke test against the GreenHelix sandbox. You deploy to production on a Friday afternoon and go home. By Saturday morning, a retry loop has created 47 duplicate escrows, a performance escrow released funds against stale metrics, and a settlement webhook silently failed for six hours because the endpoint returned 503 and nobody was watching. The system was never tested for these failures because the traditional testing pyramid -- unit tests at the bottom, integration tests in the middle, end-to-end tests at the top -- was not designed for autonomous agents that make financial decisions across unreliable networks against counterparties that may themselves be failing. This guide rebuilds the testing pyramid for agent commerce, then layers production observability, chaos testing, alerting, and CI/CD on top. Every pattern is backed by working Python code, grounded in the 260-test suite that ships with the GreenHelix gateway, and designed to be copied directly into your project.

---

## Table of Contents

1. [The Testing Pyramid for Agent Systems](#chapter-1-the-testing-pyramid-for-agent-systems)
2. [Tool-Level Testing Patterns](#chapter-2-tool-level-testing-patterns)
3. [Workflow & Integration Testing](#chapter-3-workflow--integration-testing)
4. [Chaos Testing for Agent Commerce](#chapter-4-chaos-testing-for-agent-commerce)
5. [Production Observability](#chapter-5-production-observability)
6. [Alerting & Incident Response](#chapter-6-alerting--incident-response)
7. [CI/CD for Agent Systems](#chapter-7-cicd-for-agent-systems)
8. [What to Do Next](#chapter-8-what-to-do-next)

---

## Chapter 1: The Testing Pyramid for Agent Systems

### Why the Traditional Pyramid Breaks

The standard testing pyramid assumes your code calls functions that return values. Unit tests verify individual functions. Integration tests verify that modules compose correctly. End-to-end tests verify the full user flow. This model works when the system under test is deterministic, when function calls do not have financial side effects, and when failure modes are limited to "returns wrong value" or "throws exception."

Agent commerce systems violate all three assumptions. A call to `create_escrow` locks real funds. A call to `release_escrow` transfers real money. A retry that fires twice creates two escrows, not one error. The failure modes are not "wrong return value" -- they are "agent paid twice for the same work," "escrow timed out but funds are still locked," and "settlement succeeded on the gateway but the webhook notification was lost." Traditional unit tests cannot catch these failures because they test the code in isolation from the financial state machine. Traditional end-to-end tests cannot catch them because they run the happy path once and call it done.

### The Agent Testing Pyramid

Agent commerce requires a four-layer testing pyramid that maps to the actual failure modes:

```
                    ╱╲
                   ╱  ╲
                  ╱Chaos╲           Layer 4: Chaos tests
                 ╱  Tests ╲         (failure injection, timeouts,
                ╱──────────╲        concurrent load)
               ╱ Multi-Agent ╲      Layer 3: Multi-agent workflow tests
              ╱   Workflows   ╲     (sagas, rollbacks, webhook delivery)
             ╱─────────────────╲
            ╱   Tool Contract    ╲   Layer 2: Tool-level contract tests
           ╱      Tests           ╲  (schema, idempotency, permissions)
          ╱────────────────────────╲
         ╱    Deterministic Mocks    ╲ Layer 1: Mock-based unit tests
        ╱      (fast, offline)        ╲ (business logic, validation)
       ╱───────────────────────────────╲
```

**Layer 1: Deterministic mocks** test your business logic without hitting the gateway. They run in milliseconds and catch logic errors: wrong amount calculations, missing trust checks, incorrect state transitions. These are 60% of your tests.

**Layer 2: Tool contract tests** verify that each GreenHelix tool accepts the expected input schema, returns the expected output shape, and produces the correct error codes for invalid input. These run against the sandbox and catch API contract changes. These are 25% of your tests.

**Layer 3: Multi-agent workflow tests** verify complete business flows: marketplace listing through escrow release through settlement. They test the saga pattern (multi-step rollback on failure) and webhook delivery. These are 10% of your tests.

**Layer 4: Chaos tests** inject failures -- network timeouts, random tool errors, concurrent duplicate requests -- and verify that the system recovers without financial inconsistency. These are 5% of your tests but catch the bugs that cost the most money.

### The AgentTestHarness

Every test in this guide uses the `AgentTestHarness` class. It manages fixtures, provides deterministic mocks for Layer 1, and switches to sandbox mode for Layers 2-4.

```python
import pytest
import time
import json
import uuid
import requests
from unittest.mock import MagicMock, patch
from typing import Optional
from dataclasses import dataclass, field


@dataclass
class MockResponse:
    """Deterministic mock for a GreenHelix tool response."""
    tool: str
    status: str = "success"
    data: dict = field(default_factory=dict)
    error_code: Optional[str] = None
    error_message: Optional[str] = None

    def to_dict(self) -> dict:
        result = {"status": self.status}
        if self.status == "success":
            result.update(self.data)
        else:
            result["error"] = {
                "code": self.error_code or "unknown_error",
                "message": self.error_message or "An error occurred",
            }
        return result


class AgentTestHarness:
    """Test harness for GreenHelix agent commerce systems.

    Manages fixtures, mocks, and sandbox connections for all four
    layers of the agent testing pyramid.

    Usage:
        harness = AgentTestHarness(
            api_key="test-key",
            agent_id="test-agent",
            base_url="https://sandbox.greenhelix.net/v1",
        )

        # Layer 1: deterministic mocks
        harness.mock_tool("get_balance", {"balance": "100.00"})
        result = harness.execute("get_balance", {})
        assert result["balance"] == "100.00"

        # Layer 2+: sandbox mode
        harness.use_sandbox()
        result = harness.execute("get_balance", {})
    """

    def __init__(
        self,
        api_key: str,
        agent_id: str,
        base_url: str = "https://sandbox.greenhelix.net/v1",
    ):
        self.api_key = api_key
        self.agent_id = agent_id
        self.base_url = base_url
        self._mocks: dict[str, MockResponse] = {}
        self._call_log: list[dict] = []
        self._sandbox_mode = False
        self._session = requests.Session()
        self._session.headers.update({
            "Content-Type": "application/json",
            "Authorization": f"Bearer {api_key}",
        })

    # ── Mode Control ───────────────────────────────────────────

    def use_mocks(self):
        """Switch to deterministic mock mode (Layer 1)."""
        self._sandbox_mode = False

    def use_sandbox(self):
        """Switch to live sandbox mode (Layer 2+)."""
        self._sandbox_mode = True

    # ── Mock Registration ──────────────────────────────────────

    def mock_tool(self, tool: str, data: dict, status: str = "success"):
        """Register a deterministic mock response for a tool."""
        self._mocks[tool] = MockResponse(tool=tool, status=status, data=data)

    def mock_tool_error(
        self, tool: str, error_code: str, error_message: str
    ):
        """Register a deterministic error response for a tool."""
        self._mocks[tool] = MockResponse(
            tool=tool,
            status="error",
            error_code=error_code,
            error_message=error_message,
        )

    def mock_tool_sequence(self, tool: str, responses: list[dict]):
        """Register a sequence of responses for successive calls."""
        self._mock_sequences = getattr(self, "_mock_sequences", {})
        self._mock_sequences[tool] = list(responses)

    # ── Execution ──────────────────────────────────────────────

    def execute(self, tool: str, input_data: dict) -> dict:
        """Execute a tool against mocks or sandbox."""
        call_record = {
            "tool": tool,
            "input": input_data,
            "timestamp": time.time(),
        }

        if self._sandbox_mode:
            resp = self._session.post(
                f"{self.base_url}/v1",
                json={"tool": tool, "input": input_data},
            )
            resp.raise_for_status()
            result = resp.json()
        else:
            # Check sequences first
            sequences = getattr(self, "_mock_sequences", {})
            if tool in sequences and sequences[tool]:
                result = sequences[tool].pop(0)
            elif tool in self._mocks:
                result = self._mocks[tool].to_dict()
            else:
                raise ValueError(
                    f"No mock registered for tool '{tool}'. "
                    f"Register with harness.mock_tool('{tool}', {{...}})"
                )

        call_record["result"] = result
        self._call_log.append(call_record)
        return result

    # ── Assertions ─────────────────────────────────────────────

    def assert_tool_called(self, tool: str, times: Optional[int] = None):
        """Assert that a tool was called, optionally a specific number of times."""
        calls = [c for c in self._call_log if c["tool"] == tool]
        assert len(calls) > 0, f"Tool '{tool}' was never called"
        if times is not None:
            assert len(calls) == times, (
                f"Tool '{tool}' called {len(calls)} times, expected {times}"
            )

    def assert_tool_not_called(self, tool: str):
        """Assert that a tool was never called."""
        calls = [c for c in self._call_log if c["tool"] == tool]
        assert len(calls) == 0, (
            f"Tool '{tool}' was called {len(calls)} times, expected 0"
        )

    def assert_call_order(self, tools: list[str]):
        """Assert that tools were called in a specific order."""
        called_tools = [c["tool"] for c in self._call_log]
        idx = 0
        for tool in tools:
            try:
                idx = called_tools.index(tool, idx) + 1
            except ValueError:
                assert False, (
                    f"Expected '{tool}' after position {idx} in call log. "
                    f"Actual order: {called_tools}"
                )

    def get_calls(self, tool: Optional[str] = None) -> list[dict]:
        """Get call log, optionally filtered by tool name."""
        if tool:
            return [c for c in self._call_log if c["tool"] == tool]
        return list(self._call_log)

    def reset(self):
        """Clear all mocks and call history."""
        self._mocks.clear()
        self._call_log.clear()
        if hasattr(self, "_mock_sequences"):
            self._mock_sequences.clear()
```

### The conftest.py: Reusable Fixtures

Drop this `conftest.py` into your test directory. Every test file in this guide imports from it.

```python
# tests/conftest.py
import os
import uuid
import pytest


@pytest.fixture
def api_key():
    """API key for sandbox testing. Uses env var or test default."""
    return os.environ.get("GREENHELIX_API_KEY", "test-api-key-sandbox")


@pytest.fixture
def base_url():
    """Sandbox URL for integration tests."""
    return os.environ.get(
        "GREENHELIX_BASE_URL", "https://sandbox.greenhelix.net/v1"
    )


@pytest.fixture
def agent_id():
    """Unique agent ID per test run to prevent collision."""
    return f"test-agent-{uuid.uuid4().hex[:12]}"


@pytest.fixture
def buyer_id():
    """Unique buyer agent ID."""
    return f"test-buyer-{uuid.uuid4().hex[:12]}"


@pytest.fixture
def seller_id():
    """Unique seller agent ID."""
    return f"test-seller-{uuid.uuid4().hex[:12]}"


@pytest.fixture
def harness(api_key, agent_id, base_url):
    """AgentTestHarness in mock mode. Call harness.use_sandbox() for live."""
    h = AgentTestHarness(
        api_key=api_key,
        agent_id=agent_id,
        base_url=base_url,
    )
    h.use_mocks()
    return h


@pytest.fixture
def sandbox_harness(api_key, agent_id, base_url):
    """AgentTestHarness in sandbox mode for integration tests."""
    h = AgentTestHarness(
        api_key=api_key,
        agent_id=agent_id,
        base_url=base_url,
    )
    h.use_sandbox()
    return h


@pytest.fixture
def mock_session():
    """Pre-configured requests.Session mock for unit tests."""
    session = MagicMock()
    response = MagicMock()
    response.status_code = 200
    response.json.return_value = {"status": "success"}
    response.raise_for_status.return_value = None
    session.post.return_value = response
    return session


@pytest.fixture
def mock_response():
    """Factory fixture for creating MockResponse objects."""
    def _make(tool, data=None, status="success", error_code=None):
        return MockResponse(
            tool=tool,
            status=status,
            data=data or {},
            error_code=error_code,
        )
    return _make


# ── Per-class fixtures for isolated test suites ─────────────

class AgentFixtures:
    """Mixin providing standard mocks for agent commerce tests."""

    @pytest.fixture(autouse=True)
    def setup_agent_mocks(self, harness):
        """Pre-register common mocks for every test in the class."""
        self.harness = harness
        harness.mock_tool("get_balance", {"balance": "500.00", "currency": "USD"})
        harness.mock_tool("create_wallet", {"wallet_id": "w-test-001", "status": "active"})
        harness.mock_tool("register_agent", {"agent_id": harness.agent_id, "status": "registered"})
        harness.mock_tool("get_trust_score", {"agent_id": "any", "score": 0.85})
        harness.mock_tool("get_budget_status", {
            "daily_limit": "100.00",
            "spent_today": "25.00",
            "remaining": "75.00",
        })


class EscrowFixtures(AgentFixtures):
    """Extended fixtures for escrow-related tests."""

    @pytest.fixture(autouse=True)
    def setup_escrow_mocks(self, harness):
        """Add escrow mocks on top of agent mocks."""
        self.escrow_id = f"escrow-{uuid.uuid4().hex[:8]}"
        harness.mock_tool("create_escrow", {
            "escrow_id": self.escrow_id,
            "status": "funded",
            "amount": "50.00",
        })
        harness.mock_tool("release_escrow", {
            "escrow_id": self.escrow_id,
            "status": "released",
        })
        harness.mock_tool("cancel_escrow", {
            "escrow_id": self.escrow_id,
            "status": "cancelled",
        })
```

### Pattern: Deterministic Mocks vs. Sandbox Testing

The harness supports both modes. Use mocks for business logic tests (fast, no network, deterministic). Use sandbox for contract and integration tests (real API, real state, real latency).

```python
class TestBudgetGuardrails(AgentFixtures):
    """Layer 1: Test budget logic with deterministic mocks."""

    def test_blocks_escrow_when_over_budget(self, harness):
        """Escrow creation should be blocked when daily budget is exhausted."""
        harness.mock_tool("get_budget_status", {
            "daily_limit": "100.00",
            "spent_today": "99.00",
            "remaining": "1.00",
        })
        budget = harness.execute("get_budget_status", {})
        remaining = float(budget["remaining"])
        escrow_amount = 50.00

        # Business logic: do not create escrow if amount > remaining
        assert escrow_amount > remaining
        harness.assert_tool_not_called("create_escrow")

    def test_allows_escrow_within_budget(self, harness):
        """Escrow creation should proceed when budget allows."""
        budget = harness.execute("get_budget_status", {})
        remaining = float(budget["remaining"])
        escrow_amount = 25.00

        assert escrow_amount <= remaining
        harness.execute("create_escrow", {
            "payer_agent_id": harness.agent_id,
            "payee_agent_id": "seller-001",
            "amount": str(escrow_amount),
        })
        harness.assert_tool_called("create_escrow", times=1)


@pytest.mark.sandbox
class TestBudgetGuardrailsSandbox:
    """Layer 2: Verify budget tools against live sandbox."""

    def test_budget_cap_enforced(self, sandbox_harness):
        """Sandbox should reject escrows exceeding the budget cap."""
        h = sandbox_harness
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "100.00"})
        h.execute("set_budget_cap", {
            "agent_id": h.agent_id,
            "daily_limit": "10.00",
        })
        # This should fail because escrow exceeds daily limit
        result = h.execute("create_escrow", {
            "payer_agent_id": h.agent_id,
            "payee_agent_id": "seller-test",
            "amount": "50.00",
        })
        # Gateway enforces budget at the tool level
        assert result.get("status") in ("error", "rejected")
```

---

## Chapter 2: Tool-Level Testing Patterns

### The Tool Contract Test

Every GreenHelix tool has an implicit contract: it accepts a specific input schema, returns a specific output shape, and produces documented error codes for invalid input. A tool contract test verifies all three. When the gateway updates an API version or adds a required field, your contract tests break before your production code does.

```python
class ToolContract:
    """Defines the expected contract for a GreenHelix tool.

    Used by contract tests to verify schema, output shape,
    and error behavior against the sandbox.
    """

    def __init__(
        self,
        tool: str,
        required_fields: list[str],
        output_fields: list[str],
        error_cases: dict[str, dict],
    ):
        self.tool = tool
        self.required_fields = required_fields
        self.output_fields = output_fields
        self.error_cases = error_cases  # {case_name: {input: ..., expected_error: ...}}


# ── Contract definitions for core tools ────────────────────

BILLING_CONTRACTS = {
    "get_balance": ToolContract(
        tool="get_balance",
        required_fields=[],
        output_fields=["balance", "currency"],
        error_cases={
            "no_wallet": {
                "input": {},
                "expected_error": "wallet_not_found",
            },
        },
    ),
    "deposit": ToolContract(
        tool="deposit",
        required_fields=["amount"],
        output_fields=["balance", "transaction_id"],
        error_cases={
            "negative_amount": {
                "input": {"amount": "-10.00"},
                "expected_error": "invalid_amount",
            },
            "zero_amount": {
                "input": {"amount": "0"},
                "expected_error": "invalid_amount",
            },
        },
    ),
    "set_budget_cap": ToolContract(
        tool="set_budget_cap",
        required_fields=["agent_id", "daily_limit"],
        output_fields=["agent_id", "daily_limit"],
        error_cases={
            "negative_limit": {
                "input": {"agent_id": "test", "daily_limit": "-50.00"},
                "expected_error": "invalid_amount",
            },
        },
    ),
}

PAYMENT_CONTRACTS = {
    "create_escrow": ToolContract(
        tool="create_escrow",
        required_fields=["payer_agent_id", "payee_agent_id", "amount"],
        output_fields=["escrow_id", "status", "amount"],
        error_cases={
            "insufficient_funds": {
                "input": {
                    "payer_agent_id": "buyer",
                    "payee_agent_id": "seller",
                    "amount": "999999.00",
                },
                "expected_error": "insufficient_funds",
            },
            "self_escrow": {
                "input": {
                    "payer_agent_id": "same-agent",
                    "payee_agent_id": "same-agent",
                    "amount": "10.00",
                },
                "expected_error": "invalid_escrow",
            },
        },
    ),
    "release_escrow": ToolContract(
        tool="release_escrow",
        required_fields=["escrow_id"],
        output_fields=["escrow_id", "status"],
        error_cases={
            "nonexistent": {
                "input": {"escrow_id": "escrow-does-not-exist"},
                "expected_error": "escrow_not_found",
            },
        },
    ),
}

IDENTITY_CONTRACTS = {
    "register_agent": ToolContract(
        tool="register_agent",
        required_fields=["agent_id", "public_key", "name"],
        output_fields=["agent_id", "status"],
        error_cases={
            "missing_key": {
                "input": {"agent_id": "test", "name": "Test"},
                "expected_error": "missing_field",
            },
        },
    ),
    "get_trust_score": ToolContract(
        tool="get_trust_score",
        required_fields=["agent_id"],
        output_fields=["agent_id", "score"],
        error_cases={
            "nonexistent_agent": {
                "input": {"agent_id": "agent-that-does-not-exist-xyz"},
                "expected_error": "agent_not_found",
            },
        },
    ),
}

MARKETPLACE_CONTRACTS = {
    "register_service": ToolContract(
        tool="register_service",
        required_fields=["name", "description", "endpoint", "price", "tags", "category"],
        output_fields=["service_id"],
        error_cases={
            "missing_name": {
                "input": {
                    "description": "test",
                    "endpoint": "agent://test",
                    "price": 10.0,
                    "tags": [],
                    "category": "test",
                },
                "expected_error": "missing_field",
            },
        },
    ),
    "search_services": ToolContract(
        tool="search_services",
        required_fields=["query"],
        output_fields=["services"],
        error_cases={
            "empty_query": {
                "input": {"query": ""},
                "expected_error": "invalid_query",
            },
        },
    ),
}
```

### Running Contract Tests

```python
@pytest.mark.sandbox
class TestBillingContracts:
    """Layer 2: Verify billing tool contracts against sandbox."""

    @pytest.fixture(autouse=True)
    def setup_wallet(self, sandbox_harness):
        self.harness = sandbox_harness
        self.harness.execute("create_wallet", {})
        self.harness.execute("deposit", {"amount": "100.00"})

    @pytest.mark.parametrize("tool_name", BILLING_CONTRACTS.keys())
    def test_output_shape(self, tool_name):
        """Every billing tool returns expected output fields."""
        contract = BILLING_CONTRACTS[tool_name]
        # Build minimal valid input
        valid_input = {}
        if tool_name == "deposit":
            valid_input = {"amount": "10.00"}
        elif tool_name == "set_budget_cap":
            valid_input = {
                "agent_id": self.harness.agent_id,
                "daily_limit": "50.00",
            }
        result = self.harness.execute(tool_name, valid_input)
        for expected_field in contract.output_fields:
            assert expected_field in result, (
                f"Tool '{tool_name}' missing output field '{expected_field}'. "
                f"Got: {list(result.keys())}"
            )

    @pytest.mark.parametrize("tool_name", BILLING_CONTRACTS.keys())
    def test_error_cases(self, tool_name):
        """Every billing tool returns correct error codes for invalid input."""
        contract = BILLING_CONTRACTS[tool_name]
        for case_name, case in contract.error_cases.items():
            result = self.harness.execute(tool_name, case["input"])
            assert result.get("status") == "error" or "error" in result, (
                f"Tool '{tool_name}' case '{case_name}' should have failed. "
                f"Got: {result}"
            )


@pytest.mark.sandbox
class TestPaymentContracts:
    """Layer 2: Verify payment tool contracts against sandbox."""

    @pytest.fixture(autouse=True)
    def setup_accounts(self, sandbox_harness, buyer_id, seller_id):
        self.harness = sandbox_harness
        self.buyer_id = buyer_id
        self.seller_id = seller_id

    @pytest.mark.parametrize("tool_name", PAYMENT_CONTRACTS.keys())
    def test_error_cases(self, tool_name):
        """Every payment tool returns correct errors for invalid input."""
        contract = PAYMENT_CONTRACTS[tool_name]
        for case_name, case in contract.error_cases.items():
            result = self.harness.execute(tool_name, case["input"])
            assert result.get("status") == "error" or "error" in result, (
                f"Payment tool '{tool_name}' case '{case_name}' did not fail"
            )
```

### Idempotency Testing for Payment Tools

Payment tools must be idempotent. Calling `create_escrow` twice with the same parameters should not create two escrows. Calling `release_escrow` twice should not transfer funds twice. This pattern tests idempotency by submitting duplicate requests and verifying financial consistency (P1, P7).

```python
@pytest.mark.sandbox
class TestPaymentIdempotency:
    """Layer 2: Verify payment tools handle duplicate calls safely."""

    def test_duplicate_escrow_creation(self, sandbox_harness, buyer_id, seller_id):
        """Creating the same escrow twice should return the same escrow_id."""
        h = sandbox_harness
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "200.00"})

        escrow_params = {
            "payer_agent_id": buyer_id,
            "payee_agent_id": seller_id,
            "amount": "50.00",
            "description": "Idempotency test escrow",
            "idempotency_key": f"idem-{uuid.uuid4().hex[:8]}",
        }

        result_1 = h.execute("create_escrow", escrow_params)
        result_2 = h.execute("create_escrow", escrow_params)

        # Same idempotency key should return same escrow
        assert result_1["escrow_id"] == result_2["escrow_id"]

        # Balance should only be debited once
        balance = h.execute("get_balance", {})
        assert float(balance["balance"]) == 150.00

    def test_duplicate_release(self, sandbox_harness):
        """Releasing the same escrow twice should not double-pay."""
        h = sandbox_harness
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "100.00"})

        escrow = h.execute("create_escrow", {
            "payer_agent_id": h.agent_id,
            "payee_agent_id": "seller-test",
            "amount": "25.00",
        })
        escrow_id = escrow["escrow_id"]

        release_1 = h.execute("release_escrow", {"escrow_id": escrow_id})
        release_2 = h.execute("release_escrow", {"escrow_id": escrow_id})

        # Second release should be a no-op or return already_released
        assert release_1.get("status") == "released"
        assert release_2.get("status") in ("released", "already_released")
```

### Permission Boundary Testing

Agents should only be able to operate on their own resources. A buyer should not release an escrow created by a different buyer. A seller should not cancel an escrow that is not addressed to them. Permission boundary tests verify these invariants (P7).

```python
@pytest.mark.sandbox
class TestPermissionBoundaries:
    """Layer 2: Verify agents cannot access other agents' resources."""

    def test_cannot_release_others_escrow(self, sandbox_harness):
        """Agent A cannot release an escrow created by Agent B."""
        h = sandbox_harness
        # Agent A creates an escrow
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "100.00"})
        escrow = h.execute("create_escrow", {
            "payer_agent_id": h.agent_id,
            "payee_agent_id": "seller-x",
            "amount": "10.00",
        })

        # Agent B (different harness/session) tries to release it
        other = AgentTestHarness(
            api_key=h.api_key,
            agent_id="attacker-agent",
            base_url=h.base_url,
        )
        other.use_sandbox()
        result = other.execute("release_escrow", {
            "escrow_id": escrow["escrow_id"],
        })
        assert result.get("status") == "error"

    def test_cannot_read_others_balance(self, sandbox_harness):
        """Agent A cannot read Agent B's wallet balance."""
        h = sandbox_harness
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "100.00"})

        other = AgentTestHarness(
            api_key=h.api_key,
            agent_id="other-agent",
            base_url=h.base_url,
        )
        other.use_sandbox()
        result = other.execute("get_balance", {})
        # Should return the other agent's balance (0), not our 100
        balance = float(result.get("balance", 0))
        assert balance != 100.00

    def test_cannot_cancel_others_escrow(self, sandbox_harness):
        """Seller cannot cancel an escrow -- only the buyer can."""
        h = sandbox_harness
        h.execute("create_wallet", {})
        h.execute("deposit", {"amount": "50.00"})
        escrow = h.execute("create_escrow", {
            "payer_agent_id": h.agent_id,
            "payee_agent_id": "seller-y",
            "amount": "10.00",
        })

        seller = AgentTestHarness(
            api_key=h.api_key,
            agent_id="seller-y",
            base_url=h.base_url,
        )
        seller.use_sandbox()
        result = seller.execute("cancel_escrow", {
            "escrow_id": escrow["escrow_id"],
        })
        assert result.get("status") == "error"
```

---

## Chapter 3: Workflow & Integration Testing

### The Saga Test Pattern

Agent commerce workflows are sagas: multi-step operations where each step has a compensating action. If step 3 fails, steps 1 and 2 must be rolled back. The saga test pattern verifies both the happy path and every possible failure point.

```python
class MarketplaceSaga:
    """Implements the complete marketplace listing workflow as a testable saga.

    Steps:
        1. Seller registers service on marketplace
        2. Buyer discovers service via search
        3. Buyer checks seller trust score
        4. Buyer creates escrow
        5. Seller performs work (simulated)
        6. Buyer releases escrow
        7. Buyer rates service
        8. Settlement completes

    Compensating actions:
        Step 4 fails → no cleanup needed (funds not locked)
        Step 5 fails → cancel escrow (return funds to buyer)
        Step 6 fails → open dispute
    """

    def __init__(self, harness: AgentTestHarness, buyer_id: str, seller_id: str):
        self.harness = harness
        self.buyer_id = buyer_id
        self.seller_id = seller_id
        self.state = {"step": 0, "completed_steps": []}

    def run(self) -> dict:
        """Execute the full saga, rolling back on failure."""
        try:
            # Step 1: Register service
            service = self.harness.execute("register_service", {
                "name": "Test Summarization Service",
                "description": "Summarizes documents for testing",
                "endpoint": f"agent://{self.seller_id}",
                "price": 25.00,
                "tags": ["test", "summarization"],
                "category": "data-processing",
            })
            self.state["service_id"] = service["service_id"]
            self.state["completed_steps"].append("register_service")

            # Step 2: Discover service
            results = self.harness.execute("search_services", {
                "query": "test summarization",
            })
            assert len(results.get("services", [])) > 0
            self.state["completed_steps"].append("discover_service")

            # Step 3: Trust check
            trust = self.harness.execute("get_trust_score", {
                "agent_id": self.seller_id,
            })
            if trust.get("score", 0) < 0.5:
                return {"status": "aborted", "reason": "low_trust"}
            self.state["completed_steps"].append("trust_check")

            # Step 4: Create escrow
            escrow = self.harness.execute("create_escrow", {
                "payer_agent_id": self.buyer_id,
                "payee_agent_id": self.seller_id,
                "amount": "25.00",
                "description": "Saga test escrow",
            })
            self.state["escrow_id"] = escrow["escrow_id"]
            self.state["completed_steps"].append("create_escrow")

            # Step 5: Simulate work (in real tests, call seller endpoint)
            work_result = {"quality": 0.95, "documents_processed": 500}
            self.state["completed_steps"].append("work_completed")

            # Step 6: Release escrow
            release = self.harness.execute("release_escrow", {
                "escrow_id": escrow["escrow_id"],
            })
            self.state["completed_steps"].append("release_escrow")

            # Step 7: Rate service
            self.harness.execute("rate_service", {
                "service_id": service["service_id"],
                "rating": 5,
            })
            self.state["completed_steps"].append("rate_service")

            return {"status": "completed", "state": self.state}

        except Exception as e:
            return self._compensate(str(e))

    def _compensate(self, error: str) -> dict:
        """Roll back completed steps on failure."""
        if "create_escrow" in self.state["completed_steps"]:
            escrow_id = self.state.get("escrow_id")
            if escrow_id and "release_escrow" not in self.state["completed_steps"]:
                self.harness.execute("cancel_escrow", {
                    "escrow_id": escrow_id,
                })
                self.state["completed_steps"].append("compensate:cancel_escrow")

        return {
            "status": "rolled_back",
            "error": error,
            "state": self.state,
        }
```

### Testing the Saga

```python
class TestMarketplaceSaga(EscrowFixtures):
    """Layer 3: Full marketplace workflow with rollback verification."""

    def test_happy_path(self, harness, buyer_id, seller_id):
        """Complete saga executes all 7 steps."""
        harness.mock_tool("register_service", {
            "service_id": "svc-test-001",
        })
        harness.mock_tool("search_services", {
            "services": [{"name": "Test Service", "agent_id": seller_id}],
        })
        harness.mock_tool("rate_service", {"status": "rated"})

        saga = MarketplaceSaga(harness, buyer_id, seller_id)
        result = saga.run()

        assert result["status"] == "completed"
        assert len(result["state"]["completed_steps"]) == 7
        harness.assert_call_order([
            "register_service",
            "search_services",
            "get_trust_score",
            "create_escrow",
            "release_escrow",
            "rate_service",
        ])

    def test_rollback_on_escrow_failure(self, harness, buyer_id, seller_id):
        """Failed escrow creation does not leave orphaned state."""
        harness.mock_tool("register_service", {"service_id": "svc-test-002"})
        harness.mock_tool("search_services", {
            "services": [{"name": "Test", "agent_id": seller_id}],
        })
        harness.mock_tool_error(
            "create_escrow", "insufficient_funds", "Not enough balance"
        )

        saga = MarketplaceSaga(harness, buyer_id, seller_id)
        result = saga.run()

        assert result["status"] == "rolled_back"
        assert "create_escrow" not in result["state"]["completed_steps"]
        harness.assert_tool_not_called("release_escrow")

    def test_rollback_cancels_escrow_on_work_failure(self, harness, buyer_id, seller_id):
        """Failed work step triggers escrow cancellation."""
        harness.mock_tool("register_service", {"service_id": "svc-test-003"})
        harness.mock_tool("search_services", {
            "services": [{"name": "Test", "agent_id": seller_id}],
        })

        saga = MarketplaceSaga(harness, buyer_id, seller_id)
        # Simulate work failure by injecting error after escrow creation
        original_execute = harness.execute

        call_count = {"n": 0}
        def failing_execute(tool, input_data):
            call_count["n"] += 1
            if tool == "release_escrow":
                raise RuntimeError("Simulated work verification failure")
            return original_execute(tool, input_data)

        harness.execute = failing_execute
        result = saga.run()

        assert result["status"] == "rolled_back"
        assert "compensate:cancel_escrow" in result["state"]["completed_steps"]
```

### Subscription Lifecycle Testing

Subscriptions are stateful workflows: create, renew, pause, cancel. Each transition must be tested, including edge cases like renewal with insufficient balance (P2, P6).

```python
class TestSubscriptionLifecycle(AgentFixtures):
    """Layer 3: Verify subscription state transitions."""

    def test_full_lifecycle(self, harness):
        """Create → renew → cancel lifecycle."""
        sub_id = f"sub-{uuid.uuid4().hex[:8]}"
        harness.mock_tool("create_subscription", {
            "subscription_id": sub_id,
            "status": "active",
            "next_payment_date": "2026-05-06",
        })
        harness.mock_

…(truncated)
