# Greenhelix Trading Bot Risk Service

> Trading Bot Risk-as-a-Service: Real-Time Portfolio Risk Monitoring for Multi-Exchange Operations. Build a cross-exchange, cross-strategy real-time portfolio risk monitoring system with webhooks, event bus, and SLA compliance enforcement. Covers drawdown alerts, correlation monitoring, liquidation proximity, circuit breakers, and production deployment.

- Skill: `lord1egypt/greenhelix-trading-bot-risk-service` (Agent Skill)
- Install (CLI): `npx skillmds@latest add lord1egypt/greenhelix-trading-bot-risk-service`
- Raw SKILL.md: https://api.skillmd.com/api/skills/lord1egypt/greenhelix-trading-bot-risk-service/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- License: MIT
- Author: Lord1Egypt (https://skillmd.com/u/lord1egypt)
- Updated: 2026-09-08
- Page: https://skillmd.com/skills/lord1egypt/greenhelix-trading-bot-risk-service

---

# Trading Bot Risk-as-a-Service: Real-Time Portfolio Risk Monitoring for Multi-Exchange Operations

> **Notice**: This is an educational guide with illustrative code examples.
> It does not execute code or install dependencies.
> All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which
> provides 500 free credits — no API key required to get started.
>
> **Referenced credentials** (you supply these in your own environment):
> - `AGENT_SIGNING_KEY`: Cryptographic signing key for agent identity (Ed25519 key pair for request signing)


If you run trading bots across multiple exchanges, you have a visibility problem. Binance shows you Binance positions. Coinbase shows you Coinbase positions. Kraken shows you Kraken positions. Nobody shows you the aggregate. Nobody tells you that your mean-reversion strategy on Binance and your momentum strategy on Coinbase have become 0.94 correlated over the last four hours -- meaning a single market move will hit both simultaneously. Nobody warns you that your combined leverage across three exchanges has crept from 3x to 7x because each exchange only sees its own slice. The May 2025 cascade liquidations on BitMEX made this concrete: bots running isolated risk checks per exchange missed the aggregate exposure building across venues, and when BTC dropped 12% in ninety minutes, the cross-exchange margin calls arrived simultaneously. Operators who had $200K in aggregate equity across four exchanges discovered they were effectively running 11x leverage when positions were netted. The liquidation cascade took 47 seconds from first margin call to full wipeout. This guide builds a complete, production-grade risk monitoring system that sits above your exchange connections and provides a unified view of portfolio risk. It uses GreenHelix's event bus for real-time data aggregation, webhooks for alert delivery, and SLA compliance monitoring for operational health tracking. You will build drawdown trackers, correlation monitors, liquidation proximity engines, and circuit breakers -- all wired together through a single risk management layer that treats your entire multi-exchange operation as one portfolio.
1. [Risk Architecture Overview](#chapter-1-risk-architecture-overview)
2. [The RiskMonitor Class](#chapter-2-the-riskmonitor-class)

## What You'll Learn
- Chapter 1: Risk Architecture Overview
- Chapter 2: The RiskMonitor Class
- Chapter 3: Drawdown Monitoring and Alerts
- Chapter 4: Correlation Monitoring
- Chapter 5: Liquidation Proximity Engine
- Chapter 6: Circuit Breaker Implementation
- Chapter 7: Multi-Exchange Aggregation
- Next Steps
- What's Next

## Full Guide

# Trading Bot Risk-as-a-Service: Real-Time Portfolio Risk Monitoring for Multi-Exchange Operations

If you run trading bots across multiple exchanges, you have a visibility problem. Binance shows you Binance positions. Coinbase shows you Coinbase positions. Kraken shows you Kraken positions. Nobody shows you the aggregate. Nobody tells you that your mean-reversion strategy on Binance and your momentum strategy on Coinbase have become 0.94 correlated over the last four hours -- meaning a single market move will hit both simultaneously. Nobody warns you that your combined leverage across three exchanges has crept from 3x to 7x because each exchange only sees its own slice. The May 2025 cascade liquidations on BitMEX made this concrete: bots running isolated risk checks per exchange missed the aggregate exposure building across venues, and when BTC dropped 12% in ninety minutes, the cross-exchange margin calls arrived simultaneously. Operators who had $200K in aggregate equity across four exchanges discovered they were effectively running 11x leverage when positions were netted. The liquidation cascade took 47 seconds from first margin call to full wipeout. This guide builds a complete, production-grade risk monitoring system that sits above your exchange connections and provides a unified view of portfolio risk. It uses GreenHelix's event bus for real-time data aggregation, webhooks for alert delivery, and SLA compliance monitoring for operational health tracking. You will build drawdown trackers, correlation monitors, liquidation proximity engines, and circuit breakers -- all wired together through a single risk management layer that treats your entire multi-exchange operation as one portfolio.

---

## Table of Contents

1. [Risk Architecture Overview](#chapter-1-risk-architecture-overview)
2. [The RiskMonitor Class](#chapter-2-the-riskmonitor-class)
3. [Drawdown Monitoring and Alerts](#chapter-3-drawdown-monitoring-and-alerts)
4. [Correlation Monitoring](#chapter-4-correlation-monitoring)
5. [Liquidation Proximity Engine](#chapter-5-liquidation-proximity-engine)
6. [Circuit Breaker Implementation](#chapter-6-circuit-breaker-implementation)
7. [Multi-Exchange Aggregation](#chapter-7-multi-exchange-aggregation)
9. [What's Next](#whats-next)

---

## Chapter 1: Risk Architecture Overview

### Why Centralized Risk Monitoring Matters

Every professional trading desk has a risk management layer that sits above the execution layer. The risk system sees all positions, all strategies, all venues. It computes aggregate exposure, monitors correlation regimes, and has the authority to halt trading when thresholds are breached. Retail and semi-professional bot operators skip this layer because building it is hard and the exchanges do not provide the APIs to do it natively across venues. The result is that a $500K multi-exchange bot operation runs with less risk infrastructure than a single-desk prop trader at a regional firm.

The failure modes are specific and predictable:

- **Cross-exchange leverage accumulation**: Each exchange computes margin independently. A 3x position on Binance and a 3x position on Coinbase in correlated assets is effectively 6x aggregate leverage against your total equity -- but neither exchange reports it that way.
- **Correlation regime changes**: Two strategies that were uncorrelated during backtesting become highly correlated during market stress. The diversification benefit you assumed in your position sizing disappears exactly when you need it.
- **Cascading liquidations**: A liquidation on one exchange reduces your total equity, which increases your leverage ratio on other exchanges, which triggers further liquidations. This feedback loop operates faster than human reaction time.
- **Silent risk drift**: Without continuous monitoring, risk parameters drift over days and weeks. A portfolio that started at 2x aggregate leverage creeps to 5x through a series of individually reasonable position additions.

### Architecture

The system has four layers:

```
+------------------+    +------------------+    +------------------+
|  Binance Bot     |    |  Coinbase Bot    |    |  Kraken Bot      |
|  (execution)     |    |  (execution)     |    |  (execution)     |
+--------+---------+    +--------+---------+    +--------+---------+
         |                       |                       |
         |   publish_event       |   publish_event       |   publish_event
         v                       v                       v
+------------------------------------------------------------------------+
|                     GreenHelix Event Bus                                |
|  (real-time event ingestion, schema validation, ordering)              |
+--------+---------------------+-----------------------+-----------------+
         |                     |                       |
         |   get_events        |   webhooks            |   get_sla_compliance
         v                     v                       v
+------------------------------------------------------------------------+
|                     Risk Engine                                        |
|  +----------------+  +------------------+  +------------------------+  |
|  | DrawdownTracker|  | CorrelationMon   |  | LiquidationProximity   |  |
|  +----------------+  +------------------+  +------------------------+  |
|  +----------------+  +------------------+  +------------------------+  |
|  | CircuitBreaker |  | ExchangeAggr     |  | AlertManager           |  |
|  +----------------+  +------------------+  +------------------------+  |
+--------+---------------------------------------------------------------+
         |
         |   send_message / register_webhook
         v
+------------------------------------------------------------------------+
|                     Alert Delivery                                      |
|  Slack, PagerDuty, Telegram, Email, SMS                                |
+------------------------------------------------------------------------+
```

### GreenHelix Tools Used

This guide uses seven GreenHelix tools:

| Tool | Purpose |
|------|---------|
| `register_agent` | Register the risk monitoring agent |
| `register_webhook` | Set up alert delivery endpoints |
| `publish_event` | Bots publish position and trade events |
| `get_events` | Risk engine retrieves events for analysis |
| `get_sla_compliance` | Monitor risk engine uptime and latency |
| `submit_metrics` | Report risk metrics for observability |
| `send_message` | Deliver alerts to operators |

### Risk Metrics Hierarchy

Risk metrics are computed at four levels, each aggregating from the level below:

1. **Position-level**: Per-position P&L, unrealized P&L, distance to liquidation, margin utilization
2. **Strategy-level**: Strategy drawdown, strategy Sharpe ratio (rolling), strategy exposure, number of open positions
3. **Portfolio-level**: Aggregate drawdown, cross-strategy correlation matrix, net exposure by asset, aggregate leverage
4. **Fleet-level**: Total equity at risk, number of active strategies, circuit breaker status, SLA compliance score

Each level publishes events to the GreenHelix event bus, and higher levels consume events from lower levels. This creates a clean data flow where the risk engine never polls exchanges directly -- it only reads from the event bus.

### Event Types for Risk Monitoring

Define one event type per risk signal:

| Event Type | Trigger | Key Fields |
|---|---|---|
| `risk.portfolio_check` | Periodic risk scan completes | total_equity, aggregate_leverage, alerts |
| `risk.drawdown_alert` | Drawdown exceeds threshold | level, threshold_pct, current_pct, trigger_timeframe |
| `risk.correlation_matrix` | Correlation matrix updated | matrix, strategy_count |
| `risk.correlation_regime_change` | Correlation shift detected | avg_change, max_change |
| `risk.liquidation_proximity` | Liquidation check completes | positions, min_distance_pct |
| `risk.circuit_breaker_transition` | State change | from, to, reason, trip_count |
| `risk.circuit_breaker_config` | Triggers updated | triggers |
| `risk.aggregate_snapshot` | Cross-exchange aggregation | total_equity, aggregate_leverage, net_exposure |
| `risk.thresholds_updated` | Risk thresholds changed | thresholds |
| `risk.key_rotated` | Agent key rotation | new_public_key |

Each event is signed with the risk agent's Ed25519 private key. The signature covers the canonical JSON serialization of the payload (keys sorted, no whitespace), ensuring that risk events cannot be fabricated or tampered with after the fact. This matters when you need to prove to auditors or investors that a circuit breaker tripped at a specific time for a specific reason.

### Verifying the API Connection

Before building the full risk engine, verify that your GreenHelix API credentials work:

```bash
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "register_agent",
    "input": {
      "agent_id": "risk-monitor-prod-01",
      "public_key": "'"$PUBLIC_KEY_B64"'",
      "name": "Portfolio Risk Monitor"
    }
  }'
```

```python
import requests

API_BASE = "https://api.greenhelix.net/v1"
API_KEY = "your-api-key"

response = requests.post(
    f"{API_BASE}/v1",
    headers={
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
    },
    json={
        "tool": "register_agent",
        "input": {
            "agent_id": "risk-monitor-prod-01",
            "public_key": "your-public-key-base64",
            "name": "Portfolio Risk Monitor",
        },
    },
)
print(response.json())
```

If both return a success response with your agent ID, the foundation is in place.

---

## Chapter 2: The RiskMonitor Class

### Core Infrastructure

The `RiskMonitor` class is the foundation for all risk monitoring in this guide. Every subsequent chapter builds on it.

```bash
# Generate an Ed25519 keypair for the risk monitor agent
python3 -c "
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
from cryptography.hazmat.primitives import serialization
import base64

private_key = Ed25519PrivateKey.generate()
public_key = private_key.public_key()

priv_bytes = private_key.private_bytes(
    encoding=serialization.Encoding.Raw,
    format=serialization.PrivateFormat.Raw,
    encryption_algorithm=serialization.NoEncryption()
)
pub_bytes = public_key.public_bytes(
    encoding=serialization.Encoding.Raw,
    format=serialization.PublicFormat.Raw
)

print(f'PRIVATE_KEY_B64={base64.b64encode(priv_bytes).decode()}')
print(f'PUBLIC_KEY_B64={base64.b64encode(pub_bytes).decode()}')
"
```

```bash
# Register the risk monitor agent
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "register_agent",
    "input": {
      "agent_id": "risk-monitor-prod-01",
      "public_key": "'"$PUBLIC_KEY_B64"'",
      "name": "Portfolio Risk Monitor"
    }
  }'
```

### Python Implementation

```python
import json
import time
import base64
import hashlib
from datetime import datetime, timezone
from typing import Dict, List, Optional, Any

import requests
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
from cryptography.hazmat.primitives import serialization


class RiskMonitor:
    """Cross-exchange portfolio risk monitoring using GreenHelix APIs."""

    def __init__(
        self,
        api_key: str,
        agent_id: str,
        private_key_b64: str,
        base_url: str = "https://api.greenhelix.net/v1",
    ):
        self.api_key = api_key
        self.agent_id = agent_id
        self.base_url = base_url
        self.session = requests.Session()
        self.session.headers.update({
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json",
        })

        # Load Ed25519 private key
        key_bytes = base64.b64decode(private_key_b64)
        self._private_key = Ed25519PrivateKey.from_private_bytes(key_bytes)
        self._public_key = self._private_key.public_key()

        # Risk thresholds (defaults)
        self.thresholds = {
            "max_drawdown_pct": 15.0,
            "max_correlation": 0.85,
            "max_leverage": 5.0,
            "min_margin_ratio": 0.20,
            "max_single_position_pct": 25.0,
        }

    def _execute(self, tool: str, input_data: dict) -> dict:
        """Execute a GreenHelix tool."""
        response = self.session.post(
            f"{self.base_url}/v1",
            json={"tool": tool, "input": input_data},
        )
        response.raise_for_status()
        return response.json()

    def _sign(self, payload: dict) -> str:
        """Sign a payload with Ed25519."""
        canonical = json.dumps(payload, sort_keys=True, separators=(",", ":"))
        signature = self._private_key.sign(canonical.encode("utf-8"))
        return base64.b64encode(signature).decode()

    def register_risk_agent(self) -> dict:
        """Register this risk monitor as a GreenHelix agent."""
        pub_bytes = self._public_key.public_bytes(
            encoding=serialization.Encoding.Raw,
            format=serialization.PublicFormat.Raw,
        )
        return self._execute("register_agent", {
            "agent_id": self.agent_id,
            "public_key": base64.b64encode(pub_bytes).decode(),
            "name": f"Risk Monitor ({self.agent_id})",
        })

    def configure_thresholds(
        self,
        max_drawdown_pct: float = 15.0,
        max_correlation: float = 0.85,
        max_leverage: float = 5.0,
        min_margin_ratio: float = 0.20,
        max_single_position_pct: float = 25.0,
    ) -> dict:
        """Set risk thresholds and publish them as a configuration event."""
        self.thresholds = {
            "max_drawdown_pct": max_drawdown_pct,
            "max_correlation": max_correlation,
            "max_leverage": max_leverage,
            "min_margin_ratio": min_margin_ratio,
            "max_single_position_pct": max_single_position_pct,
        }

        event_payload = {
            "agent_id": self.agent_id,
            "thresholds": self.thresholds,
            "timestamp": datetime.now(timezone.utc).isoformat(),
        }
        event_payload["signature"] = self._sign(self.thresholds)

        return self._execute("publish_event", {
            "event_type": "risk.thresholds_updated",
            "payload": event_payload,
        })

    def check_portfolio_risk(self, positions: List[dict]) -> dict:
        """
        Run all risk checks against current positions.

        Returns a dict with risk metrics and any triggered alerts.
        """
        total_equity = sum(p.get("equity", 0) for p in positions)
        total_notional = sum(abs(p.get("notional", 0)) for p in positions)
        leverage = total_notional / total_equity if total_equity > 0 else float("inf")

        alerts = []

        # Check leverage
        if leverage > self.thresholds["max_leverage"]:
            alerts.append({
                "type": "leverage_breach",
                "current": round(leverage, 2),
                "threshold": self.thresholds["max_leverage"],
                "severity": "critical",
            })

        # Check concentration
        for pos in positions:
            if total_notional > 0:
                concentration = abs(pos.get("notional", 0)) / total_notional * 100
                if concentration > self.thresholds["max_single_position_pct"]:
                    alerts.append({
                        "type": "concentration_breach",
                        "position": pos.get("symbol", "unknown"),
                        "current_pct": round(concentration, 2),
                        "threshold": self.thresholds["max_single_position_pct"],
                        "severity": "warning",
                    })

        risk_summary = {
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "total_equity": total_equity,
            "total_notional": total_notional,
            "aggregate_leverage": round(leverage, 2),
            "position_count": len(positions),
            "alerts": alerts,
            "status": "critical" if any(a["severity"] == "critical" for a in alerts)
                      else "warning" if alerts
                      else "healthy",
        }

        # Publish risk check result as event
        event_payload = {
            "agent_id": self.agent_id,
            **risk_summary,
        }
        event_payload["signature"] = self._sign(risk_summary)

        self._execute("publish_event", {
            "event_type": "risk.portfolio_check",
            "payload": event_payload,
        })

        return risk_summary
```

### Configuring Thresholds via curl

```bash
# Publish a threshold configuration event
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "publish_event",
    "input": {
      "event_type": "risk.thresholds_updated",
      "payload": {
        "agent_id": "risk-monitor-prod-01",
        "thresholds": {
          "max_drawdown_pct": 15.0,
          "max_correlation": 0.85,
          "max_leverage": 5.0,
          "min_margin_ratio": 0.20,
          "max_single_position_pct": 25.0
        },
        "timestamp": "2026-04-07T12:00:00Z",
        "signature": "base64-ed25519-signature"
      }
    }
  }'
```

### Usage Example

```python
monitor = RiskMonitor(
    api_key="your-api-key",
    agent_id="risk-monitor-prod-01",
    private_key_b64="your-private-key-base64",
)

# Register and configure
monitor.register_risk_agent()
monitor.configure_thresholds(
    max_drawdown_pct=12.0,
    max_leverage=4.0,
    min_margin_ratio=0.25,
)

# Check risk against current positions
positions = [
    {"symbol": "BTC-USD", "notional": 50000, "equity": 15000, "exchange": "binance"},
    {"symbol": "ETH-USD", "notional": 30000, "equity": 10000, "exchange": "coinbase"},
    {"symbol": "SOL-USD", "notional": 20000, "equity": 8000, "exchange": "kraken"},
]

result = monitor.check_portfolio_risk(positions)
print(f"Status: {result['status']}, Leverage: {result['aggregate_leverage']}x")
```

The `RiskMonitor` class provides the transport layer and basic risk checks. The following chapters build specialized risk engines that integrate with it.

---

## Chapter 3: Drawdown Monitoring and Alerts

### Why Drawdown Is the Primary Risk Metric

Drawdown measures peak-to-trough decline. Unlike volatility (which is symmetric) or VaR (which is model-dependent), drawdown directly answers the question operators care about: "How much have I lost from my best point?" A 15% drawdown means your portfolio is worth 85% of its highest recorded value. Drawdown is also the metric most directly tied to survival -- a 50% drawdown requires a 100% gain to recover, and a 75% drawdown requires a 300% gain. Bot operators who monitor drawdown at multiple timeframes catch problems before they become existential.

### The DrawdownTracker Class

```python
import time
from collections import defaultdict
from datetime import datetime, timezone
from typing import Dict, List, Optional, Tuple


class DrawdownTracker:
    """Tracks drawdown across multiple timeframes with alert escalation."""

    # Timeframe windows in seconds
    TIMEFRAMES = {
        "1h": 3600,
        "4h": 14400,
        "24h": 86400,
        "7d": 604800,
    }

    # Alert escalation levels
    ALERT_LEVELS = {
        "warning": 5.0,
        "critical": 8.0,
        "emergency": 12.0,
        "circuit_breaker": 15.0,
    }

    def __init__(self, risk_monitor: "RiskMonitor"):
        self.risk_monitor = risk_monitor
        self._equity_history: List[Tuple[float, float]] = []  # (timestamp, equity)
        self._peaks: Dict[str, float] = {}  # timeframe -> peak equity
        self._current_drawdowns: Dict[str, float] = {}
        self._alert_history: List[dict] = []

    def update(self, equity: float, timestamp: Optional[float] = None) -> dict:
        """
        Record a new equity value and compute drawdowns across all timeframes.

        Returns the current drawdown state.
        """
        ts = timestamp or time.time()
        self._equity_history.append((ts, equity))

        # Prune history older than 7 days
        cutoff = ts - self.TIMEFRAMES["7d"]
        self._equity_history = [
            (t, e) for t, e in self._equity_history if t >= cutoff
        ]

        drawdowns = {}
        for tf_name, tf_seconds in self.TIMEFRAMES.items():
            tf_cutoff = ts - tf_seconds
            tf_values = [e for t, e in self._equity_history if t >= tf_cutoff]

            if not tf_values:
                drawdowns[tf_name] = 0.0
                continue

            peak = max(tf_values)
            self._peaks[tf_name] = peak
            current_dd = ((peak - equity) / peak) * 100 if peak > 0 else 0.0
            drawdowns[tf_name] = round(current_dd, 4)

        self._current_drawdowns = drawdowns
        return drawdowns

    def check_alerts(self) -> List[dict]:
        """
        Check current drawdowns against alert thresholds.

        Returns a list of triggered alerts, sorted by severity.
        """
        alerts = []
        worst_dd = max(self._current_drawdowns.values()) if self._current_drawdowns else 0.0

        for level_name, level_threshold in sorted(
            self.ALERT_LEVELS.items(), key=lambda x: x[1]
        ):
            if worst_dd >= level_threshold:
                # Find which timeframe triggered
                trigger_tf = max(
                    self._current_drawdowns, key=self._current_drawdowns.get
                )

                alert = {
                    "level": level_name,
                    "threshold_pct": level_threshold,
                    "current_pct": round(worst_dd, 4),
                    "trigger_timeframe": trigger_tf,
                    "timestamp": datetime.now(timezone.utc).isoformat(),
                    "peak_equity": self._peaks.get(trigger_tf, 0),
                }
                alerts.append(alert)

        if alerts:
            self._alert_history.extend(alerts)
            # Publish the most severe alert
            worst_alert = alerts[-1]
            payload = {
                "agent_id": self.risk_monitor.agent_id,
                "alert_type": "drawdown",
                **worst_alert,
            }
            payload["signature"] = self.risk_monitor._sign(worst_alert)

            self.risk_monitor._execute("publish_event", {
                "event_type": "risk.drawdown_alert",
                "payload": payload,
            })

        return alerts

    def get_history(self, timeframe: str = "24h") -> List[dict]:
        """Get drawdown history for a specific timeframe."""
        tf_seconds = self.TIMEFRAMES.get(timeframe, 86400)
        cutoff = time.time() - tf_seconds
        return [
            {"timestamp": t, "equity": e}
            for t, e in self._equity_history
            if t >= cutoff
        ]
```

### Webhook-Based Alert Delivery

Register a webhook to receive drawdown alerts in real time:

```bash
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "register_webhook",
    "input": {
      "url": "https://your-ops-server.example.com/risk/drawdown",
      "event_types": [
        "risk.drawdown_alert",
        "risk.portfolio_check"
      ],
      "secret": "your-webhook-hmac-secret"
    }
  }'
```

```python
# Register webhooks for drawdown alerts
monitor.register_risk_agent()

webhook_result = monitor._execute("register_webhook", {
    "url": "https://your-ops-server.example.com/risk/drawdown",
    "event_types": ["risk.drawdown_alert", "risk.portfolio_check"],
    "secret": "your-webhook-hmac-secret",
})
print(f"Webhook registered: {webhook_result}")
```

### Alert Escalation in Practice

The four-level escalation model maps directly to operational responses:

| Level | Threshold | Action |
|-------|-----------|--------|
| Warning | 5% drawdown | Log and notify via Slack |
| Critical | 8% drawdown | Page the on-call operator, reduce position sizes by 50% |
| Emergency | 12% drawdown | Halt new position opens, begin unwinding existing positions |
| Circuit Breaker | 15% drawdown | Cancel all open orders, close all positions, halt all strategies |

### Running the Drawdown Loop

```python
tracker = DrawdownTracker(risk_monitor=monitor)

# Simulate a monitoring loop (in production, this runs continuously)
def run_drawdown_monitor(monitor: RiskMonitor, tracker: DrawdownTracker):
    """Poll equity every 10 seconds and check drawdown alerts."""
    while True:
        # Fetch latest equity from your exchange aggregation layer
        events = monitor._execute("get_events", {
            "event_type": "exchange.equity_update",
            "limit": 1,
            "order": "desc",
        })

        if events.get("events"):
            equity = events["events"][0]["payload"].get("total_equity", 0)
            drawdowns = tracker.update(equity)
            alerts = tracker.check_alerts()

            if alerts:
                worst = alerts[-1]
                print(
                    f"[{worst['level'].upper()}] Drawdown: {worst['current_pct']}% "
                    f"(threshold: {worst['threshold_pct']}%, "
                    f"timeframe: {worst['trigger_timeframe']})"
                )

                # Send alert via messaging
                if worst["level"] in ("emergency", "circuit_breaker"):
                    monitor._execute("send_message", {
                        "to": "ops-team",
                        "subject": f"RISK ALERT: {worst['level'].upper()} drawdown",
                        "body": (
                            f"Drawdown: {worst['current_pct']}% "
                            f"(threshold: {worst['threshold_pct']}%)\n"
                            f"Timeframe: {worst['trigger_timeframe']}\n"
                            f"Peak equity: ${worst['peak_equity']:,.2f}"
                        ),
                    })

        time.sleep(10)
```

```bash
# Query drawdown alert history via curl
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "get_events",
    "input": {
      "event_type": "risk.drawdown_alert",
      "limit": 20,
      "order": "desc"
    }
  }'
```

### Drawdown Severity Classification

Not all drawdowns are equal. A 10% drawdown in a trending market during high volatility is qualitatively different from a 10% drawdown in a ranging market during low volatility. The tracker uses the drawdown across multiple timeframes to classify severity:

- **Shallow, slow drawdown** (24h drawdown > 1h drawdown): The portfolio is grinding lower. This is typically caused by a losing trend-following position. Monitor but do not panic -- the strategy may recover when the trend reverses.
- **Deep, fast drawdown** (1h drawdown > 24h drawdown): The portfolio is in rapid decline. This pattern is characteristic of liquidation cascades, flash crashes, or correlated position blowups. Immediate attention required.
- **Oscillating drawdown** (1h and 4h drawdowns similar, both < 24h): The portfolio is whipsawing. This pattern is common during high-volatility news events. Reduce position sizes rather than closing positions -- the whipsaw may resolve in either direction.

The drawdown tracker publishes this classification in the event payload, allowing downstream consumers (the circuit breaker, alert manager) to tailor their response to the type of drawdown, not just the magnitude.

The drawdown tracker is intentionally simple. It uses empirical peaks, not modeled ones. It does not try to predict future drawdown -- it measures what has already happened and triggers actions based on predefined thresholds. Predictive models add complexity without adding reliability under stress conditions.

---

## Chapter 4: Correlation Monitoring

### Why Correlation Kills Portfolios

Diversification is the only free lunch in finance -- until it isn't. Portfolio theory assumes that combining uncorrelated strategies reduces total risk. If Strategy A has a Sharpe of 1.5 and Strategy B has a Sharpe of 1.5, and they are uncorrelated, the combined portfolio has a Sharpe of roughly 2.1. But correlation is not static. During market stress, correlations spike. The 2020 COVID crash saw BTC-ETH correlation go from 0.6 to 0.95 in 48 hours. The 2022 LUNA collapse pushed correlations across all crypto assets above 0.9 for weeks. Strategies that appeared diversified during calm markets became a single concentrated bet during the exact conditions where diversification mattered most.

Monitoring correlation in real time -- and detecting regime changes -- lets you reduce exposure before the diversification illusion costs you capital.

### The CorrelationMonitor Class

```python
import math
from collections import deque
from datetime import datetime, timezone
from typing import Dict, List, Optional, Tuple


class CorrelationMonitor:
    """
    Rolling correlation computation between strategies
    with regime change detection.
    """

    def __init__(
        self,
        risk_monitor: "RiskMonitor",
        window_size: int = 100,
        regime_change_threshold: float = 0.3,
    ):
        self.risk_monitor = risk_monitor
        self.window_size = window_size
        self.regime_change_threshold = regime_change_threshold

        # Returns buffer per strategy
        self._returns: Dict[str, deque] = {}
        self._correlation_history: List[dict] = []
        self._previous_matrix: Optional[Dict[str, Dict[str, float]]] = None

    def add_returns(self, strategy_id: str, returns_value: float) -> None:
        """Add a return observation for a strategy."""
        if strategy_id not in self._returns:
            self._returns[strategy_id] = deque(maxlen=self.window_size)
        self._returns[strategy_id].append(returns_value)

    def _pearson(self, x: List[float], y: List[float]) -> float:
        """Compute Pearson correlation between two series."""
        n = min(len(x), len(y))
        if n < 10:
            return 0.0

        x, y = x[:n], y[:n]
        mean_x = sum(x) / n
        mean_y = sum(y) / n

        cov = sum((xi - mean_x) * (yi - mean_y) for xi, yi in zip(x, y))
        std_x = math.sqrt(sum((xi - mean_x) ** 2 for xi in x))
        std_y = math.sqrt(sum((yi - mean_y) ** 2 for yi in y))

        if std_x == 0 or std_y == 0:
            return 0.0

        return cov / (std_x * std_y)

    def compute_matrix(self) -> Dict[str, Dict[str, float]]:
        """Compute the full correlation matrix across all strategies."""
        strategies = sorted(self._returns.keys())
        matrix = {}

        for s1 in strategies:
            matrix[s1] = {}
            for s2 in strategies:
                if s1 == s2:
                    matrix[s1][s2] = 1.0
                else:
                    corr = self._pearson(
                        list(self._returns[s1]),
                        list(self._returns[s2]),
                    )
                    matrix[s1][s2] = round(corr, 4)

        # Publish the matrix as an event
        payload = {
            "agent_id": self.risk_monitor.agent_id,
            "matrix": matrix,
            "strategy_count": len(strategies),
            "timestamp": datetime.now(timezone.utc).isoformat(),
        }
        payload["signature"] = self.risk_monitor._sign({"matrix": matrix})

        self.risk_monitor._execute("publish_event", {
            "event_type": "risk.correlation_matrix",
            "payload": payload,
        })

        self._correlation_history.append({
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "matrix": matrix,
        })

        return matrix

    def detect_regime_change(self) -> Optional[dict]:
        """
        Detect if the correlation regime has changed significantly.

        Compares current matrix to previous matrix. A regime change is
        detected when the average absolute change in correlations exceeds
        the threshold.
        """
        current = self.compute_matrix()

        if self._previous_matrix is None:
            self._previous_matrix = current
            return None

        strategies = sorted(current.keys())
        changes = []

        for s1 in strategies:
            for s2 in strategies:
                if s1 >= s2:
                    continue
                prev = self._previous_matrix.get(s1, {}).get(s2, 0.0)
                curr = current.get(s1, {}).get(s2, 0.0)
                changes.append(abs(curr - prev))

        if not changes:
            self._previous_matrix = current
            return None

        avg_change = sum(changes) / len(changes)
        max_change = max(changes)

        self._previous_matrix = current

        if avg_change > self.regime_change_threshold:
            alert = {
                "type": "correlation_regime_change",
                "avg_change": round(avg_change, 4),
                "max_change": round(max_change, 4),
                "threshold": self.regime_change_threshold,
                "timestamp": datetime.now(timezone.utc).isoformat(),
                "matrix": current,
            }

            # Publish regime change event
            event_payload = {
                "agent_id": self.risk_monitor.agent_id,
                **alert,
            }
            event_payload["signature"] = self.risk_monitor._sign(alert)

            self.risk_monitor._execute("publish_event", {
                "event_type": "risk.correlation_regime_change",
                "payload": event_payload,
            })

            return alert

        return None
```

### Querying Correlation History via curl

```bash
# Get recent correlation matrix events
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "get_events",
    "input": {
      "event_type": "risk.correlation_matrix",
      "limit": 10,
      "order": "desc"
    }
  }'
```

```bash
# Get regime change alerts
curl -X POST https://sandbox.greenhelix.net/v1 \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "tool": "get_events",
    "input": {
      "event_type": "risk.correlation_regime_change",
      "limit": 5,
      "order": "desc"
    }
  }'
```

### Usage Example

```python
corr_monitor = CorrelationMonitor(
    risk_monitor=monitor,
    window_size=100,
    regime_change_threshold=0.3,
)

# Feed returns from each strategy (in production, these come from trade events)
import random
for i in range(120):
    corr_monitor.add_returns("mean-reversion-btc", random.gauss(0.001, 0.02))
    corr_monitor.add_returns("momentum-eth", random.gauss(0.0005, 0.025))
    corr_monitor.add_returns("arb-sol-perps", random.gauss(0.0008, 0.015))

matrix = corr_monitor.compute_matrix()
for s1, row in matrix.items():
    print(f"{s1}: {row}")

regime = corr_monitor.detect_regime_change()
if regime:
    print(f"REGIME CHANGE detected: avg change {regime['avg_change']}")
```

### Interpreting the Correlation Matrix

The correlation matrix is symmetric with 1.0 on the diagonal. Off-diagonal values range from -1.0 (perfectly inverse) to +1.0 (perfectly correlated). For risk management purposes:

- **Below 0.3**: Strategies are effectively uncorrelated. Diversification benefit is real.
- **0.3 to 0.6**: Moderate correlation. Some diversification benefit remains, but not full.
- **0.6 to 0.85**: High correlation. The strategies will draw down together in stressed markets. Position sizes should account for this.
- **Above 0.85**: The strategies are effectively the same bet. The circuit breaker should consider triggering if this persists.

### Practical Considerations for Correlation Monitoring

**Window size selection**: The `window_size` parameter controls how many return observations are used to compute correlation. Smaller windows (30-50) are more responsive to recent correlation changes but produce noisier estimates. Larger windows (100-200) are more stable but slower to detect regime changes. A common compromise is to run two correlation monitors in parallel -- one with a 50-observation window for early detection and one with a 150-observation window for confirmation.

**Return frequency**: The correlation monitor needs returns at a consistent frequency. If Strategy A reports returns every minute and Strategy B reports every 15 minutes, the correlation estimate will be unreliable. Normalize to the lowest common frequency across all strategies. For most crypto bot operations, 5-minute or 15-minute returns provide a good balance between responsiveness and stability.

**Handling missing data**: If a strategy stops reporting returns (e.g., because the bot crashed or the exchange is down), the correlation monitor should flag this rather than computing correlation with stale data. A strategy that has not reported in 30

…(truncated)
