Trading Bot Risk-as-a-Service: Real-Time Portfolio Risk Monitoring for Multi-Exchange Operations
Notice: This is an educational guide with illustrative code examples. It does not execute code or install dependencies. All examples use the GreenHelix sandbox (https://sandbox.greenhelix.net) which provides 500 free credits — no API key required to get started.
Referenced credentials (you supply these in your own environment):
AGENT_SIGNING_KEY: Cryptographic signing key for agent identity (Ed25519 key pair for request signing)
If you run trading bots across multiple exchanges, you have a visibility problem. Binance shows you Binance positions. Coinbase shows you Coinbase positions. Kraken shows you Kraken positions. Nobody shows you the aggregate. Nobody tells you that your mean-reversion strategy on Binance and your momentum strategy on Coinbase have become 0.94 correlated over the last four hours -- meaning a single market move will hit both simultaneously. Nobody warns you that your combined leverage across three exchanges has crept from 3x to 7x because each exchange only sees its own slice. The May 2025 cascade liquidations on BitMEX made this concrete: bots running isolated risk checks per exchange missed the aggregate exposure building across venues, and when BTC dropped 12% in ninety minutes, the cross-exchange margin calls arrived simultaneously. Operators who had $200K in aggregate equity across four exchanges discovered they were effectively running 11x leverage when positions were netted. The liquidation cascade took 47 seconds from first margin call to full wipeout. This guide builds a complete, production-grade risk monitoring system that sits above your exchange connections and provides a unified view of portfolio risk. It uses GreenHelix's event bus for real-time data aggregation, webhooks for alert delivery, and SLA compliance monitoring for operational health tracking. You will build drawdown trackers, correlation monitors, liquidation proximity engines, and circuit breakers -- all wired together through a single risk management layer that treats your entire multi-exchange operation as one portfolio.
What You'll Learn
- Chapter 1: Risk Architecture Overview
- Chapter 2: The RiskMonitor Class
- Chapter 3: Drawdown Monitoring and Alerts
- Chapter 4: Correlation Monitoring
- Chapter 5: Liquidation Proximity Engine
- Chapter 6: Circuit Breaker Implementation
- Chapter 7: Multi-Exchange Aggregation
- Next Steps
- What's Next
Full Guide
Trading Bot Risk-as-a-Service: Real-Time Portfolio Risk Monitoring for Multi-Exchange Operations
If you run trading bots across multiple exchanges, you have a visibility problem. Binance shows you Binance positions. Coinbase shows you Coinbase positions. Kraken shows you Kraken positions. Nobody shows you the aggregate. Nobody tells you that your mean-reversion strategy on Binance and your momentum strategy on Coinbase have become 0.94 correlated over the last four hours -- meaning a single market move will hit both simultaneously. Nobody warns you that your combined leverage across three exchanges has crept from 3x to 7x because each exchange only sees its own slice. The May 2025 cascade liquidations on BitMEX made this concrete: bots running isolated risk checks per exchange missed the aggregate exposure building across venues, and when BTC dropped 12% in ninety minutes, the cross-exchange margin calls arrived simultaneously. Operators who had $200K in aggregate equity across four exchanges discovered they were effectively running 11x leverage when positions were netted. The liquidation cascade took 47 seconds from first margin call to full wipeout. This guide builds a complete, production-grade risk monitoring system that sits above your exchange connections and provides a unified view of portfolio risk. It uses GreenHelix's event bus for real-time data aggregation, webhooks for alert delivery, and SLA compliance monitoring for operational health tracking. You will build drawdown trackers, correlation monitors, liquidation proximity engines, and circuit breakers -- all wired together through a single risk management layer that treats your entire multi-exchange operation as one portfolio.
Table of Contents
- Risk Architecture Overview
- The RiskMonitor Class
- Drawdown Monitoring and Alerts
- Correlation Monitoring
- Liquidation Proximity Engine
- Circuit Breaker Implementation
- Multi-Exchange Aggregation
- What's Next
Chapter 1: Risk Architecture Overview
Why Centralized Risk Monitoring Matters
Every professional trading desk has a risk management layer that sits above the execution layer. The risk system sees all positions, all strategies, all venues. It computes aggregate exposure, monitors correlation regimes, and has the authority to halt trading when thresholds are breached. Retail and semi-professional bot operators skip this layer because building it is hard and the exchanges do not provide the APIs to do it natively across venues. The result is that a $500K multi-exchange bot operation runs with less risk infrastructure than a single-desk prop trader at a regional firm.
The failure modes are specific and predictable:
- Cross-exchange leverage accumulation: Each exchange computes margin independently. A 3x position on Binance and a 3x position on Coinbase in correlated assets is effectively 6x aggregate leverage against your total equity -- but neither exchange reports it that way.
- Correlation regime changes: Two strategies that were uncorrelated during backtesting become highly correlated during market stress. The diversification benefit you assumed in your position sizing disappears exactly when you need it.
- Cascading liquidations: A liquidation on one exchange reduces your total equity, which increases your leverage ratio on other exchanges, which triggers further liquidations. This feedback loop operates faster than human reaction time.
- Silent risk drift: Without continuous monitoring, risk parameters drift over days and weeks. A portfolio that started at 2x aggregate leverage creeps to 5x through a series of individually reasonable position additions.
Architecture
The system has four layers:
+------------------+ +------------------+ +------------------+
| Binance Bot | | Coinbase Bot | | Kraken Bot |
| (execution) | | (execution) | | (execution) |
+--------+---------+ +--------+---------+ +--------+---------+
| | |
| publish_event | publish_event | publish_event
v v v
+------------------------------------------------------------------------+
| GreenHelix Event Bus |
| (real-time event ingestion, schema validation, ordering) |
+--------+---------------------+-----------------------+-----------------+
| | |
| get_events | webhooks | get_sla_compliance
v v v
+------------------------------------------------------------------------+
| Risk Engine |
| +----------------+ +------------------+ +------------------------+ |
| | DrawdownTracker| | CorrelationMon | | LiquidationProximity | |
| +----------------+ +------------------+ +------------------------+ |
| +----------------+ +------------------+ +------------------------+ |
| | CircuitBreaker | | ExchangeAggr | | AlertManager | |
| +----------------+ +------------------+ +------------------------+ |
+--------+---------------------------------------------------------------+
|
| send_message / register_webhook
v
+------------------------------------------------------------------------+
| Alert Delivery |
| Slack, PagerDuty, Telegram, Email, SMS |
+------------------------------------------------------------------------+
GreenHelix Tools Used
This guide uses seven GreenHelix tools:
| Tool | Purpose |
|---|---|
register_agent |
Register the risk monitoring agent |
register_webhook |
Set up alert delivery endpoints |
publish_event |
Bots publish position and trade events |
get_events |
Risk engine retrieves events for analysis |
get_sla_compliance |
Monitor risk engine uptime and latency |
submit_metrics |
Report risk metrics for observability |
send_message |
Deliver alerts to operators |
Risk Metrics Hierarchy
Risk metrics are computed at four levels, each aggregating from the level below:
- Position-level: Per-position P&L, unrealized P&L, distance to liquidation, margin utilization
- Strategy-level: Strategy drawdown, strategy Sharpe ratio (rolling), strategy exposure, number of open positions
- Portfolio-level: Aggregate drawdown, cross-strategy correlation matrix, net exposure by asset, aggregate leverage
- Fleet-level: Total equity at risk, number of active strategies, circuit breaker status, SLA compliance score
Each level publishes events to the GreenHelix event bus, and higher levels consume events from lower levels. This creates a clean data flow where the risk engine never polls exchanges directly -- it only reads from the event bus.
Event Types for Risk Monitoring
Define one event type per risk signal:
| Event Type | Trigger | Key Fields |
|---|---|---|
risk.portfolio_check |
Periodic risk scan completes | total_equity, aggregate_leverage, alerts |
risk.drawdown_alert |
Drawdown exceeds threshold | level, threshold_pct, current_pct, trigger_timeframe |
risk.correlation_matrix |
Correlation matrix updated | matrix, strategy_count |
risk.correlation_regime_change |
Correlation shift detected | avg_change, max_change |
risk.liquidation_proximity |
Liquidation check completes | positions, min_distance_pct |
risk.circuit_breaker_transition |
State change | from, to, reason, trip_count |
risk.circuit_breaker_config |
Triggers updated | triggers |
risk.aggregate_snapshot |
Cross-exchange aggregation | total_equity, aggregate_leverage, net_exposure |
risk.thresholds_updated |
Risk thresholds changed | thresholds |
risk.key_rotated |
Agent key rotation | new_public_key |
Each event is signed with the risk agent's Ed25519 private key. The signature covers the canonical JSON serialization of the payload (keys sorted, no whitespace), ensuring that risk events cannot be fabricated or tampered with after the fact. This matters when you need to prove to auditors or investors that a circuit breaker tripped at a specific time for a specific reason.
Verifying the API Connection
Before building the full risk engine, verify that your GreenHelix API credentials work:
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "register_agent",
"input": {
"agent_id": "risk-monitor-prod-01",
"public_key": "'"$PUBLIC_KEY_B64"'",
"name": "Portfolio Risk Monitor"
}
}'
import requests
API_BASE = "https://api.greenhelix.net/v1"
API_KEY = "your-api-key"
response = requests.post(
f"{API_BASE}/v1",
headers={
"Authorization": f"Bearer {API_KEY}",
"Content-Type": "application/json",
},
json={
"tool": "register_agent",
"input": {
"agent_id": "risk-monitor-prod-01",
"public_key": "your-public-key-base64",
"name": "Portfolio Risk Monitor",
},
},
)
print(response.json())
If both return a success response with your agent ID, the foundation is in place.
Chapter 2: The RiskMonitor Class
Core Infrastructure
The RiskMonitor class is the foundation for all risk monitoring in this guide. Every subsequent chapter builds on it.
# Generate an Ed25519 keypair for the risk monitor agent
python3 -c "
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
from cryptography.hazmat.primitives import serialization
import base64
private_key = Ed25519PrivateKey.generate()
public_key = private_key.public_key()
priv_bytes = private_key.private_bytes(
encoding=serialization.Encoding.Raw,
format=serialization.PrivateFormat.Raw,
encryption_algorithm=serialization.NoEncryption()
)
pub_bytes = public_key.public_bytes(
encoding=serialization.Encoding.Raw,
format=serialization.PublicFormat.Raw
)
print(f'PRIVATE_KEY_B64={base64.b64encode(priv_bytes).decode()}')
print(f'PUBLIC_KEY_B64={base64.b64encode(pub_bytes).decode()}')
"
# Register the risk monitor agent
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "register_agent",
"input": {
"agent_id": "risk-monitor-prod-01",
"public_key": "'"$PUBLIC_KEY_B64"'",
"name": "Portfolio Risk Monitor"
}
}'
Python Implementation
import json
import time
import base64
import hashlib
from datetime import datetime, timezone
from typing import Dict, List, Optional, Any
import requests
from cryptography.hazmat.primitives.asymmetric.ed25519 import Ed25519PrivateKey
from cryptography.hazmat.primitives import serialization
class RiskMonitor:
"""Cross-exchange portfolio risk monitoring using GreenHelix APIs."""
def __init__(
self,
api_key: str,
agent_id: str,
private_key_b64: str,
base_url: str = "https://api.greenhelix.net/v1",
):
self.api_key = api_key
self.agent_id = agent_id
self.base_url = base_url
self.session = requests.Session()
self.session.headers.update({
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
})
# Load Ed25519 private key
key_bytes = base64.b64decode(private_key_b64)
self._private_key = Ed25519PrivateKey.from_private_bytes(key_bytes)
self._public_key = self._private_key.public_key()
# Risk thresholds (defaults)
self.thresholds = {
"max_drawdown_pct": 15.0,
"max_correlation": 0.85,
"max_leverage": 5.0,
"min_margin_ratio": 0.20,
"max_single_position_pct": 25.0,
}
def _execute(self, tool: str, input_data: dict) -> dict:
"""Execute a GreenHelix tool."""
response = self.session.post(
f"{self.base_url}/v1",
json={"tool": tool, "input": input_data},
)
response.raise_for_status()
return response.json()
def _sign(self, payload: dict) -> str:
"""Sign a payload with Ed25519."""
canonical = json.dumps(payload, sort_keys=True, separators=(",", ":"))
signature = self._private_key.sign(canonical.encode("utf-8"))
return base64.b64encode(signature).decode()
def register_risk_agent(self) -> dict:
"""Register this risk monitor as a GreenHelix agent."""
pub_bytes = self._public_key.public_bytes(
encoding=serialization.Encoding.Raw,
format=serialization.PublicFormat.Raw,
)
return self._execute("register_agent", {
"agent_id": self.agent_id,
"public_key": base64.b64encode(pub_bytes).decode(),
"name": f"Risk Monitor ({self.agent_id})",
})
def configure_thresholds(
self,
max_drawdown_pct: float = 15.0,
max_correlation: float = 0.85,
max_leverage: float = 5.0,
min_margin_ratio: float = 0.20,
max_single_position_pct: float = 25.0,
) -> dict:
"""Set risk thresholds and publish them as a configuration event."""
self.thresholds = {
"max_drawdown_pct": max_drawdown_pct,
"max_correlation": max_correlation,
"max_leverage": max_leverage,
"min_margin_ratio": min_margin_ratio,
"max_single_position_pct": max_single_position_pct,
}
event_payload = {
"agent_id": self.agent_id,
"thresholds": self.thresholds,
"timestamp": datetime.now(timezone.utc).isoformat(),
}
event_payload["signature"] = self._sign(self.thresholds)
return self._execute("publish_event", {
"event_type": "risk.thresholds_updated",
"payload": event_payload,
})
def check_portfolio_risk(self, positions: List[dict]) -> dict:
"""
Run all risk checks against current positions.
Returns a dict with risk metrics and any triggered alerts.
"""
total_equity = sum(p.get("equity", 0) for p in positions)
total_notional = sum(abs(p.get("notional", 0)) for p in positions)
leverage = total_notional / total_equity if total_equity > 0 else float("inf")
alerts = []
# Check leverage
if leverage > self.thresholds["max_leverage"]:
alerts.append({
"type": "leverage_breach",
"current": round(leverage, 2),
"threshold": self.thresholds["max_leverage"],
"severity": "critical",
})
# Check concentration
for pos in positions:
if total_notional > 0:
concentration = abs(pos.get("notional", 0)) / total_notional * 100
if concentration > self.thresholds["max_single_position_pct"]:
alerts.append({
"type": "concentration_breach",
"position": pos.get("symbol", "unknown"),
"current_pct": round(concentration, 2),
"threshold": self.thresholds["max_single_position_pct"],
"severity": "warning",
})
risk_summary = {
"timestamp": datetime.now(timezone.utc).isoformat(),
"total_equity": total_equity,
"total_notional": total_notional,
"aggregate_leverage": round(leverage, 2),
"position_count": len(positions),
"alerts": alerts,
"status": "critical" if any(a["severity"] == "critical" for a in alerts)
else "warning" if alerts
else "healthy",
}
# Publish risk check result as event
event_payload = {
"agent_id": self.agent_id,
**risk_summary,
}
event_payload["signature"] = self._sign(risk_summary)
self._execute("publish_event", {
"event_type": "risk.portfolio_check",
"payload": event_payload,
})
return risk_summary
Configuring Thresholds via curl
# Publish a threshold configuration event
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "publish_event",
"input": {
"event_type": "risk.thresholds_updated",
"payload": {
"agent_id": "risk-monitor-prod-01",
"thresholds": {
"max_drawdown_pct": 15.0,
"max_correlation": 0.85,
"max_leverage": 5.0,
"min_margin_ratio": 0.20,
"max_single_position_pct": 25.0
},
"timestamp": "2026-04-07T12:00:00Z",
"signature": "base64-ed25519-signature"
}
}
}'
Usage Example
monitor = RiskMonitor(
api_key="your-api-key",
agent_id="risk-monitor-prod-01",
private_key_b64="your-private-key-base64",
)
# Register and configure
monitor.register_risk_agent()
monitor.configure_thresholds(
max_drawdown_pct=12.0,
max_leverage=4.0,
min_margin_ratio=0.25,
)
# Check risk against current positions
positions = [
{"symbol": "BTC-USD", "notional": 50000, "equity": 15000, "exchange": "binance"},
{"symbol": "ETH-USD", "notional": 30000, "equity": 10000, "exchange": "coinbase"},
{"symbol": "SOL-USD", "notional": 20000, "equity": 8000, "exchange": "kraken"},
]
result = monitor.check_portfolio_risk(positions)
print(f"Status: {result['status']}, Leverage: {result['aggregate_leverage']}x")
The RiskMonitor class provides the transport layer and basic risk checks. The following chapters build specialized risk engines that integrate with it.
Chapter 3: Drawdown Monitoring and Alerts
Why Drawdown Is the Primary Risk Metric
Drawdown measures peak-to-trough decline. Unlike volatility (which is symmetric) or VaR (which is model-dependent), drawdown directly answers the question operators care about: "How much have I lost from my best point?" A 15% drawdown means your portfolio is worth 85% of its highest recorded value. Drawdown is also the metric most directly tied to survival -- a 50% drawdown requires a 100% gain to recover, and a 75% drawdown requires a 300% gain. Bot operators who monitor drawdown at multiple timeframes catch problems before they become existential.
The DrawdownTracker Class
import time
from collections import defaultdict
from datetime import datetime, timezone
from typing import Dict, List, Optional, Tuple
class DrawdownTracker:
"""Tracks drawdown across multiple timeframes with alert escalation."""
# Timeframe windows in seconds
TIMEFRAMES = {
"1h": 3600,
"4h": 14400,
"24h": 86400,
"7d": 604800,
}
# Alert escalation levels
ALERT_LEVELS = {
"warning": 5.0,
"critical": 8.0,
"emergency": 12.0,
"circuit_breaker": 15.0,
}
def __init__(self, risk_monitor: "RiskMonitor"):
self.risk_monitor = risk_monitor
self._equity_history: List[Tuple[float, float]] = [] # (timestamp, equity)
self._peaks: Dict[str, float] = {} # timeframe -> peak equity
self._current_drawdowns: Dict[str, float] = {}
self._alert_history: List[dict] = []
def update(self, equity: float, timestamp: Optional[float] = None) -> dict:
"""
Record a new equity value and compute drawdowns across all timeframes.
Returns the current drawdown state.
"""
ts = timestamp or time.time()
self._equity_history.append((ts, equity))
# Prune history older than 7 days
cutoff = ts - self.TIMEFRAMES["7d"]
self._equity_history = [
(t, e) for t, e in self._equity_history if t >= cutoff
]
drawdowns = {}
for tf_name, tf_seconds in self.TIMEFRAMES.items():
tf_cutoff = ts - tf_seconds
tf_values = [e for t, e in self._equity_history if t >= tf_cutoff]
if not tf_values:
drawdowns[tf_name] = 0.0
continue
peak = max(tf_values)
self._peaks[tf_name] = peak
current_dd = ((peak - equity) / peak) * 100 if peak > 0 else 0.0
drawdowns[tf_name] = round(current_dd, 4)
self._current_drawdowns = drawdowns
return drawdowns
def check_alerts(self) -> List[dict]:
"""
Check current drawdowns against alert thresholds.
Returns a list of triggered alerts, sorted by severity.
"""
alerts = []
worst_dd = max(self._current_drawdowns.values()) if self._current_drawdowns else 0.0
for level_name, level_threshold in sorted(
self.ALERT_LEVELS.items(), key=lambda x: x[1]
):
if worst_dd >= level_threshold:
# Find which timeframe triggered
trigger_tf = max(
self._current_drawdowns, key=self._current_drawdowns.get
)
alert = {
"level": level_name,
"threshold_pct": level_threshold,
"current_pct": round(worst_dd, 4),
"trigger_timeframe": trigger_tf,
"timestamp": datetime.now(timezone.utc).isoformat(),
"peak_equity": self._peaks.get(trigger_tf, 0),
}
alerts.append(alert)
if alerts:
self._alert_history.extend(alerts)
# Publish the most severe alert
worst_alert = alerts[-1]
payload = {
"agent_id": self.risk_monitor.agent_id,
"alert_type": "drawdown",
**worst_alert,
}
payload["signature"] = self.risk_monitor._sign(worst_alert)
self.risk_monitor._execute("publish_event", {
"event_type": "risk.drawdown_alert",
"payload": payload,
})
return alerts
def get_history(self, timeframe: str = "24h") -> List[dict]:
"""Get drawdown history for a specific timeframe."""
tf_seconds = self.TIMEFRAMES.get(timeframe, 86400)
cutoff = time.time() - tf_seconds
return [
{"timestamp": t, "equity": e}
for t, e in self._equity_history
if t >= cutoff
]
Webhook-Based Alert Delivery
Register a webhook to receive drawdown alerts in real time:
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "register_webhook",
"input": {
"url": "https://your-ops-server.example.com/risk/drawdown",
"event_types": [
"risk.drawdown_alert",
"risk.portfolio_check"
],
"secret": "your-webhook-hmac-secret"
}
}'
# Register webhooks for drawdown alerts
monitor.register_risk_agent()
webhook_result = monitor._execute("register_webhook", {
"url": "https://your-ops-server.example.com/risk/drawdown",
"event_types": ["risk.drawdown_alert", "risk.portfolio_check"],
"secret": "your-webhook-hmac-secret",
})
print(f"Webhook registered: {webhook_result}")
Alert Escalation in Practice
The four-level escalation model maps directly to operational responses:
| Level | Threshold | Action |
|---|---|---|
| Warning | 5% drawdown | Log and notify via Slack |
| Critical | 8% drawdown | Page the on-call operator, reduce position sizes by 50% |
| Emergency | 12% drawdown | Halt new position opens, begin unwinding existing positions |
| Circuit Breaker | 15% drawdown | Cancel all open orders, close all positions, halt all strategies |
Running the Drawdown Loop
tracker = DrawdownTracker(risk_monitor=monitor)
# Simulate a monitoring loop (in production, this runs continuously)
def run_drawdown_monitor(monitor: RiskMonitor, tracker: DrawdownTracker):
"""Poll equity every 10 seconds and check drawdown alerts."""
while True:
# Fetch latest equity from your exchange aggregation layer
events = monitor._execute("get_events", {
"event_type": "exchange.equity_update",
"limit": 1,
"order": "desc",
})
if events.get("events"):
equity = events["events"][0]["payload"].get("total_equity", 0)
drawdowns = tracker.update(equity)
alerts = tracker.check_alerts()
if alerts:
worst = alerts[-1]
print(
f"[{worst['level'].upper()}] Drawdown: {worst['current_pct']}% "
f"(threshold: {worst['threshold_pct']}%, "
f"timeframe: {worst['trigger_timeframe']})"
)
# Send alert via messaging
if worst["level"] in ("emergency", "circuit_breaker"):
monitor._execute("send_message", {
"to": "ops-team",
"subject": f"RISK ALERT: {worst['level'].upper()} drawdown",
"body": (
f"Drawdown: {worst['current_pct']}% "
f"(threshold: {worst['threshold_pct']}%)\n"
f"Timeframe: {worst['trigger_timeframe']}\n"
f"Peak equity: ${worst['peak_equity']:,.2f}"
),
})
time.sleep(10)
# Query drawdown alert history via curl
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "get_events",
"input": {
"event_type": "risk.drawdown_alert",
"limit": 20,
"order": "desc"
}
}'
Drawdown Severity Classification
Not all drawdowns are equal. A 10% drawdown in a trending market during high volatility is qualitatively different from a 10% drawdown in a ranging market during low volatility. The tracker uses the drawdown across multiple timeframes to classify severity:
- Shallow, slow drawdown (24h drawdown > 1h drawdown): The portfolio is grinding lower. This is typically caused by a losing trend-following position. Monitor but do not panic -- the strategy may recover when the trend reverses.
- Deep, fast drawdown (1h drawdown > 24h drawdown): The portfolio is in rapid decline. This pattern is characteristic of liquidation cascades, flash crashes, or correlated position blowups. Immediate attention required.
- Oscillating drawdown (1h and 4h drawdowns similar, both < 24h): The portfolio is whipsawing. This pattern is common during high-volatility news events. Reduce position sizes rather than closing positions -- the whipsaw may resolve in either direction.
The drawdown tracker publishes this classification in the event payload, allowing downstream consumers (the circuit breaker, alert manager) to tailor their response to the type of drawdown, not just the magnitude.
The drawdown tracker is intentionally simple. It uses empirical peaks, not modeled ones. It does not try to predict future drawdown -- it measures what has already happened and triggers actions based on predefined thresholds. Predictive models add complexity without adding reliability under stress conditions.
Chapter 4: Correlation Monitoring
Why Correlation Kills Portfolios
Diversification is the only free lunch in finance -- until it isn't. Portfolio theory assumes that combining uncorrelated strategies reduces total risk. If Strategy A has a Sharpe of 1.5 and Strategy B has a Sharpe of 1.5, and they are uncorrelated, the combined portfolio has a Sharpe of roughly 2.1. But correlation is not static. During market stress, correlations spike. The 2020 COVID crash saw BTC-ETH correlation go from 0.6 to 0.95 in 48 hours. The 2022 LUNA collapse pushed correlations across all crypto assets above 0.9 for weeks. Strategies that appeared diversified during calm markets became a single concentrated bet during the exact conditions where diversification mattered most.
Monitoring correlation in real time -- and detecting regime changes -- lets you reduce exposure before the diversification illusion costs you capital.
The CorrelationMonitor Class
import math
from collections import deque
from datetime import datetime, timezone
from typing import Dict, List, Optional, Tuple
class CorrelationMonitor:
"""
Rolling correlation computation between strategies
with regime change detection.
"""
def __init__(
self,
risk_monitor: "RiskMonitor",
window_size: int = 100,
regime_change_threshold: float = 0.3,
):
self.risk_monitor = risk_monitor
self.window_size = window_size
self.regime_change_threshold = regime_change_threshold
# Returns buffer per strategy
self._returns: Dict[str, deque] = {}
self._correlation_history: List[dict] = []
self._previous_matrix: Optional[Dict[str, Dict[str, float]]] = None
def add_returns(self, strategy_id: str, returns_value: float) -> None:
"""Add a return observation for a strategy."""
if strategy_id not in self._returns:
self._returns[strategy_id] = deque(maxlen=self.window_size)
self._returns[strategy_id].append(returns_value)
def _pearson(self, x: List[float], y: List[float]) -> float:
"""Compute Pearson correlation between two series."""
n = min(len(x), len(y))
if n < 10:
return 0.0
x, y = x[:n], y[:n]
mean_x = sum(x) / n
mean_y = sum(y) / n
cov = sum((xi - mean_x) * (yi - mean_y) for xi, yi in zip(x, y))
std_x = math.sqrt(sum((xi - mean_x) ** 2 for xi in x))
std_y = math.sqrt(sum((yi - mean_y) ** 2 for yi in y))
if std_x == 0 or std_y == 0:
return 0.0
return cov / (std_x * std_y)
def compute_matrix(self) -> Dict[str, Dict[str, float]]:
"""Compute the full correlation matrix across all strategies."""
strategies = sorted(self._returns.keys())
matrix = {}
for s1 in strategies:
matrix[s1] = {}
for s2 in strategies:
if s1 == s2:
matrix[s1][s2] = 1.0
else:
corr = self._pearson(
list(self._returns[s1]),
list(self._returns[s2]),
)
matrix[s1][s2] = round(corr, 4)
# Publish the matrix as an event
payload = {
"agent_id": self.risk_monitor.agent_id,
"matrix": matrix,
"strategy_count": len(strategies),
"timestamp": datetime.now(timezone.utc).isoformat(),
}
payload["signature"] = self.risk_monitor._sign({"matrix": matrix})
self.risk_monitor._execute("publish_event", {
"event_type": "risk.correlation_matrix",
"payload": payload,
})
self._correlation_history.append({
"timestamp": datetime.now(timezone.utc).isoformat(),
"matrix": matrix,
})
return matrix
def detect_regime_change(self) -> Optional[dict]:
"""
Detect if the correlation regime has changed significantly.
Compares current matrix to previous matrix. A regime change is
detected when the average absolute change in correlations exceeds
the threshold.
"""
current = self.compute_matrix()
if self._previous_matrix is None:
self._previous_matrix = current
return None
strategies = sorted(current.keys())
changes = []
for s1 in strategies:
for s2 in strategies:
if s1 >= s2:
continue
prev = self._previous_matrix.get(s1, {}).get(s2, 0.0)
curr = current.get(s1, {}).get(s2, 0.0)
changes.append(abs(curr - prev))
if not changes:
self._previous_matrix = current
return None
avg_change = sum(changes) / len(changes)
max_change = max(changes)
self._previous_matrix = current
if avg_change > self.regime_change_threshold:
alert = {
"type": "correlation_regime_change",
"avg_change": round(avg_change, 4),
"max_change": round(max_change, 4),
"threshold": self.regime_change_threshold,
"timestamp": datetime.now(timezone.utc).isoformat(),
"matrix": current,
}
# Publish regime change event
event_payload = {
"agent_id": self.risk_monitor.agent_id,
**alert,
}
event_payload["signature"] = self.risk_monitor._sign(alert)
self.risk_monitor._execute("publish_event", {
"event_type": "risk.correlation_regime_change",
"payload": event_payload,
})
return alert
return None
Querying Correlation History via curl
# Get recent correlation matrix events
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "get_events",
"input": {
"event_type": "risk.correlation_matrix",
"limit": 10,
"order": "desc"
}
}'
# Get regime change alerts
curl -X POST https://sandbox.greenhelix.net/v1 \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"tool": "get_events",
"input": {
"event_type": "risk.correlation_regime_change",
"limit": 5,
"order": "desc"
}
}'
Usage Example
corr_monitor = CorrelationMonitor(
risk_monitor=monitor,
window_size=100,
regime_change_threshold=0.3,
)
# Feed returns from each strategy (in production, these come from trade events)
import random
for i in range(120):
corr_monitor.add_returns("mean-reversion-btc", random.gauss(0.001, 0.02))
corr_monitor.add_returns("momentum-eth", random.gauss(0.0005, 0.025))
corr_monitor.add_returns("arb-sol-perps", random.gauss(0.0008, 0.015))
matrix = corr_monitor.compute_matrix()
for s1, row in matrix.items():
print(f"{s1}: {row}")
regime = corr_monitor.detect_regime_change()
if regime:
print(f"REGIME CHANGE detected: avg change {regime['avg_change']}")
Interpreting the Correlation Matrix
The correlation matrix is symmetric with 1.0 on the diagonal. Off-diagonal values range from -1.0 (perfectly inverse) to +1.0 (perfectly correlated). For risk management purposes:
- Below 0.3: Strategies are effectively uncorrelated. Diversification benefit is real.
- 0.3 to 0.6: Moderate correlation. Some diversification benefit remains, but not full.
- 0.6 to 0.85: High correlation. The strategies will draw down together in stressed markets. Position sizes should account for this.
- Above 0.85: The strategies are effectively the same bet. The circuit breaker should consider triggering if this persists.
Practical Considerations for Correlation Monitoring
Window size selection: The window_size parameter controls how many return observations are used to compute correlation. Smaller windows (30-50) are more responsive to recent correlation changes but produce noisier estimates. Larger windows (100-200) are more stable but slower to detect regime changes. A common compromise is to run two correlation monitors in parallel -- one with a 50-observation window for early detection and one with a 150-observation window for confirmation.
Return frequency: The correlation monitor needs returns at a consistent frequency. If Strategy A reports returns every minute and Strategy B reports every 15 minutes, the correlation estimate will be unreliable. Normalize to the lowest common frequency across all strategies. For most crypto bot operations, 5-minute or 15-minute returns provide a good balance between responsiveness and stability.
Handling missing data: If a strategy stops reporting returns (e.g., because the bot crashed or the exchange is down), the correlation monitor should flag this rather than computing correlation with stale data. A strategy that has not reported in 30
…(truncated)