# Ml4t Backtest Overfitting

> Detect and prevent overfitting to historical data via multiple testing corrections and pre-registration. Use when evaluating strategy variants to ensure performance is not a data-mining artifact.

- Skill: `ml4t/ml4t-backtest-overfitting` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ml4t/ml4t-backtest-overfitting`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ml4t/ml4t-backtest-overfitting/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: ml4t (https://skillmd.com/u/ml4t)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ml4t/ml4t-backtest-overfitting

---

# Backtest Overfitting

Testing many strategies on the same data guarantees finding one that looks profitable by chance. With 100 independent trials at p < 0.05, you expect five false positives.

## The Problem

Every parameter you tune, every feature you try, and every universe filter you adjust is an implicit trial. A researcher who reports a Sharpe ratio of 2.0 after exploring 200 configurations has not found alpha - they have found the luckiest draw from a noise distribution. Correcting for the number of trials, by haircut here and by Deflated Sharpe Ratio in `ml4t-deflated-sharpe`, is what separates the two. Without it, most published backtests are statistically meaningless.

## The Pattern

### WRONG

```python
# Tune until something looks good
best_sharpe = 0
for lookback in [5, 10, 21, 63, 126, 252]:
    for top_k in [5, 10, 20, 50]:
        result = backtest(lookback=lookback, top_k=top_k)
        sharpe = result["sharpe"]
        if sharpe > best_sharpe:
            best_sharpe = sharpe
            best_params = (lookback, top_k)

print(f"Best Sharpe: {best_sharpe:.2f}")  # meaningless without correction
```

### CORRECT

```python
import numpy as np
from scipy.stats import norm

results = []
for lookback in [5, 10, 21, 63, 126, 252]:
    for top_k in [5, 10, 20, 50]:
        result = backtest(lookback=lookback, top_k=top_k)
        results.append(result["sharpe"])

# Sharpe haircut: subtract the selection bound (Bailey & Lopez de Prado).
n_trials = len(results)
best_sharpe = max(results)
sharpe_std = np.std(results)
expected_max = sharpe_std * (
    (1 - np.euler_gamma) * norm.ppf(1 - 1 / n_trials)
    + np.euler_gamma * norm.ppf(1 - 1 / (n_trials * np.e))
)
# This is NOT the Deflated Sharpe Ratio, which is a probability that also
# takes sample size, skew and kurtosis: see ml4t-deflated-sharpe.
haircut = best_sharpe - expected_max
print(f"Observed: {best_sharpe:.2f}, After haircut: {haircut:.2f}, Trials: {n_trials}")
```

## Probability of Backtest Overfitting (PBO)

PBO uses combinatorial CV (CSCV; see `cpcv` skill) to generate multiple train/test paths, then checks how often the IS-best strategy underperforms OOS:

```python
import numpy as np
from itertools import combinations

n_groups, n_test = 8, 4  # CSCV splits into complementary halves, not 2-of-8
groups = np.array_split(np.arange(len(data)), n_groups)
ranks = []  # relative OOS rank of the strategy chosen in sample
for test_g in combinations(range(n_groups), n_test):
    test = np.concatenate([groups[g] for g in test_g])
    train = np.concatenate([groups[g] for g in range(n_groups) if g not in test_g])
    is_sharpe = [backtest(p, data[train])["sharpe"] for p in param_grid]
    oos_sharpe = [backtest(p, data[test])["sharpe"] for p in param_grid]
    best_is = int(np.argmax(is_sharpe))  # the one you would have shipped
    rank = np.argsort(oos_sharpe)[::-1].tolist().index(best_is)
    ranks.append(rank / (len(param_grid) - 1))  # 0 = best OOS, 1 = worst

pbo = np.mean(np.array(ranks) > 0.5)  # how often the IS winner is below median
```

## Red Flags

| Signal | Concern |
|--------|---------|
| Sharpe > 2.0 on daily data | Almost certainly overfit or leakage |
| OOS matches IS within 5% | Data leakage, not genuine alpha |
| Complex model barely beats simple | Extra parameters fit noise |
| Performance cliff after 2020 | Regime-specific overfitting |

## Guardrails

- Document total configurations tested - each is a trial. Separate exploration from confirmation.
- Pre-register hypothesis and success threshold in version control before any backtest.
- Minimum 5 years daily data (~1,250 observations) for Sharpe estimation.
- If the haircut Sharpe is negative, the strategy has no statistical evidence of alpha.

## Production Implementation

`ml4t-diagnostic` provides validated implementations of both corrections:

```python
from ml4t.diagnostic.evaluation.stats import compute_pbo, benjamini_hochberg_fdr
from ml4t.diagnostic.splitters import CombinatorialCV

cpcv = CombinatorialCV(n_groups=8, n_test_groups=4, embargo_size=5)  # halves, as above
pbo = compute_pbo(np.array(is_sharpes), np.array(oos_sharpes))
rejected = benjamini_hochberg_fdr(p_values, alpha=0.05)
```

## Checklist

- [ ] Strategy hypothesis committed to git BEFORE any backtest
- [ ] Total trials documented (including informal exploration); multiple-testing correction applied
- [ ] Sharpe haircut applied, and DSR from ml4t-deflated-sharpe reported with it
- [ ] PBO calculated from combinatorial CV folds (PBO < 0.50 required)
- [ ] True holdout set preserved and used exactly once

