Anomaly Detection
Aeon provides anomaly detection methods for identifying unusual patterns in time series at both series and collection levels.
Collection Anomaly Detectors
Detect anomalous time series within a collection:
Series Anomaly Detectors
Detect anomalous points or subsequences within a single time series.
Distance-Based Methods
Use similarity metrics to identify anomalies:
CBLOF - Cluster-Based Local Outlier Factor
- Clusters data, identifies outliers based on cluster properties
- Use when: Anomalies form sparse clusters
KMeansAD - K-means based anomaly detection
- Distance to nearest cluster center indicates anomaly
- Use when: Normal patterns cluster well
LeftSTAMPi - Left STAMP incremental
- Matrix profile for online anomaly detection
- Use when: Streaming data, need online detection
STOMP - Scalable Time series Ordered-search Matrix Profile
- Computes matrix profile for subsequence anomalies
- Use when: Discord discovery, motif detection
MERLIN - Matrix profile-based method
- Efficient matrix profile computation
- Use when: Large time series, need scalability
LOF - Local Outlier Factor adapted for time series
- Density-based outlier detection
- Use when: Anomalies in low-density regions
ROCKAD - ROCKET-based semi-supervised detection
- Uses ROCKET features for anomaly identification
- Use when: Have some labeled data, want feature-based approach
Distribution-Based Methods
Analyze statistical distributions:
Isolation-Based Methods
Use isolation principles:
IsolationForest - Random forest-based isolation
- Anomalies easier to isolate than normal points
- Use when: High-dimensional data, no assumptions about distribution
OneClassSVM - Support vector machine for novelty detection
- Learns boundary around normal data
- Use when: Well-defined normal region, need robust boundary
STRAY - Streaming Robust Anomaly Detection
- Robust to data distribution changes
- Use when: Streaming data, distribution shifts
External Library Integration
PyODAdapter - Bridges PyOD library to aeon
- Access 40+ PyOD anomaly detectors
- Use when: Need specific PyOD algorithm
Quick Start
from aeon.anomaly_detection import STOMP
import numpy as np
# Create time series with anomaly
y = np.concatenate([
np.sin(np.linspace(0, 10, 100)),
[5.0], # Anomaly spike
np.sin(np.linspace(10, 20, 100))
])
# Detect anomalies
detector = STOMP(window_size=10)
anomaly_scores = detector.fit_predict(y)
# Higher scores indicate more anomalous points
threshold = np.percentile(anomaly_scores, 95)
anomalies = anomaly_scores > threshold
Point vs Subsequence Anomalies
Point anomalies: Single unusual values
- Use: COPOD, DWT_MLEAD, IsolationForest
Subsequence anomalies (discords): Unusual patterns
- Use: STOMP, LeftSTAMPi, MERLIN
Collective anomalies: Groups of points forming unusual pattern
- Use: Matrix profile methods, clustering-based
Evaluation Metrics
Specialized metrics for anomaly detection:
from aeon.benchmarking.metrics.anomaly_detection import (
range_precision,
range_recall,
range_f_score,
roc_auc_score
)
# Range-based metrics account for window detection
precision = range_precision(y_true, y_pred, alpha=0.5)
recall = range_recall(y_true, y_pred, alpha=0.5)
f1 = range_f_score(y_true, y_pred, alpha=0.5)
Algorithm Selection
- Speed priority: KMeansAD, IsolationForest
- Accuracy priority: STOMP, COPOD
- Streaming data: LeftSTAMPi, STRAY
- Discord discovery: STOMP, MERLIN
- Multi-dimensional: COPOD, PyODAdapter
- Semi-supervised: ROCKAD, OneClassSVM
- No training data: IsolationForest, STOMP
Best Practices
- Normalize data: Many methods sensitive to scale
- Choose window size: For matrix profile methods, window size critical
- Set threshold: Use percentile-based or domain-specific thresholds
- Validate results: Visualize detections to verify meaningfulness
- Handle seasonality: Detrend/deseasonalize before detection
1---2name: anomaly-detection3description: Aeon provides anomaly detection methods for identifying unusual patterns in time series at both series and collection levels.4---5# Anomaly Detection67Aeon provides anomaly detection methods for identifying unusual patterns in time series at both series and collection levels.89## Collection Anomaly Detectors1011Detect anomalous time series within a collection:1213- `ClassificationAdapter` - Adapts classifiers for anomaly detection14 - Train on normal data, flag outliers during prediction15 - **Use when**: Have labeled normal data, want classification-based approach1617- `OutlierDetectionAdapter` - Wraps sklearn outlier detectors18 - Works with IsolationForest, LOF, OneClassSVM19 - **Use when**: Want to use sklearn anomaly detectors on collections2021## Series Anomaly Detectors2223Detect anomalous points or subsequences within a single time series.2425### Distance-Based Methods2627Use similarity metrics to identify anomalies:2829- `CBLOF` - Cluster-Based Local Outlier Factor30 - Clusters data, identifies outliers based on cluster properties31 - **Use when**: Anomalies form sparse clusters3233- `KMeansAD` - K-means based anomaly detection34 - Distance to nearest cluster center indicates anomaly35 - **Use when**: Normal patterns cluster well3637- `LeftSTAMPi` - Left STAMP incremental38 - Matrix profile for online anomaly detection39 - **Use when**: Streaming data, need online detection4041- `STOMP` - Scalable Time series Ordered-search Matrix Profile42 - Computes matrix profile for subsequence anomalies43 - **Use when**: Discord discovery, motif detection4445- `MERLIN` - Matrix profile-based method46 - Efficient matrix profile computation47 - **Use when**: Large time series, need scalability4849- `LOF` - Local Outlier Factor adapted for time series50 - Density-based outlier detection51 - **Use when**: Anomalies in low-density regions5253- `ROCKAD` - ROCKET-based semi-supervised detection54 - Uses ROCKET features for anomaly identification55 - **Use when**: Have some labeled data, want feature-based approach5657### Distribution-Based Methods5859Analyze statistical distributions:6061- `COPOD` - Copula-Based Outlier Detection62 - Models marginal and joint distributions63 - **Use when**: Multi-dimensional time series, complex dependencies6465- `DWT_MLEAD` - Discrete Wavelet Transform Multi-Level Anomaly Detection66 - Decomposes series into frequency bands67 - **Use when**: Anomalies at specific frequencies6869### Isolation-Based Methods7071Use isolation principles:7273- `IsolationForest` - Random forest-based isolation74 - Anomalies easier to isolate than normal points75 - **Use when**: High-dimensional data, no assumptions about distribution7677- `OneClassSVM` - Support vector machine for novelty detection78 - Learns boundary around normal data79 - **Use when**: Well-defined normal region, need robust boundary8081- `STRAY` - Streaming Robust Anomaly Detection82 - Robust to data distribution changes83 - **Use when**: Streaming data, distribution shifts8485### External Library Integration8687- `PyODAdapter` - Bridges PyOD library to aeon88 - Access 40+ PyOD anomaly detectors89 - **Use when**: Need specific PyOD algorithm9091## Quick Start9293```python94from aeon.anomaly_detection import STOMP95import numpy as np9697# Create time series with anomaly98y = np.concatenate([99 np.sin(np.linspace(0, 10, 100)),100 [5.0], # Anomaly spike101 np.sin(np.linspace(10, 20, 100))102])103104# Detect anomalies105detector = STOMP(window_size=10)106anomaly_scores = detector.fit_predict(y)107108# Higher scores indicate more anomalous points109threshold = np.percentile(anomaly_scores, 95)110anomalies = anomaly_scores > threshold111```112113## Point vs Subsequence Anomalies114115- **Point anomalies**: Single unusual values116 - Use: COPOD, DWT_MLEAD, IsolationForest117118- **Subsequence anomalies** (discords): Unusual patterns119 - Use: STOMP, LeftSTAMPi, MERLIN120121- **Collective anomalies**: Groups of points forming unusual pattern122 - Use: Matrix profile methods, clustering-based123124## Evaluation Metrics125126Specialized metrics for anomaly detection:127128```python129from aeon.benchmarking.metrics.anomaly_detection import (130 range_precision,131 range_recall,132 range_f_score,133 roc_auc_score134)135136# Range-based metrics account for window detection137precision = range_precision(y_true, y_pred, alpha=0.5)138recall = range_recall(y_true, y_pred, alpha=0.5)139f1 = range_f_score(y_true, y_pred, alpha=0.5)140```141142## Algorithm Selection143144- **Speed priority**: KMeansAD, IsolationForest145- **Accuracy priority**: STOMP, COPOD146- **Streaming data**: LeftSTAMPi, STRAY147- **Discord discovery**: STOMP, MERLIN148- **Multi-dimensional**: COPOD, PyODAdapter149- **Semi-supervised**: ROCKAD, OneClassSVM150- **No training data**: IsolationForest, STOMP151152## Best Practices1531541. **Normalize data**: Many methods sensitive to scale1552. **Choose window size**: For matrix profile methods, window size critical1563. **Set threshold**: Use percentile-based or domain-specific thresholds1574. **Validate results**: Visualize detections to verify meaningfulness1585. **Handle seasonality**: Detrend/deseasonalize before detection