# Bubbleml Eval

> bubbleml-eval

- Skill: `qhjqhj00/bubbleml-eval` (Agent Skill)
- Install (CLI): `npx skillmds@latest add qhjqhj00/bubbleml-eval`
- Raw SKILL.md: https://api.skillmd.com/api/skills/qhjqhj00/bubbleml-eval/raw
- Safety review: pending (external: skill-scanner PASS, skillspector PASS)
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: qhjqhj00 (https://skillmd.com/u/qhjqhj00)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/qhjqhj00/bubbleml-eval

---


# bubbleml-eval

> BubbleML: A Multi-Physics Dataset and Benchmarks for Machine Learning — Sheikh Md Shakeel Hassan et al. (2023) (arXiv:2307.14623, 2023)

## What this evaluates

This benchmark evaluates machine learning models on two distinct scientific tasks using the BubbleML dataset: predicting optical flow for bubble dynamics and solving multiphysics PDEs for temperature and velocity field propagation. It probes a model's ability to capture non-rigid object motion, sharp physical interfaces, and long-horizon temporal dynamics in phase-change simulations.

## Datasets

- **BubbleML** — total 2500; splits: train (2000), val (500); repo https://github.com/HPCForge/BubbleML

## Metrics

- `end-point error (EPE)` **(primary)** — range: other
  - Average L2 distance between predicted and ground truth flow vectors across all pixels. Lower is better.
- `RMSE` — range: other
  - Root Mean Squared Error between predicted and ground truth physical fields (temperature or velocity). Lower is better.
- `IRMSE` — range: other
  - RMSE computed exclusively along liquid-vapor bubble interfaces to penalize misalignment at sharp physical boundaries.

## Input / output format

**Input**: Optical flow: consecutive image frames tracking bubble positions. SciML: past solution fields (velocity and/or temperature) at k consecutive timesteps.

**Output**: Optical flow: 2D velocity vector field per pixel in Middlebury flow format. SciML: predicted temperature and/or velocity field at the next timestep.

## Scoring recipe

```python
def compute_epe(pred_flow, gt_flow):
    return np.mean(np.sqrt(np.sum((pred_flow - gt_flow)**2, axis=-1)))

def compute_rmse(pred, gt):
    return np.sqrt(np.mean((pred - gt)**2))

def compute_irmse(pred, gt, interface_mask):
    pred_masked = pred[interface_mask]
    gt_masked = gt[interface_mask]
    return np.sqrt(np.mean((pred_masked - gt_masked)**2))
```

## Common pitfalls

- High error rates at bubble boundaries due to sharp temperature/velocity gradients and non-rigid deformation.
- Auto-regressive rollout models suffer from error accumulation over time, degrading long-horizon predictions.
- Over-fitting to the boiling dataset during fine-tuning harms generalization to other optical flow benchmarks.

## Evidence (verbatim from paper)

> To assess the performance of the trained models, we measure the end-point error and Table 2 summarizes the results for one dataset. We draw inspiration from PDEBench and adopt a large set of metrics that include the Root Mean Squared Error (RMSE), Max Squared Error, Relative Error, Boundary RMSE (BRMSE), and low/mid/high Fourier errors. We incorporate an additional physics metric: the RMSE along bubble interfaces (IRMSE).

## Citation

```bibtex
@misc{hassan2023bubbleml,
  title={BubbleML: A Multi-Physics Dataset and Benchmarks for Machine Learning},
  author={Sheikh Md Shakeel Hassan et al. (2023)},
  year={2023},
  note={arXiv:2307.14623}
}
```

- arXiv: 2307.14623

