Albucore Benchmarks
Before designing a performance comparison, read ../performance-optimization/SKILL.md and
../../../docs/performance-optimization.md completely. Extend the benchmark along the dimension that controls the
candidate, such as label density for bincount, table and channel layout for LUTs, or output size and dtype for random
generation.
Use exactly one CPU thread per process for every candidate, following the thread controls in the canonical performance guide. Benchmark additional thread counts only when the user explicitly requests thread scaling.
Layout
benchmarks/- Python timing scripts. Run from repo root:uv run python benchmarks/<script>.py.benchmarks/timing.py- Sharedmedian_mshelper for scripts executed aspython benchmarks/foo.py../benchmark.sh- Dataset-driven runner; expects an externalbenchmarkpackage that is not always present in-tree. Prefer synthetic scripts for CI-style checks.benchmarks/benchmark_router_synthetic.py- Times public routers on syntheticuint8andfloat32arrays: HWC, plus NHWC formean,std, andmean_stdonly.benchmarks/compare_router_json.py- Builds a Markdown table from two JSON outputs.benchmarks/benchmark_resize3d_tensor.py- Times direct Tensor, zero-copy Tensor→NumPy→Tensor, and publicresize3droutes for contiguous and channel-last-strided CPUCDHWTensors.benchmarks/benchmark_warp_affine3d.py- Times full single-volume NumPyDHWCaffine paths, including the NumPy→Torch bridge and public router.benchmarks/benchmark_warp_affine3d_tensor.py- Times native Torch affine-grid, manual-grid and coverage-fill probes, and public single-volumeCDHWrouting.
Canonical Shape Grid
Benchmark shape sweeps use channel-last Albucore conventions.
HWC images:
128x160with 1, 3, 9 channels - small / warm-cache, non-square.240x320with 1, 3, 9 channels - mid-size crop, non-square.480x640with 1, 3, 9 channels - typical augmentation training crop, non-square.768x1024with 1, 3, 9 channels - high-res / full-image pass, non-square.
Use non-square H/W pairs so height-width swaps fail visibly. Avoid square-only benchmark grids.
DHWC volumes:
16x128x160x1,16x128x160x3- thin slab, non-square in-plane.32x128x160x1,32x128x160x3- common nnU-Net patch depth.64x128x160x3- deeper slab.96x128x160x1- deep single-channel slab.48x240x320x3- large in-plane, multi-channel.
For resize3d, also include C=5, unit input/output spatial axes, and an explicit D*C value on both sides of the OpenCV encoded-channel boundary. Time its public NumPy route end-to-end, including channel packing, Torch conversions, and output repair. For Tensor input, sweep contiguous and channel-last-strided CDHW, direct interpolation, the zero-copy bridge, and the public router. Use uv run python benchmarks/benchmark_resize3d.py --quick and uv run python benchmarks/benchmark_resize3d_tensor.py --quick while iterating; record any resulting routing decision in docs/numkong-performance.md or a focused report under benchmarks/results/.
For warp_affine3d, benchmark only one volume per call: NumPy DHWC or CPU Tensor CDHW. The full matrix uses
uint8/float32, C=1/3/5/9, canonical output sizes including a unit output axis, nearest/trilinear interpolation,
one 3×4 forward matrix per scenario, and zero/nonzero fill. NumPy timings use contiguous inputs; Tensor timings add
contiguous and channel-last-strided inputs. Test the equivalent homogeneous 4×4 representation as a contract, not a
timing route. Run uv run python benchmarks/benchmark_warp_affine3d.py --quick --threads 1 and uv run python benchmarks/benchmark_warp_affine3d_tensor.py --quick --threads 1. A manual grid, coverage sampler, tiled route, or
native extension remains a diagnostic candidate until it has exact correctness parity and a sustained full-path win.
Channel choices: 1 for grayscale, 3 for RGB / 3-channel, and 9 for hyperspectral paths that exceed MAX_OPENCV_WORKING_CHANNELS=4.
Compare the current tree with a previous release
uv run python benchmarks/benchmark_router_synthetic.py \
--output-json benchmarks/results/router-current.json
uv run --no-project --with albucore==<previous-version> --with opencv-python-headless \
--with numkong --with stringzilla --with numpy \
python benchmarks/benchmark_router_synthetic.py \
--output-json benchmarks/results/router-previous.json
uv run python benchmarks/compare_router_json.py \
benchmarks/results/router-current.json \
benchmarks/results/router-previous.json \
benchmarks/results/REPORT_router_current_vs_previous.md
Replace <previous-version> with the release that answers the current question. Use --quick for smaller
shape/channel grids while iterating, and do not accumulate version-specific baselines in the repository.
Docs
- Current NumKong route decisions:
docs/numkong-performance.md - Generated benchmark evidence:
benchmarks/results/ - General performance policy:
docs/performance-optimization.md