# Ta Bench

> Benchmarking TA-Lib — ta_bench, ta_bench_direct, ta_bench_stream, ta_bench_icount and scripts/stream_ab.py, what each ratio actually compares (the six-binary build matrix), streaming vs batch, instruction counts, and the --shape= input corpus. Use when running a benchmark in this repo, interpreting a speedup or ratio, or defending a performance number.

- Skill: `ta-lib/ta-bench` (Agent Skill)
- Install (CLI): `npx skillmds@latest add ta-lib/ta-bench`
- Raw SKILL.md: https://api.skillmd.com/api/skills/ta-lib/ta-bench/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: ta-lib (https://skillmd.com/u/ta-lib)
- Updated: 2026-09-17
- Page: https://skillmd.com/skills/ta-lib/ta-bench

---


# Benchmarking TA-Lib

```bash
# Full pipeline (builds everything, regens, tests, benchmarks)
scripts/regtest.py

# Benchmark specific indicators (trustworthy — isolated, high iterations)
cd bin && ./ta_bench --language=cref,c --function=RSI,SMA --points=100000 --iters=500

# Full benchmark (noisy — use for overview, verify outliers in isolation)
cd bin && ./ta_bench --language=cref,c --points=100000 --iters=200
```

**Gotcha:** `ta_ref_serve` is statically linked — rebuild when `libta-lib.a`
changes or benchmarks are invalid. `regtest.py` handles this automatically.

Both hand-written benches report the **spread** of their own repeated passes,
because a bare median is silent about whether the box was quiet enough for it
to mean anything — at `--iters=50` the same five functions read 0.57–0.81x, at
`--iters=200` they read 1.00x. Read the spread before the ratio. `--max-spread=N`
(percent, default 25) exits non-zero when the run is too noisy to interpret, and
`ta_bench_direct --jsonl=PATH` appends a run record for tracking over time.

`ta_bench_direct`'s ratio is `ta_bench_cg` (single TU, `-flto`) over
`libta-lib.a` (separate TUs, no LTO) — **a build-configuration difference, not
an algorithm one**, which is why binary layout alone can move it further than
the old ±10% colour band. It now colours only outside `--no-signal` (default
1.20x) and only when the row's own spread is narrower than the effect claimed.
`--reps=N` samples both arms instead of just the reference.

`ta_bench` sends `no_output:1`, so servers return timings without serialising
the output arrays — it only ever reads `timing_ns`. Without it a 100k-point run
spends ~97% of its wall clock formatting and parsing JSON nobody looks at.
Anything that needs the values (`--codegen`, `--xlang-hash`, `server_verify`)
simply omits the flag. `cref` is a frozen binary and predates it, so runs
including `cref` stay slower than C-only ones.

## The same source, six binaries

Every benchmark ratio in this tree compares two *builds*, and they are not the
same build. Measured `.text` on x86-64 gcc:

| binary | build | TU model | bytes |
|---|---|---|---|
| `libta-lib.a` | CMake Release | separate TUs, no LTO | 2,890,487 |
| `libta-lib.so` | CMake Release | separate TUs, PIC | 2,870,694 |
| `ta_codegen_serve_c` | `gcc -O3 -flto` | single TU + ta_abstract | 3,940,182 |
| `ta_bench_stream` | `gcc -O3 -flto` | single TU + streaming | 1,955,238 |
| `ta_bench_cg` | `gcc -O3 -flto` | single TU, indicators only | 1,021,939 |
| autotools `libta-lib` | libtool | separate TUs, no LTO | not built here |


3.9x between the extremes, from identical source. The two build flags that
all three build systems must keep in step are stated in the root `CLAUDE.md`.

Which tool measures which:

- `ta_bench_direct` — C-ref column is `libta-lib.a`, C column is `ta_bench_cg`.
  Its ratio is therefore rows 1 vs 5 above.
- `ta_bench --language=c` — `ta_codegen_serve_c` (row 3), *not* `ta_bench_cg`.
- `ta_bench --language=cref` — `ta_ref_serve`, the frozen pre-cutover source.
  Different code, not just a different build; the only cross-*version* number.
- `ta_bench_stream` — itself, both arms, which is why its speedup column is the
  one ratio here that isn't cross-configuration.
- `ta_bench_icount` — `libta-lib.a` (row 1), the shipped build. The only
  instrument here that reports no ratio at all: one absolute instruction count
  per entry point. See below.

Consequences worth internalising before quoting any number: a function's ns from
`ta_bench_direct` and from `ta_bench --language=c` are not comparable; a ratio
near 1.0 in `ta_bench_direct` means single-TU + LTO bought nothing for that
function, not that the two are the same code path; and layout alone moves these
ratios further than the old ±10% colour band allowed, which is why the band is
now 1.20x and spread-gated.

### A/B of one change has a trap the table above does not show

The three `-flto` rows are **single-TU** builds — `ta_bench_stream.c` alone
`#include`s 181 `.c` files — so `-flto` there is nearly a no-op and the
compiler re-decides inlining for every function on every build. In an A/B those
decisions can differ **between the arms**. On #252 `TA_ATR_Update` was outlined
in one arm and inlined in the other, and `ta_bench_stream` reported +16% for a
function that is unchanged in the shipped build. Two defences:

- **Carry a control function** whose generated code is byte-identical across
  the arms, and believe nothing smaller than its excursion. RSI once read −55%
  on identical emitted C; 107 unchanged functions put that run's real floor at
  ±15%.
- **When the mechanism is memory traffic or aliasing, measure the shipped
  build.** A throwaway harness linking `libta-lib.a`, one function per process,
  min of N, read ±3% on the same change. That is an ad-hoc check, not a seventh
  instrument — do not add it to `bin/`.

**Is it worth optimizing at all?** The non-inlined call floor — a separate-TU
function that does nothing but write `*out` — is **1.06 ns**. That is ~30% of
SMA's 3.5 ns `Update`, ~15% of MFI's 6.9 ns, and ~1.5% of HT_TRENDLINE's 69 ns.
A saving well under that floor on a cheap function is one no caller can
observe.

## Counting instructions instead of timing

`ta_bench_icount` + `scripts/bench_icount.py` is the nightly regression net: the
`icount` job of `dev-nightly`, never a PR or a push. A job there rather than its
own workflow so its baseline commit is a product of the run that goes green,
which is what keeps "green dev-nightly means mergeable dev" true. It runs every
C entry point once under callgrind (batch, `_Open`, `_OpenAndFill`, `_Update`,
`_Peek`) and compares the retired-instruction count against
`.github/perf/icount-baseline-<arch>.tsv`.

```bash
scripts/bench_icount.py                       # build, measure, compare (needs valgrind)
scripts/bench_icount.py --update-baseline     # ... and lower any row it beat
scripts/bench_icount.py --accept=SMA          # ... and RAISE only SMA's rows
scripts/bench_icount.py --accept=SMA/batch    # ... and only that one tier
scripts/bench_icount.py --no-build --function=RSI,SMA   # narrowed: report only
```

Why it exists next to five timing tools: a count is **exact**. Two runs of the
same binary on a loaded runner and an idle one agree to the instruction, which
is what lets a 10% threshold gate anything on a shared vCPU. The timing tools
all concede that ground: `--max-spread=25`, `--no-signal=1.20`,
`--min-ratio=0.35`.

What a count cannot see, and where it actively misleads:

- Out-of-order execution, port pressure, dependency-chain latency: all free.
- Valgrind over-charges branches, so a branchless rewrite can read here as a
  regression while being a win on hardware.
- **A percentage from this tool is not a speed figure and never goes in a
  release note.** Use it for the algorithmic class (a lost fast path, an extra
  pass over the window, an un-inlined call) and the devbox for magnitudes.

**The baseline only ever moves down.** A passing nightly lowers a row it beat and
holds every row it did not, so a regression under the threshold is never
absorbed: three nights of +9% is +30% against a baseline that never moved, and
the gate catches it. A wholesale nightly rewrite would have read green three
times and lost the drift. Raising a row takes `--accept`, and `--accept` names the rows (workflow
`mode=accept` + the `accept` input). Accepting one deliberate regression does
not re-baseline the corpus: every row not named keeps the monotone rule, so what
the other thousand entry points accumulated survives. A failure on a row nobody
named still fails the run, so `--accept=SMA` cannot absorb a regression in RSI.
`--accept=ALL` exists for a toolchain change and says what it does.

Two more properties to hold on to. `--function` narrows the run, and the
allocating tiers (`open`, `openfill`) then shift by a few hundred instructions
because the heap history they see is different; that is why a narrowed run
reports and never gates. And the baseline is per (architecture, compiler): on a
mismatch the script refuses to compare rather than printing a thousand false
rows.

## Streaming vs batch

`ta_bench_stream` answers the question streaming has to justify itself on: is
`TA_<NAME>_Update` actually cheaper than recomputing the last bar with the batch
call? Its `speedup` column is `batch_last_ns / update_ns` — above 1 means
streaming wins. Both halves are measured in one TU, one input, one layout, so
unlike `ta_bench_direct`'s ratio it is not comparing two build configurations.

```bash
cd bin && ./ta_bench_stream --points=20000 --iters=50
./ta_bench_stream --points=20000 --iters=50 --min-ratio=0.35   # exits 1 if any func is below
```

`ta_bench_stream` is **C only**. For the Rust, Java and C# streaming tiers,
`scripts/stream_ab.py` A/Bs `update` (or `peek`) per bar — or `open`, which times
the whole warm-up instead of one bar — between the working tree
and a git revision — same generated harness compiled against two copies of the
library, interleaved rounds with alternating arm order, every streaming function
so the untouched ones are the control. It reads only the generated Rust crate,
Java fragments and C# library (no ta_abstract, no servers, no C build) and
derives every call from the emitted signatures, so adding an indicator needs no
edit there.

Every arm asserts a **floor** (`FUNC_FLOOR`, 170 against 176 today) on how many
functions it parsed. That is not a style check: #278 recased the Rust and Java
stream APIs and both arms' regexes kept the old spelling, so each parsed zero —
Rust's until `c308e789`, Java's for a further day. Nothing but a person trying
to use the tool noticed either. If an arm dies saying it is under the floor, the
generated API moved: fix the parser, do not lower the floor.

The **C# arm** pins `TieredCompilation` off in the harness project, so every
method is fully optimised on its first call. The cost, stated rather than
hidden: dynamic PGO is off with it, so the C# ns columns are static-opt numbers
rather than what a long-lived process settles at. Both arms get the same
treatment, so the change column — the output — is unaffected, and the ns columns
were never comparable across invocations anyway. Note also that `TALib.csproj`
sets `TreatWarningsAsErrors`, so a `--base` from an older revision has to compile
*warning-clean* under today's SDK, a higher bar than the other two arms impose.

There is still **no C# row in `ta_bench`** — `ta_bench --mode=open` answers
`unsupported_mode` for it — so batch-vs-stream ratios remain a three-language
table.

```bash
scripts/stream_ab.py --base=origin/dev                                  # all three
scripts/stream_ab.py --base=HEAD~1 --lang=rust --call=peek --mark=MIN,MAX
scripts/stream_ab.py --base=origin/dev --lang=csharp --call=peek
scripts/stream_ab.py --base=origin/dev --call=open --mark=BBANDS,STDDEV   # the Open tier
```

Current shape: median ~1.6x, but **~25 stream slower than
batch** and another ~50 sit under 1.5x. Recursive/multi-stage state wins big
(`HT_TRENDLINE` ~24x, `TRIX`/`TEMA` ~16x); window-recomputers and stateless
patterns lose (`AVGDEV`, `MAVP`, `MIDPRICE`, `WILLR`, CDL*) because the handle
buys nothing and costs indirection. Those losers overlap the rolling-extremum
family — see the corpus note below.

`--min-ratio` is a cliff detector, not a quality bar: run to run the worst ratio
moves 0.42–0.50 and the worst function's *name* changes, so a threshold near 1.0
just flaps. 0.35 has headroom while still failing on a real regression.

## Benchmark input corpus

Some indicators have input-dependent cost, so which series you measure on is
part of the measurement. `src/tools/ta_bench/bench_corpus.h` holds the corpus —
one deterministic generator, shared by `ta_bench`, `ta_bench_direct` and the
generated `ta_bench_cg` / `ta_bench_stream`. Select a class with `--shape=`:

```bash
cd bin && ./ta_bench --list-shapes            # the input classes and what each reaches

# random walk (default: the historical seed-42 series) and GBM — the acceptance gate
./ta_bench --language=cref,c --function=WILLR --shape=randwalk --iters=500
./ta_bench --language=cref,c --function=WILLR --shape=gbm      --iters=500

# alternating trend/chop legs — the class rolling min/max degrades on
for s in trend-chop-0.5p trend-chop-1p trend-chop-2p trend-chop-4p; do
  ./ta_bench --language=cref,c --function=WILLR --shape=$s --period=30 --iters=500
done
```

The rolling min/max caches the window extremum and rescans the window when that
extremum is the bar dropping out of it, so its cost depends on how often that
happens. On a zero-drift walk the rate decays as ~1/sqrt(period); on a trending
leg it is set by the drift/noise ratio instead and barely moves with the period,
so the two separate further the longer the window (1.1x the rescan rate at
period 14, 3x at period 200). `randwalk` alone cannot see that — issue #147.

The tail shapes are not peers: `constant` is the worst case at `2*(period-1)`
comparisons per bar, exactly twice `mono-up`/`mono-down`. Flat input pins both
extrema because the rescan compares with strict `>`/`<` and leaves the cached
index on `trailingIdx`, so the `>=`/`<=` fast-path arms never run; a monotone
ramp pins only one of the two.

**Which tier that still describes** matters, because #147 replaced half of it.
The batch tier of MIN, MAX, MINMAX, MIDPOINT, MIDPRICE and WILLR is now a Van
Herk / Gil-Werman block scan: branchless, a fixed number of comparisons per bar
at any period, input-independent. So for those six, `constant`, `mono-*` and
`trend-chop-*` all cost the same through `ta_bench --language=c` (the batch
call) and the shape sweep says nothing about them. The rescan — and everything
above — is still what STOCH and STOCHF run, and still what the *streaming* tier
of all six runs, which is what `ta_bench --shape=... --mode=open` and
`ta_bench_stream`'s `update_ns` measure. Reach for the shape sweep when the arm
under test is one of those; for the six functions' batch arm it is inert.

`--shape` is opt-in and `randwalk` reproduces the pre-corpus series bit for bit,
so a default run costs and measures exactly what it did before. `--seed` picks
the stream; `--regime-period` the window the trend/chop regime length is relative
to (defaults to `--period` when given, else 14); `--trend-strength` the trend-leg
drift in per-bar standard deviations (default 0.5 — sweep it to see how the cost
responds to trend/noise). `--verify-corpus` checks every shape is reproducible
and produces valid OHLC, at the `--points` you pass it.

`--list-shapes` groups the classes by what they are for, and the grouping
matters. The rescan rate depends only on the *rank order* of the bars, so
`randwalk-lo`, `randwalk-hi` and `gbm` cannot move it however much they change
the magnitudes — measured within 1% of `randwalk` at period 14/30/200. They are
controls, useful for numerical-conditioning questions (deadbands, cancellation,
ratio-based indicators), not stressors. Only `trend-chop-*` varies the rescan
rate; `mono-*` and `constant` are the analytic tail.

One documented exemption in `--verify-corpus`: the walk family floors `low` at
1.0 but leaves `close` unclamped, so `low <= min(open,close)` fails on 32 bars of
`randwalk` at n=100000 (11 with a negative close). That is inherited from the
pre-corpus generator and is preserved deliberately — clamping `close` would break
the byte-for-byte reproduction of the historical seed-42 series, which matters
more on a timing-only corpus. Every other predicate holds for every shape.

The corpus is timing-only — it is never hashed and is unrelated to
`fuzz_data.h`, whose `FUZZ_*` shape list is iterated by `--fuzz-064` /
`--xlang-hash`. Keep it that way: adding a shape there changes what those gates
compare (see the note at `test_variants.c:148`).

