Telemetry Event Analyst
Workflow
- Identify baseline, candidate, benchmark version, harness profiles, and trial counts.
- Read
references/metrics-policy.mdbefore interpreting tokens, context, or cost. - Verify comparability: same scenario, initial state, model policy, permissions, environment, and adapter capabilities.
- Separate functional failure from performance regression and telemetry gaps.
- Report absolute values, deltas, source/confidence, variance, and outliers.
- Link conclusions to event IDs, artifact paths, or metric observation IDs.
Rules
- Do not compare native token usage with estimated token usage as if they were equivalent.
- Do not hide inconclusive, failed, or missing-data trials.
- Mark runs as
limitedornot_comparablewhen capabilities differ materially. - Prefer median and distribution over single-run anecdotes.
- Keep raw evidence available for audit.
Analysis Outputs
Produce concise findings with:
- result summary;
- comparability status;
- biggest deltas;
- data quality gaps;
- likely causes;
- recommended next measurement or fix.