voice-ai-platform-eval
Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms — Andrés et al. (2025) (arXiv:2511.04133, 2025)
What this evaluates
This benchmark evaluates the quality of commercial voice AI testing platforms across two independent dimensions: simulation quality (how realistically platforms generate test conversations) and evaluation accuracy (how accurately platforms assess conversation quality against human ground truth). It probes whether automated testing systems can reliably replace human quality assurance in high-stakes voice AI deployments.
Datasets
- Custom Voice AI Testing Benchmark — total ?; splits: simulation_test_cases (-1), evaluation_golden_set (-1)
Metrics
F1-score(primary) — range: [0, 1]- Harmonic mean of precision and recall: 2 * (precision * recall) / (precision + recall). Calculated by comparing platform predictions against human consensus ground truth (majority vote for binary, median for continuous metrics).
Elo rating— range: [0, 1]- Adapted from chess rankings, updates platform scores based on pairwise comparison outcomes and opponent strength. Points gained/lost depend on the expected win probability derived from score differences.
League score— range: [0, 1]- Counts wins as 1 point and losses as 0 point per pairwise comparison, providing a transparent frequency-of-preference metric.
Input / output format
Input: Simulation: Structured scenario description and persona specifications. Evaluation: Audio transcripts of conversations between a testing agent and a subject agent, provided as standardized transcript files via an observability API.
Output: Simulation: Human pairwise choices (A/B) for specific quality metrics. Evaluation: Platform-generated binary/continuous scores or labels for each metric, compared to human consensus.
Scoring recipe
def score_simulation(pairwise_results):
league = {p: 0 for p in platforms}
for a, b, winner in pairwise_results:
league[winner] += 1
elo = update_elo_ratings(league, opponent_strengths)
return pca_aggregate(elo, league)
def score_evaluation(platform_preds, human_gt):
tp = sum(1 for p, g in zip(platform_preds, human_gt) if p == 1 and g == 1)
fp = sum(1 for p, g in zip(platform_preds, human_gt) if p == 1 and g == 0)
fn = sum(1 for p, g in zip(platform_preds, human_gt) if p == 0 and g == 1)
precision = tp / (tp + fp) if (tp + fp) > 0 else 0
recall = tp / (tp + fn) if (tp + fn) > 0 else 0
return 2 * (precision * recall) / (precision + recall) if (precision + recall) > 0 else 0
Common pitfalls
- Confusing simulation quality (testing agent generation) with evaluation accuracy (platform's assessment of conversations), despite the framework explicitly isolating them.
- Assuming platform simulation strength correlates with evaluation strength; the observability API ensures platforms evaluate identical transcripts independently of their simulation capabilities.
- Using absolute ratings instead of pairwise comparisons for simulation quality, which the paper notes suffers from subjective interpretation and low inter-rater reliability.
Evidence (verbatim from paper)
For binary metrics, we calculate precision, recall, F1-score, and accuracy across all platform-metric-recording combinations. The F1-score proves particularly valuab
Citation
@misc{andres2025testing,
title={Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms},
author={Andrés et al. (2025)},
year={2025},
note={arXiv:2511.04133}
}
- arXiv: 2511.04133