deepscientist-eval
DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively — Weng et al. (2025) (arXiv:2509.26603, 2025)
What this evaluates
Evaluates an autonomous AI research system's ability to progressively advance state-of-the-art methods across three distinct AI tasks: agent failure attribution, LLM inference acceleration, and AI text detection. It also assesses the scientific quality of the AI-generated research papers through automated and human peer review.
Datasets
- Who&When benchmark — total ?; splits: test (-1)
- MBPP — total ?; splits: test (-1)
- AI Text Detection dataset — total ?; splits: test (-1)
Metrics
Accuracy(primary) — range: percent- Percentage of correctly attributed agent failures in handcraft and algorithm-generated settings.
Tokens/second— range: other- Throughput metric measuring the number of tokens generated per second during LLM inference.
AUROC(primary) — range: [0, 1]- Area under the Receiver Operating Characteristic curve for binary classification of human vs. AI-generated text.
Latency— range: other- Inference time measured in milliseconds.
Rating— range: [0, 5]- Average score (1-5 scale) assigned by reviewers across Soundness, Presentation, Contribution, and overall Rating dimensions.
Input / output format
Input: Baseline SOTA method code and execution logs, task descriptions, and human supervision feedback for task advancement; generated research papers for quality evaluation.
Output: New method code/algorithm, performance metrics (Accuracy, Tokens/second, AUROC, Latency), and full research paper text.
Scoring recipe
For tasks:
acc = (correct_predictions / total_samples) * 100
auroc = compute_roc_auc(y_true, y_pred)
tokens_sec = total_tokens / total_inference_time
latency = mean_inference_time_ms
For papers:
rating = mean(review_scores['Soundness'], review_scores['Presentation'], review_scores['Contribution'], review_scores['Rating'])
Common pitfalls
- The system's low scientific progress rate (1-3%) reflects expected exploration behavior, not system failure.
- Distinguishing genuine scientific discoveries from engineering optimizations requires careful protocol adherence.
- Human evaluation relies on a small committee (3 reviewers), limiting statistical generalizability.
Evidence (verbatim from paper)
As shown in Figure [1] and 3, this entire discovery trajectory demonstrates DeepScientist’s ability for advancing frontier-pushing scientific findings progressively, establishing a new SOTA with a 7.9% higher AUROC while also doubling the inference speed.
Citation
@misc{weng2025deepscientist,
title={DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively},
author={Weng et al. (2025)},
year={2025},
note={arXiv:2509.26603}
}
- arXiv: 2509.26603