Treereview Peer Review Eval

This benchmark evaluates an LLM's ability to perform deep, structured scientific peer review by generating comprehensive reviews and actionable feedback comments. It probes the model's capacity for hierarchical question decomposition, context-aware analysis of long documents, and alignment with human reviewer judgments across multiple quality dimensions. Use when the user wants to benchmark on TreeReview Benchmark (ICLR-2024, NeurIPS-2023, Nature Communications), or asks about evaluating this task. Reports Overall Quality (LLM-as-Judge).

qhjqhj00 5f323b6 4.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/treereview-peer-review-eval commit 5f323b6dab

Frequently asked questions

npx skillmds add qhjqhj00/treereview-peer-review-eval