Benchmark deep research agents across factual, quality, and process dimensions with MiroEval

Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.

agentskillexchange Updated 28 repo stars

File contents

Benchmark deep research agents across factual, quality, and process dimensions with MiroEval

Score deep research agents on benchmark tasks using factual verification, report-quality scoring, and process evaluation before model or workflow changes ship.

Prerequisites

Python, uv, model result JSON, required API keys for judge and retrieval services

Installation

No source-backed install or usage instructions could be extracted automatically. Review the upstream project before running this skill in a sensitive workflow.

Documentation

Source

agentskillexchange/skills/tree/main/skills/benchmark-deep-research-agents-across-factual-quality-and-process-dimensions-with-miroeval commit b3bacb980b

Frequently asked questions

npx skillmds@latest add agentskillexchange/benchmark-deep-research-agents-across-factual-quality-and-pr