LLM Benchmark Analyst

search and analyze llm benchmark results within a fixed benchmark universe, then produce evidence-based model strength and weakness reports or domain-leader summaries. use when comparing a model across benchmarks, ranking the best models by domain, explaining what a benchmark measures, checking predecessor-vs-current progress, or writing benchmark reports that must prioritize exact model version, evaluation date, benchmark variant, score semantics, sub-scores, and benchmark defect warnings. works with browser, web, and multimodal extraction for text, table, canvas, or image-only leaderboards.

dvcrn Updated 32 repo stars

File contents

dvcrn/openclaw-skills-marketplace/tree/main/plugins/chekhovin--llm-benchmark-analyst/skills/llm-benchmark-analyst commit 107e4f3499

Frequently asked questions

npx skillmds@latest add dvcrn/llm-benchmark-analyst