Packs
4 packs@jiachen-t-wang
Curation Bench
Curation Bench from Jiachen-T-Wang/curation-bench-pro.
99 skills · pack
@gtynnn060110-hash
Environment
Environment from gtynnn060110-hash/continual-skill-bench-final.
7 skills · pack
curated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · pack
@owl-listener
Prototyping Testing
Prototyping and testing skills: wireframe specs, usability heuristics, heuristic evaluations, accessibility audits, A/B test design, and benchmark analysis.
8 skills · pack
Results for “bench”
12 skillsscore
Audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
3
bis-eval
Benchmarks energy-function-based safe control algorithms on the BIS (Benchmark of Interactive Safety) dataset, scoring safety, efficiency, and hybrid performance in human-robot and robot co-working scenarios.
3
art-eval
Benchmarks medical AI agents on synthetic EHR tasks, measuring success rates for data retrieval, temporal aggregation, and threshold-based conditional logic with exact-match scoring.
3
cab-eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
pytdc
Access AI-ready drug discovery datasets and benchmarks from Therapeutics Data Commons, covering ADME, toxicity, drug-target interactions, and molecular generation with standardized splits and evaluation metrics.
30.2k · bundle
More results
caa-eval
Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
3
benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
0
bmad-ml-lab-meeting
Run an AI Lab division meeting with research and build agents (no AI Startup agents). Use when the user requests to "start a lab meeting", "convene the lab", or "run a sprint retrospective for the lab".
0 · bundle
benchmark-methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
google-maps-reviews-api-skill
Extract structured review data from Google Maps search results using the BrowserAct API, enabling local business analysis, reputation monitoring, and competitive benchmarking.
3.7k · bundle
autoresearch
Guides users through defining goals, metrics, and scope, then runs an autonomous loop of code changes, testing, measuring, and keeping or discarding results for any programming task with a measurable outcome.
36.2k
psnr
Evaluates the trade-off between file size reduction and image fidelity when encoding radio astronomy data using JPEG2000, benchmarking both lossless and lossy compression modes to determine the compression ratio at which visual artifacts first appear.
3