Packs

4 packs

Results for “bench”

12 skills
More results
qhjqhj00
caa-eval
Benchmarks large audio-language models against adversarial audio attacks using the CAA dataset, computing WER, ROUGE-L, cosine similarity, and coherence scores to assess robustness in conversational settings.
3
rajanthar
benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
0
sirnosh
bmad-ml-lab-meeting
Run an AI Lab division meeting with research and build agents (no AI Startup agents). Use when the user requests to "start a lab meeting", "convene the lab", or "run a sprint retrospective for the lab".
0 · bundle
affaan-m
benchmark-methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
browser-act
google-maps-reviews-api-skill
Extract structured review data from Google Maps search results using the BrowserAct API, enabling local business analysis, reputation monitoring, and competitive benchmarking.
3.7k · bundle
github
autoresearch
Guides users through defining goals, metrics, and scope, then runs an autonomous loop of code changes, testing, measuring, and keeping or discarding results for any programming task with a measurable outcome.
36.2k
qhjqhj00
psnr
Evaluates the trade-off between file size reduction and image fidelity when encoding radio astronomy data using JPEG2000, benchmarking both lossless and lossy compression modes to determine the compression ratio at which visual artifacts first appear.
3