Plugins
2 pluginscurated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · plugin
@owl-listener
Prototyping Testing
Prototyping and testing skills: wireframe specs, usability heuristics, heuristic evaluations, accessibility audits, A/B test design, and benchmark analysis.
8 skills · plugin
Results for “benchmark”
64 skillsResults
Documents benchmark results for a multi-agent system and provides commands to run benchmark suites.
0 · bundle
Finance Metrics Quickref
Look up SaaS finance metrics, formulas, and benchmarks fast. Use when you need a quick metric definition, formula, or benchmark during analysis.
5.6k
Skill Benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.
542
Dataperf Benchmarks For Data Centric AI Development Arxiv 22
DataPerf: Benchmarks for Data-Centric AI Development
6
Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
Ivx Cf Benchmarking
Design and execute performance benchmarks for models and systems. Use when measuring throughput, latency, or comparing system variants.
0 · bundle
More results
Skill Creator
Guides the creation, iterative improvement, and evaluation of agent skills, including drafting, testing, benchmarking, and optimizing descriptions.
2 · bundle
Skill Creator
Create new skills, modify existing ones, and measure their performance through iterative evaluation and benchmarking.
158k · bundle
Benchmark
Measures performance baselines, detects regressions before and after PRs, and compares stack alternatives.
1
Microbenchmarking
Create, run, configure, and review BenchmarkDotNet microbenchmarks for .NET code, covering project setup, comparison strategies, and cost-aware execution.
4k · bundle
Tilegym Adding Cutile Kernel
Add a new cuTile GPU kernel operator to TileGym, covering dispatch registration, backend implementation, exports, tests, and benchmarks.
2.2k · bundle
Skill Creator
Guides the creation, modification, and evaluation of agent skills, including running benchmarks and optimizing descriptions for better triggering.
253 · bundle
Dbs Benchmark
Helps find and analyze competitors to imitate using a five-filter method, focusing on profitability and feasibility while eliminating personal bias.
Benchmark
Performance regression detection using the browse daemon. (gstack)
0
Golang Testing
Write reliable, maintainable Go tests using table-driven tests, subtests, benchmarks, fuzzing, and golden files following TDD methodology.
226k
Skill Creator
Create new skills, modify and improve existing skills, and measure skill performance through iterative evaluation and benchmarking.
1.5k · bundle
Benchmark
Measure performance baselines, detect regressions before and after PRs, and compare stack alternatives using browser, API, and build benchmarks.
226k
Golang Testing
Write reliable Go tests using table-driven tests, subtests, benchmarks, fuzzing, and coverage, following TDD with idiomatic patterns.
0
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
1
Pymoo
Solves single- and multi-objective optimization problems with NSGA-II/III, MOEA/D, and other evolutionary algorithms, including constraint handling, Pareto front analysis, and benchmark problems.
253 · bundle
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
2
Benchmark
Performance baseline measurement and regression detection. Use when measuring perf before/after a PR, setting up baselines, investigating "feels slow" reports, validating launch performance targets, or comparing your stack against alternatives.
0
Performance Benchmarking
Use when evaluating, measuring, or comparing the performance of systems, functions, or services. This skill provides a framework for establishing baselines, measuring performance, and validating that changes meet performance requirements.
0
Pymoo
Solve single and multi-objective optimization problems using NSGA-II/III, MOEA/D, and other evolutionary algorithms with customizable operators, constraint handling, and benchmark problems.
30.2k · bundle
Benchmark
Performance regression detection using the browse daemon. Establishes baselines for page load times, Core Web Vitals, and resource sizes. Compares before/after on every PR. Tracks performance trends over time. Use when: "performance", "benchmark", "page speed", "lighthouse", "web vitals", "bundle size", "load time". (gstack) Voice triggers (speech-to-text aliases): "speed test", "check performance".
0
Performance Optimization
Measures application performance first, identifies bottlenecks, and applies targeted fixes for frontend and backend systems.
69.5k
Slurm Apptainer
Curation-Bench Protocol (skill-grounded)
6 · bundle
Testing Quality Assurance
Coordinates quality assurance workflows by routing testing tasks to specialized sub-skills for API testing, performance benchmarking, test analysis, tool evaluation, and process optimization.
2 · bundle
Kpi Tracker
Set up and track KPI dashboards with targets, actuals, and trend analysis. TRIGGERS - Use when user wants to define KPIs, track metrics, or create a scorecard.
22
Kpi Tracker
Set up and track KPI dashboards with targets, actuals, and trend analysis. TRIGGERS - Use when user wants to define KPIs, track metrics, or create a scorecard.
3
Nocaps Novel Object Captioning At Scale Arxiv 1812 08658v2
Nocaps: Novel Object Captioning at Scale
6
Priority Decision System
Prioritize product work with explicit criteria, scoring, tradeoffs, and decision rationale.
0
Outcome Map
Create product objectives, key results, initiatives, guardrails, and review cadence.
0
Performance Profiler
Systematically profile Node.js, Python, and Go applications to identify CPU, memory, and I/O bottlenecks, generate flamegraphs, analyze bundle sizes, optimize database queries, and run load tests with k6 and Artillery.
20.4k · bundle
Performance
Use for identifying and resolving performance bottlenecks.
3
Social Listener
Monitor social platforms and forums for brand mentions, category discussions, and competitor activity
2 · bundle