Plugins
2 pluginscurated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · plugin
@owl-listener
Prototyping Testing
Prototyping and testing skills: wireframe specs, usability heuristics, heuristic evaluations, accessibility audits, A/B test design, and benchmark analysis.
8 skills · plugin
Results for “benchmark”
201 skillsSkill Creator
Guides users through creating, editing, and optimizing agent skills, including drafting, testing, evaluating, and improving skill performance.
19 · bundle
Pytdc
Therapeutics Data Commons. AI-ready drug discovery datasets (ADME, toxicity, DTI), benchmarks, scaffold splits, molecular oracles, for therapeutic ML and pharmacological prediction.
5 · bundle
Pymoo
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
0 · bundle
Hmmsim
Use when you need to characterize score distributions of a profile HMM on random sequences, such as calibration checks, benchmarking, or filter-behavior experiments.
0 · bundle
Paw Pa Research
Proposal research workflow that matches local case studies and gathers web evidence into an HTML research dossier. Use when the user needs proposal research, client intel, tech stack discovery, pricing benchmarks, competitive context, or case-study matching for a brief. Triggers: 'research this proposal', 'build a research dossier', 'match case studies', 'find pricing benchmarks', 'client intel for', 'what tech does X use'.
85 · bundle
Aide
AIDE file integrity monitoring reference. Database initialization, integrity checks, update workflow, aide.conf configuration, selection rules, report parsing, and production deployment with CIS benchmark compliance.
12 · bundle
Pymoo
Multi-objective optimization framework. NSGA-II, NSGA-III, MOEA/D, Pareto fronts, constraint handling, benchmarks (ZDT, DTLZ), for engineering design and optimization problems.
5 · bundle
Compete
Researching competitors and shaping positioning: feature matrices, SWOT, benchmarking, positioning maps, battle cards, win/loss, LLM brand visibility. Research only — use for strategy, not code.
65 · bundle
Ivx Cf Evaluation
Design and implement evaluation harnesses for models, agents, and code. Use when creating benchmarks, designing eval metrics, or comparing system outputs.
0 · bundle
Polar Strategy
POLAR v2.0 — ETH Alpha Hunter. The patience benchmark. Thesis exit permanently removed. Scanner enters, DSL exits. +19.8% ROE trades after removing thesis exit.
1 · bundle
Nemo Evaluator Plugin
Run evaluation tasks against a NeMo Platform server using the Evaluator plugin CLI and Python SDK.
2.2k · bundle
Arize Experiment
Creates, runs, and analyzes Arize experiments for evaluating and comparing model performance using the ax CLI.
36.2k · bundle
Skill Creator
Guides users through creating, refining, and evaluating agent skills, including drafting, testing, and optimizing descriptions for better triggering.
559 · bundle
Agent Eval
Compares coding agents head-to-head on reproducible tasks, measuring pass rate, cost, time, and consistency.
1
Autogpt Agents
Build, deploy, and manage continuous AI agents using a visual workflow editor or development toolkit.
10.4k · bundle
Overnight Eval
Launches long-running evaluation batches in isolated tmux sessions with pre-flight verification, monitoring, and post-flight analysis for unattended runs.
0
Pymoo
Framework de otimização multi-objetivo. NSGA-II, NSGA-III, MOEA/D, frentes de Pareto, tratamento de restrições, benchmarks (ZDT, DTLZ), para problemas de design e otimização em engenharia.
10 · bundle
Pytdc
Therapeutics Data Commons. Conjuntos de dados prontos para IA em descoberta de drogas (ADME, toxicidade, DTI), benchmarks, divisões de scaffold, oráculos moleculares, para ML terapêutico e predição farmacológica.
10 · bundle
Sequence Analyzer
Analyzes email sequence performance metrics. Evaluates open rates, click rates, reply rates, and conversion by step. Identifies drop-off points, benchmarks against industry averages, and recommends optimizations.
2 · bundle
Pricing Strategy
Design, optimize, and communicate SaaS pricing — tier structure, value metrics, pricing pages, and price increase strategy.
20.4k · bundle
Agent Eval
Compare coding agents head-to-head on reproducible tasks with pass rate, cost, time, and consistency metrics.
226k
Tw Lift
Delivers measurement-driven performance optimization for latency, throughput, memory, and tail behavior, with correctness preservation and regression guards.
7 · bundle
Skill Creator
Guides users through creating, editing, and optimizing agent skills, including drafting, testing, evaluating, and improving skill descriptions for better triggering.
1 · bundle
Turbo Source Benchmark
Use in TURBO mode when Codex should compare top current solutions before choosing architecture, UI components, 3D stack, mobile stack, SEO/PPC strategy, or programming patterns.
1 · bundle
LLM Evaluation
Implement comprehensive evaluation strategies for LLM applications using automated metrics, human feedback, and benchmarking. Use when testing LLM performance, measuring AI application quality, or establishing evaluation frameworks.
0
Eval Run
Launches a model evaluation batch with parameter collection, pre-flight checks, execution, and post-run analysis for interactive or foreground runs.
0
Autoresearch Prep
Scaffolds a program.md research program for autoresearch by auto-detecting codebase signals and interviewing for missing details.
1 · bundle
Bmad Ml Cypher
Dataset analysis and data quality specialist. Use when the user asks to talk to Cypher, requests the data detective, or needs dataset assessment, bias analysis, and benchmark evaluation.
0 · bundle
Pump Testing
Multi-language test infrastructure for the Pump SDK — Rust unit/integration/security/performance tests, TypeScript Jest tests, Python fuzz tests, shell test orchestration, Criterion benchmarks, and CI quality gates.
9
Arbor
Run autonomous optimization loops that iteratively improve artifacts against evaluators using hypothesis tree refinement, without overfitting.
30.2k · bundle
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
0
Benchmark
Performance baseline measurement and regression detection. Use when measuring perf before/after a PR, setting up baselines, investigating "feels slow" reports, validating launch performance targets, or comparing your stack against alternatives.
0
Pump Testing
Design and run Pump.fun SDK test infrastructure across Rust, TypeScript, Python, and Bash, including unit tests, integration tests, security tests, fuzzing, shell orchestration, Criterion benchmarks, coverage, and CI gates.
0
Cost
Evaluates a containerized framework for deploying distributed big data workloads, measuring execution time and cloud cost scaling from four to eight nodes.
3
Bmad Eval Runner
Run a skill's evals and report results. Use when the user wants to evaluate a skill, run evals, benchmark a skill, validate triggers, optimize a description, or grade skill outputs.
1 · bundle
Performance Benchmarking
Use when evaluating, measuring, or comparing the performance of systems, functions, or services. This skill provides a framework for establishing baselines, measuring performance, and validating that changes meet performance requirements.
0