Plugins
3 pluginscurated
ML Model Lifecycle
Train, evaluate, and deploy a production ML system with monitoring.
10 skills · plugin
@owl-listener
Visual Critique
Visual critique skills: hierarchy analysis, brand consistency checks against mood/voice/tokens, composition evaluation, and typography audits — with a /critique-screen command that compiles a prioritised fix list.
7 skills · plugin
@alirezarezvani
Agenthub
Multi-agent collaboration — spawn N parallel subagents that compete on code optimization, content drafts, research approaches, or any task that benefits from diverse solutions. 7 slash commands (/hub:init, /hub:spawn, /hub:status, /hub:eval, /hub:merge, /hub:board, /hub:run), agent templates, DAG-based orchestration, LLM judge mode, message board coordination.
8 skills · plugin
Results for “l-eval”
87 skillsIvx Cf Sid Evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra. Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop with evaluation gates.
0 · bundle
Ivx Sid Evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra (global). Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop. Works in any workspace; bootstraps EVALS.md from ~/.cursor/skills/sid-orchestra/templates if missing.
0 · bundle
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
0
Eval Driven Dev
Build automated evaluation pipelines for Python LLM applications using real LLM calls and structured test datasets.
36.2k · bundle
Heuristic Evaluation
Conduct expert heuristic evaluations of digital interfaces using Nielsen's 10 usability heuristics and domain-specific criteria.
1.7k
Llava Critic Learning To Evaluate Multimodal Models Arxiv 24
LLaVA-Critic: Learning to Evaluate Multimodal Models
6
More results
Setup
Set up a new autoresearch experiment interactively. Collects domain, target file, eval command, metric, direction, and evaluator.
3
Ml Engineering
Enforces rigorous ML modeling, feature engineering, training, and evaluation standards at principal-engineer level.
0
Critique
Multi-perspective dialectical reasoning with cross-evaluative synthesis. Spawns parallel evaluative lenses (STRUCTURAL, EVIDENTIAL, SCOPE, ADVERSARIAL, PRAGMATIC) that critique thesis AND critique each other's critiques, producing N-squared evaluation matrix before recursive aggregation. Triggers on /critique, /dialectic, /crosseval, requests for thorough analysis, stress-testing arguments, or finding weaknesses. Implements Hegelian refinement enhanced with interleaved multi-domain evaluation and convergent synthesis.
0 · bundle
Skill Creator
Guides the creation, modification, and evaluation of agent skills, including running benchmarks and optimizing descriptions for better triggering.
253 · bundle
Continuous Learning
Automatically evaluates Claude Code sessions to extract reusable patterns and save them as learned skills.
226k · bundle
Run
Execute the full AgentHub competition lifecycle in a single command: initialize, capture baseline, spawn agents, evaluate results, and merge the winner.
20.4k
Setup
Set up a new autoresearch experiment interactively. Collects domain, target file, eval command, metric, direction, and evaluator. Use when the user runs /ar:setup or asks to start optimizing a file with the autoresearch loop.
11
MCP Builder
Guides the creation of high-quality MCP servers that let LLMs interact with external services through well-designed tools, covering planning, implementation, testing, and evaluation.
253 · bundle
Gan Style Harness
Uses a multi-agent generator-evaluator feedback loop to build high-quality applications from a single prompt, inspired by GANs and Anthropic's harness design.
226k
Run
Run a single experiment iteration. Edit the target file, evaluate, keep or discard.
3
Blockchain
Understand blockchain technology, interact with smart contracts, and evaluate when distributed ledgers solve real problems.
1 · bundle
Run
One-shot lifecycle command that chains init → baseline → spawn → eval → merge in a single invocation.
0
Speckit Review Tests
Test coverage quality analysis — behavioral coverage, critical gap identification, test resilience evaluation.
11
Run
One-shot lifecycle command that chains init → baseline → spawn → eval → merge in a single invocation.
0
Run
One-shot lifecycle command that chains init → baseline → spawn → eval → merge in a single invocation.
3
Deobfuscating Javascript Malware
Deobfuscates malicious JavaScript code used in web-based attacks, phishing pages, and dropper scripts by reversing encoding layers, eval chains, string manipulation, and control flow obfuscation to reveal the original malicious logic.
24.6k · bundle
Scientific Critical Thinking
Evaluate scientific claims and evidence quality by assessing experimental design, identifying biases and confounders, and applying evidence grading frameworks like GRADE and Cochrane Risk of Bias.
30.2k · bundle
Levi Skill
利威尔(少年漫)认知与表达框架(压缩蒸馏):兵长洁癖战力、矮个子反差、残酷抉择 触发:进击的巨人 等。虚构;非仇恨教唆
9 · bundle
Acceptance Eval
Acceptance Eval
18 · bundle
Eval Harness
Provides a formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles to define pass/fail criteria, measure reliability with pass@k metrics, and create regression test suites.
226k
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
1
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
1
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
0
Cto Advisor
Technical leadership guidance for engineering teams, architecture decisions, and technology strategy. Use when assessing technical debt, scaling engineering teams, evaluating technologies, making architecture decisions, establishing engineering metrics, or when user mentions CTO, tech debt, technical debt, team scaling, architecture decisions, technology evaluation, engineering metrics, DORA metrics, or technology strategy.
0 · bundle
Outcome Eval
Outcome Eval
18 · bundle
Eval Harness
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles
0
Rfp Writer
Write RFPs with requirements, evaluation, and timeline. TRIGGERS - Use when user needs help with rfp-writer related tasks.
22
Arbitrage Scanner
Detect and evaluate arbitrage opportunities across sportsbooks and prediction markets. Calculate guaranteed-profit scenarios, middle bets, and cross-platform price discrepancies. Use when comparing odds across books, calculating arb percentages, evaluating middle opportunities, or building odds-comparison workflows. Also trigger for 'arb bet', 'sure bet', 'arbitrage', 'odds comparison', 'middle bet', 'risk-free bet', or 'line shopping'.
0
J Rig
>- Skill Refiner, the eval-guided improvement loop for SKILL.md files. Runs the bootstrap, score, propose, apply, and status cycle as a thin wrapper over the published @intentsolutions/refiner CLI, proposing safe, minimal, bounded SKILL.md edits and accepting an edit only when a held-out eval score strictly improves with no regression on any other case. Ships a 3-layer cost-tiered hook architecture (sinker, line, hook) that gates skill quality at edit time, end of turn, and commit time. Use when improving an existing skill, refining a SKILL.md against measured behavior, bootstrapping an eval set for a skill, or gating skill edits before they ship. Trigger with "/j-rig", "refine this skill", "bootstrap an eval set", "propose a skill edit", "promote the candidate", or "skill refiner status".
2
Idea Evaluation
Score an unbuilt business idea on desirability, viability, feasibility, distribution wedge, why-now, founder-market-fit, market size, alternatives, defensibility, capital intensity, and regulatory/ethical risk — and return a GO / ITERATE / KILL verdict with kill criteria and a next kill test. Load when the user asks to evaluate a business idea, score a startup idea, screen an idea, decide whether to pursue this venture, do an idea review, or says "is this a good business idea", "should I build this", "evaluate this startup", "screen this idea", "go/no-go on this idea", "kill or pursue". Sub-skill of `venture-exploration`. Calls `fermi` for sizing, `assumption-mapping` for hidden beliefs, optional `pre-mortem` / `adversarial-hat` for high-stakes ideas. Does NOT evaluate built products — for that use `reality-check`.
3 · bundle