Evaluate Plugin

Skill evaluation and benchmarking - test skill effectiveness with behavioral eval cases, grade results, and track quality improvements

by @laurigates 7 skills

Skills in this plugin

7
  1. Evaluate Skill · laurigates
    Evaluate a skill by running test cases and grading results. Use when testing whether a skill produces correct guidance, validating improvements, or benchmarking before release.
    0 installs
  2. Evaluate Matrix · laurigates
    Cross-model skill evals with real execution and grading — the executability gate. Use when checking whether a weak model can actually do a skill, not just comprehend it.
    0 installs
  3. Evaluate Report · laurigates
    View evaluation results and benchmark reports for a skill or plugin. Use when reviewing past eval results, comparing benchmark runs, or tracking quality trends.
    0 installs
  4. Evaluate Improve · laurigates bundle
    Suggest improvements to SKILL.md content, descriptions, or tool config from eval results. Use when raising pass rates, fixing triggering, or iterating on a skill after evaluation.
    0 installs
  5. Evaluate Legibility · laurigates bundle
    Cold-read a SKILL.md with a zero-context agent reader to check its intent is legible. Use when validating whether a skill says clearly when to invoke it and how to start.
    0 installs
  6. Evaluate Plugin Batch · laurigates
    Batch evaluate every skill in a plugin and produce a plugin-level report. Use when auditing an entire plugin's quality or validating before a release.
    0 installs
  7. Evaluate Context Engineering · laurigates bundle
    Context-engineering audit (C1-C6) of skills and always-loaded rules. Use when trimming context bloat, auditing a plugin, or after a new Claude model ships.
    0 installs