Results for “scoring-rubric”
60 skillsrubric-creator
Creates assessment rubrics with criteria, performance levels, and scoring guidelines for various assignments
6 · bundle
eval-rubric-design
Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "evaluation rubric for", "how do I measure quality of". Sub-skill of eval-output orchestrator.
3 · bundle
advanced-evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
forter-agentic-readiness-audit
Audits a website against the Forter Agentic Readiness Guide by running 25 weighted rubrics, scoring each guideline, and producing a prioritized fix report.
106 · bundle
career
Manages career workflows: resume review and scoring, job evaluation with apply recommendations, rubric development, content extraction from conversations, and pipeline quality audits.
1
benchmark-methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
More results
skill-test
Validate SKILL.md files against a quality rubric, scoring schema and instruction structure out of 100 and reporting a pass/fail verdict with specific fixes.
13
agentic-eval
Implement iterative evaluation and refinement loops for AI agent outputs, using self-critique, evaluator-optimizer patterns, and rubric-based scoring to improve quality.
36.2k
aeon-autoresearch
Generates four improved variations of any installed skill, scores them against a weighted rubric, and applies the winning version while preserving the original.
1.2k · bundle
resume-coach
Reviews resumes for formatting, design, content, and job-specific fit; scores job postings for fit and interview odds; develops rubrics; extracts content from conversations; and audits the resume generation pipeline.
1
lead-scoring
Defines ideal customer profile filters, scores inbound and outbound leads, and builds a lightweight qualification rubric to sharpen pipeline focus for founder-led sales.
20
single-point-rubric-designer
Design a single-point rubric with one criterion and open columns for evidence. Use for student self-assessment, peer feedback, teacher formative feedback, or pre-task planning. Works with any learning target, with or without a band system.
0
grading-plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
22
priority-decision-system
Prioritize product work with explicit criteria, scoring, tradeoffs, and decision rationale.
0
lead-scoring
Score and prioritize leads based on firmographic fit and behavioral engagement signals, producing ranked tiers for sales team focus. Use when the user requests lead scoring or provides relevant inputs for this workflow.
159
lead-scorer
Score raw leads as HOT/WARM/COOL based on config-driven weights from agency.config.json
2 · bundle
acreadiness-policy
Create, apply, and manage AgentRC policies to customize readiness scoring, disable checks, override impact levels, set pass-rate thresholds, and enforce CI gating.
36.2k
context-ranking
Rank an existing set of context chunks by relevance, diversity, freshness, and utility. Use when retrieval has already produced candidates that must be scored or reranked; use context-retrieval when the source corpus still needs to be searched.
159
deal-desk
Scores deal margin and risk, routes discount approval to the right human approver, and redlines terms against commercial policy for per-deal review.
20.4k · bundle
icp-scoring
Turn a pile of accounts into a stack-ranked priority list with a reason on every row. A layered score (gates first, then an evidence-weighted base rank over the signals you actually have, then bounded boosts for product usage and buyer intent) that stays fair across channels and never scores a blank field as a zero. Built for B2B GTM teams, customizable to your signals and your ICP. Trigger on "score these accounts", "rank by fit", "composite ICP score", "stack-rank my list", "who should I work first", "prioritize these leads", or any multi-signal account qualification.
0 · bundle
quality-check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
auroc
Computes the AUROC metric using torchmetrics, handling binary, multiclass, and multilabel tasks with configurable thresholds and averaging.
3
ui-score
Score a UI file's design quality 0-100 against StyleSeed's design language with per-category breakdown, worst offenders, and prioritized fix list.
42.4k
rice
RICE feature prioritization with scoring and capacity planning. Usage: /rice prioritize <features.csv> [options]
1
coherent-rubric-logic-builder
Build a five-level rubric with coherent logic for a learning target within a developmental band. Use for Manning methodology programmes where Competent = success. For general curriculum rubrics, use criterion-referenced-rubric-generator instead.
0
cx-qa-appeal-process
Use to design or audit a QA dispute and appeal workflow with timeboxes, adjudication standards, and second-level consistency so appeals improve trust instead of rewriting scores without rules. Trigger for "QA appeal process", "agents disputing scores", "who adjudicates QA disputes", "overturn rate too high", second-level review standards, or calibration erosion from ad-hoc score changes.
1
ivx-ams-content-ops
Score and iteratively improve marketing content with an expert panel until 90+. Use for content quality gates, expert panel reviews, editorial scoring, or when another skill needs a content QA loop.
0 · bundle
lareine-charter
Judges agent outputs against LaRuche's quality standards, returning a structured scorecard with scores, verdict, and actionable corrections.
2
draft-score
Lightweight ContentShake AI self-check the /draft stage can call before saving. Returns just SEO + Quality scores (no full optimization) so the writer knows whether the draft is in winning territory before /quality-check runs. Fails soft when SEMRUSH_API_KEY is unset.
0
trak-attributing-model-behavior-at-scale-arxiv-2303-14186v2
TRAK: Attributing Model Behavior at Scale
6
ivx-sid-evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra (global). Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop. Works in any workspace; bootstraps EVALS.md from ~/.cursor/skills/sid-orchestra/templates if missing.
0 · bundle
209-sql-23f1987a
Provides SQL window function examples for ranking, aggregation, lag/lead, value extraction, frame specifications, and advanced analytics.
7 · bundle
grading-plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
3
ivx-cf-sid-evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra. Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop with evaluation gates.
0 · bundle
skill-logger
Logs and scores skill usage quality, tracking output effectiveness, user satisfaction signals, and improvement opportunities. Expert in skill analytics, quality metrics, feedback loops, and continuous improvement. Activate on "skill logging", "skill quality", "skill analytics", "skill scoring", "skill performance", "skill metrics", "track skill usage", "skill improvement". NOT for creating skills (use agent-creator), skill documentation (use skill-coach), or runtime debugging (use debugger skills).
10 · bundle
strike-zone-analyst
A funnel and account-scoring diagnostic engine for any sales org. Connect a CRM and a product-analytics tool (plus optional enrichment, community, and meeting tools). Three modes. (1) FUNNEL DIAGNOSIS finds leaky conversion gates by channel with per-stage leakage, dollarized leverage points, and cohort velocity. (2) SPRINT PLANNING enriches qualified accounts into a ranked backlog with verified buying committees. (3) SCORING AUDIT finds where your scoring model is missing real ICPs. Trigger on 'funnel diagnosis', 'diagnose the funnel', 'where are we leaking', 'why is [channel] underperforming', 'conversion by channel', 'sprint planning', 'score these accounts', 'find missed ICPs', 'audit the scoring model', or any channel-level cohort-conversion, account-prioritization, or scoring-gap question.
0 · bundle