Results for “rubric-scoring”

58 skills
eli-yu-first
Rubric Creator
Creates assessment rubrics with criteria, performance levels, and scoring guidelines for various assignments
6 · bundle
dvy1987
Eval Rubric Design
Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "evaluation rubric for", "how do I measure quality of". Sub-skill of eval-output orchestrator.
3 · bundle
muratcankoylan
Advanced Evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
forter
Forter Agentic Readiness Audit
Audits a website against the Forter Agentic Readiness Guide by running 25 weighted rubrics, scoring each guideline, and producing a prioritized fix report.
106 · bundle
joshuashepherd
Career
Manages career workflows: resume review and scoring, job evaluation with apply recommendations, rubric development, content extraction from conversations, and pipeline quality audits.
1
affaan-m
Benchmark Methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
More results
tinh2
Skill Test
Validate SKILL.md files against a quality rubric, scoring schema and instruction structure out of 100 and reporting a pass/fail verdict with specific fixes.
13
github
Agentic Eval
Implement iterative evaluation and refinement loops for AI agent outputs, using self-critique, evaluator-optimizer patterns, and rubric-based scoring to improve quality.
36.2k
bankrbot
Aeon Autoresearch
Generates four improved variations of any installed skill, scores them against a weighted rubric, and applies the winning version while preserving the original.
1.2k · bundle
vvieira010-pixel
Single Point Rubric Designer
Design a single-point rubric with one criterion and open columns for evidence. Use for student self-assessment, peer feedback, teacher formative feedback, or pre-task planning. Works with any learning target, with or without a band system.
0
joshuashepherd
Resume Coach
Reviews resumes for formatting, design, content, and job-specific fit; scores job postings for fit and interview odds; develops rubrics; extracts content from conversations; and audits the resume generation pipeline.
1
vvieira010-pixel
Coherent Rubric Logic Builder
Build a five-level rubric with coherent logic for a learning target within a developmental band. Use for Manning methodology programmes where Competent = success. For general curriculum rubrics, use criterion-referenced-rubric-generator instead.
0
qhjqhj00
Auroc
Computes the AUROC metric using torchmetrics, handling binary, multiclass, and multilabel tasks with configurable thresholds and averaging.
3
lionelndong
Quality Check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
antigravity
Evaluation
Build evaluation frameworks for agent systems, covering rubric design, test set creation, and automated evaluation pipelines.
42.4k
intense-visions
Spec Craft
Spec Craft
18 · bundle
dvy1987
Eval Output
Orchestrator for the eval-output skill suite — evaluate LLM and agent outputs for quality, accuracy, helpfulness, and safety using structured rubrics and LLM-as-judge techniques. Load when the user says "evaluate this output", "score this response", "run an eval", "LLM as judge", "evaluate agent output", "how good is this response", "rate this answer", "eval this", or provides an LLM output that should be assessed for quality. Single entry point for all output evaluation workflows.
3 · bundle
galyarderlabs
Lead Scoring
Defines ideal customer profile filters, scores inbound and outbound leads, and builds a lightweight qualification rubric to sharpen pipeline focus for founder-led sales.
20
alirezarezvani
Deal Desk
Scores deal margin and risk, routes discount approval to the right human approver, and redlines terms against commercial policy for per-deal review.
20.4k · bundle
seb1n
Context Ranking
Rank an existing set of context chunks by relevance, diversity, freshness, and utility. Use when retrieval has already produced candidates that must be scored or reranked; use context-retrieval when the source corpus still needs to be searched.
159
seb1n
Lead Scoring
Score and prioritize leads based on firmographic fit and behavioral engagement signals, producing ranked tiers for sales team focus. Use when the user requests lead scoring or provides relevant inputs for this workflow.
159
ekatasingh1107
Lead Scorer
Score raw leads as HOT/WARM/COOL based on config-driven weights from agency.config.json
2 · bundle
intelli-verse-x
Ivx Cf Sid Evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra. Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop with evaluation gates.
0 · bundle
qhjqhj00
Roc
Computes the Receiver Operating Characteristic (ROC) metric using torchmetrics, supporting binary, multiclass, and multilabel tasks.
3
github
Acreadiness Policy
Create, apply, and manage AgentRC policies to customize readiness scoring, disable checks, override impact levels, set pass-rate thresholds, and enforce CI gating.
36.2k
intelli-verse-x
Ivx Sid Evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra (global). Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop. Works in any workspace; bootstraps EVALS.md from ~/.cursor/skills/sid-orchestra/templates if missing.
0 · bundle
wondelai
UX Heuristics
Evaluate and improve interface usability using heuristic analysis based on Nielsen's 10 heuristics, Krug's laws, and severity ratings.
1.6k · bundle
danielpradilla
Priority Decision System
Prioritize product work with explicit criteria, scoring, tradeoffs, and decision rationale.
0
infinition
Lareine Charter
Judges agent outputs against LaRuche's quality standards, returning a structured scorecard with scores, verdict, and actionable corrections.
2
jiachen-t-wang
Trak Attributing Model Behavior At Scale Arxiv 2303 14186v2
TRAK: Attributing Model Behavior at Scale
6
snoodleboot-io
Batch Vs Realtime Scoring
The choice is not about scale or sophistication.
2
lionelsimai
Grading Plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
22
qhjqhj00
Logauc
Computes the LogAUC metric using the torchmetrics implementation for binary, multiclass, or multilabel classification tasks.
3
lionelndong
Draft Score
Lightweight ContentShake AI self-check the /draft stage can call before saving. Returns just SEO + Quality scores (no full optimization) so the writer knows whether the draft is in winning territory before /quality-check runs. Fails soft when SEMRUSH_API_KEY is unset.
0
dvy1987
Eval Judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
thedixitjain
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2