Results for “grading”

61 skills
More results
jorcan
Agents
Evaluates execution transcripts and output files against a list of expectations, assigning pass/fail verdicts with cited evidence and critiquing the assertions themselves.
0 · bundle
samyakjhaveri
Eval Grader
Grades and classifies evaluation batch results, applying exclusions, diagnosing failure modes, computing pass rates, and generating summary tables for papers.
0
dotnet
Grade Tests
Grades individual test methods and produces a compact markdown table with a letter grade, score band, and one-line note for each test.
4k
adobe
Adobe Batch Edit Photos
Apply consistent photo adjustments across a set of images so they look like they were edited together, using Adobe tools for color, tone, cropping, and optional selective enhancements.
142
muratcankoylan
Advanced Evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
lionelndong
Quality Check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
ekatasingh1107
Lead Qualifier
Multi-dimensional lead qualification scoring. Evaluates leads against BANT criteria, firmographic fit, behavioral signals, and intent indicators. Outputs qualified/disqualified verdict with detailed reasoning.
2 · bundle
dvy1987
Eval Rubric Design
Design structured evaluation rubrics for scoring LLM and agent outputs — defining quality dimensions, scoring scales, hard gates, score descriptions, and edge cases. Load when the user asks to create an eval rubric, define evaluation criteria, design scoring dimensions, write an eval spec, or says "what should I evaluate", "design a rubric", "create eval criteria", "define quality dimensions", "evaluation rubric for", "how do I measure quality of". Sub-skill of eval-output orchestrator.
3 · bundle
lucassantana-dev
Eval
Evaluate LLM outputs systematically — benchmarks, automated metrics, human preference, and regression tracking
1 · bundle
nvidia
RAG Eval
Evaluates RAG pipelines using a filesystem-based benchmark with corpus/ and train.json, running evaluate_rag.py to tune retrieval and generation flags and interpret RAGAS metrics.
2.2k · bundle
galyarderlabs
Lead Scoring
Defines ideal customer profile filters, scores inbound and outbound leads, and builds a lightweight qualification rubric to sharpen pipeline focus for founder-led sales.
20
georgeqle
Ord Traction
Check post-launch adoption for a published ORD package and recommend iterate, graduate, or archive
1 · bundle
lionelndong
Adversarial Quality Gate
Decide whether an article deserves publication through hard checks and skeptical comparison.
0
seaworld008
Guardian
Gatekeeping Git/PR by classifying change essence and recommending granularity, naming, and strategy. Use when PR preparation or commit strategy is needed.
65 · bundle
alirezarezvani
Self Eval
Honestly evaluate AI work quality using a two-axis scoring system with mandatory devil's advocate reasoning and cross-session anti-inflation detection.
20.4k
jiachen-t-wang
Trak Attributing Model Behavior At Scale Arxiv 2303 14186v2
TRAK: Attributing Model Behavior at Scale
6
curiositech
Design Critic
Aesthetic assessment and design scoring across 6 dimensions. Use for UI critique, design review, visual quality assessment, remix suggestions. Activate on "design critique", "aesthetic review", "UI assessment", "visual quality", "design score", "remix this design". NOT for implementation (use frontend-developer), accessibility-only audits (use color-contrast-auditor), or brand identity creation.
10 · bundle
rulebase-co
Cx Reviewer Workload
Use to staff and schedule QA reviewers for sustainable throughput without grading quality collapsing — fatigue, drift, and false precision from treating reviewers like production agents. Trigger for "how many QA reviewers do we need", reviewer capacity planning, grader fatigue, throughput per reviewer, calibration after long shifts, or scores drifting over the course of a day.
1
seb1n
Context Ranking
Rank an existing set of context chunks by relevance, diversity, freshness, and utility. Use when retrieval has already produced candidates that must be scored or reranked; use context-retrieval when the source corpus still needs to be searched.
159
fukukei23
Sentaku
選択肢(A/B/C)の深掘り比較→淘汰→推奨で判断負担を下げ判断の質を上げるスキル。5段階(L1固定3点/L1.5案拡張Diverge・自動/L2評価軸マトリクス/L3複数LLM弁証論/L4過去判断照合)。 「比較して」「深掘りして」「メリデメ教えて」「お勧めは?」「徹底的に」「過去の判断と照合」「前にどう決めたっけ」「/sentaku」等で発火。teian(浅)の深掘り要求を受け取り、brainstorming(深:設計全体)と棲み分け。
0
heath-gtm
Icp Scoring
Turn a pile of accounts into a stack-ranked priority list with a reason on every row. A layered score (gates first, then an evidence-weighted base rank over the signals you actually have, then bounded boosts for product usage and buyer intent) that stays fair across channels and never scores a blank field as a zero. Built for B2B GTM teams, customizable to your signals and your ICP. Trigger on "score these accounts", "rank by fit", "composite ICP score", "stack-rank my list", "who should I work first", "prioritize these leads", or any multi-signal account qualification.
0 · bundle
intelli-verse-x
Ivx Sid Evals
PASS/FAIL eval rubrics and alignment loops for Sid Orchestra (global). Use when the user says sid evals, @sid-evals, grade this, eval gate, alignment score, or wants to stop AI slop. Works in any workspace; bootstraps EVALS.md from ~/.cursor/skills/sid-orchestra/templates if missing.
0 · bundle
antigravity
UI Score
Score a UI file's design quality 0-100 against StyleSeed's design language with per-category breakdown, worst offenders, and prioritized fix list.
42.4k
snoodleboot-io
Code Review Practices
The single highest-leverage convention in code review is prefixing every comment
2
mukul975-2
Vendor Risk Scoring
Vendor privacy risk tiering methodology for processor management. Covers scoring factors including data volume, sensitivity, transfer locations, certifications, breach history, and control maturity with weighted risk calculation and tier assignment.
228 · bundle
onourimpram
Teaching Feedback AI Boundaries
Use when designing a course, assessment, or student-feedback workflow and deciding where AI assistance is acceptable, when grading or feedback must stay within academic-integrity and FERPA/KVKK boundaries, or when AI-use rules for students need stating.
2
winbda
Badge System
Design badge and achievement systems for learning. TRIGGERS - Use when user needs help with badge-system related tasks.
3
alirezarezvani
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
owl-listener
Critique Typography
Audits typographic decisions on a screen for scale usage, readability, consistency, and token compliance, providing specific fixes.
1.7k
lionelndong
Portfolio And Measurement
Improve existing content and close the learning loop without cannibalizing new-content work.
0