Results for “grade-evidence-grading”
23 skillsMore results
grade-tests
Grades individual test methods and produces a compact markdown table with a letter grade, score band, and one-line note for each test.
4k
grading-plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
22
grading-plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
3
ivx-cf-evaluation
Design and implement evaluation harnesses for models, agents, and code. Use when creating benchmarks, designing eval metrics, or comparing system outputs.
0 · bundle
lead-qualifier
Multi-dimensional lead qualification scoring. Evaluates leads against BANT criteria, firmographic fit, behavioral signals, and intent indicators. Outputs qualified/disqualified verdict with detailed reasoning.
2 · bundle
promote
Graduate a proven pattern from auto-memory (MEMORY.md) to CLAUDE.md or .claude/rules/ for permanent enforcement.
0
verify
Combined verification — recite (description quality via cold-read prediction) + validate (schema compliance) + review (health checks). Use as a quality gate after creating notes or as periodic maintenance. Triggers on "/verify", "/verify [note]", "verify note quality", "check note health".
3 · bundle
cross-validation-strategies
Cross-validation only estimates generalization if the split mimics the gap between
2
ord-traction
Check post-launch adoption for a published ORD package and recommend iterate, graduate, or archive
1 · bundle
tw-ideate
Mines a codebase for evidence-backed improvement opportunities, forcing two escalation gates to surface breakthrough ideas, and outputs a ranked portfolio with a plan seed without implementing.
7
curriculum-learning-crossref-icml-2009-curriculum
Curriculum Learning
6
paper-review-sim
Simulates a NeurIPS/SC/ICSE-style peer review with five reviewer personas (HPC, ML, Stats, Reproducibility, Devil's Advocate) that verify every claim against actual result data before submission.
0
context-ranking
Rank an existing set of context chunks by relevance, diversity, freshness, and utility. Use when retrieval has already produced candidates that must be scored or reranked; use context-retrieval when the source corpus still needs to be searched.
159
score-analyzer
Score Analyzer
349 · bundle
spike-guidance-adversarial
Readability Review
218
dataq-evidence-standards
Use this skill to evaluate which DataQ challenges have winning evidence and which don't. Covers documented vs anecdotal evidence and the 8 high-success patterns.
1
trace
Evidence-driven tracing lane that orchestrates competing tracer hypotheses in Claude built-in team mode
1
gateguard
Fact-forcing gate that blocks Edit/Write/Bash (including MultiEdit) and demands concrete investigation (importers, data schemas, user instruction) before allowing the action. Measurably improves output quality by +2.25 points vs ungated agents.
0
svit-scaling-up-visual-instruction-tuning-arxiv-2307-04087v2
SVIT: Scaling up Visual Instruction Tuning
6
ord-align
Validate whether an OSS tool is worth building through staged namespace, existing-solution, and feasibility review
1 · bundle
review
Routes quality reviews to specialized critic agents based on file type or flags, covering peer review, code review, and manuscript polish.
3 · bundle
verify-before-done
Require fresh evidence before claiming an implementation, fix, build, or test is complete.
4