Results for “evaluator-checklist”

51 skills
More results
stribus
Speckit Checklist Agent
Generate a custom checklist for the current feature based on user requirements.
1
stribus
Speckit Implement Agent
Speckit Implement Agent
1
fradser
Retrospective
Analyzes evaluation patterns across completed plans and evolves the superpowers checklists accordingly.
580 · bundle
microsoft
M365 Agent Evaluator
Create, run, and analyze evaluation suites for Microsoft 365 Copilot declarative agents using the @microsoft/m365-copilot-eval CLI.
2.7k · bundle
racecraft-lab
Speckit Checklist
Generate a custom checklist for the current feature based on user requirements.
11
winbda
Punch List
Create punch lists for construction completion inspection. TRIGGERS - Use when user needs help with punch-list related tasks.
3
lionelsimai
Punch List
Create punch lists for construction completion inspection. TRIGGERS - Use when user needs help with punch-list related tasks.
22
winbda
Bucket List
Create bucket lists with categorization and planning. TRIGGERS - Use when user needs help with bucket-list related tasks.
3
agentskillexchange
Pre Landing Self Review
Runs a structured self-review checklist before committing substantial code changes, covering edge cases, error paths, test coverage, documentation, and code quality.
28
bankrbot
QA Checklist
Pre-ship audit checklist for Ethereum dApps covering wallet connection, button flows, contract verification, branding, RPC config, and mobile deep linking.
1.2k · bundle
owl-listener
Design QA Checklist
Create systematic QA checklists to verify that design implementations match specifications across visual accuracy, layout, interaction, content, accessibility, and cross-platform categories.
1.7k
30eggis
Testing Testing Tool Evaluator
Expert technology assessment specialist focused on evaluating, testing, and recommending tools, software, and platforms for business use and productivity optimization
2
lionelsimai
Grading Plan
Design grading plans. TRIGGERS - Use when user needs help with grading-plan related tasks.
22
construct-ai-primary
Output Validation Checklist
Use before delivering any work product to validate quality. This skill provides a standardized checklist that every agent must complete before marking work as done, ensuring consistent quality across all deliverables regardless of agent or company.
0
b4san
Requirement Checklist
Generate comprehensive checklists to validate requirements quality (Unit Tests for Specs). Critically examine the spec for clarity, completeness, and consistency before implementation starts.
2
snoodleboot-io
Post Implementation Checklist
Comprehensive checklist for documenting follow-up work and testing needs after implementation
2
dvy1987
Setup Evaluation
Validate process decomposition and architecture design quality before execution begins. Load when the setup-evaluator agent fires (automatic for agent-chain tasks), or when user says "evaluate this setup", "check the decomposition", "validate the architecture", "is this plan sound", "review the agent design". Catches structural errors, missing knowledge, unrealistic step ordering, and topology mismatches. Does NOT modify — only evaluates.
3 · bundle
jeffallan
Playwright Expert
Write robust, maintainable end-to-end tests with Playwright using Page Object Model, proper selectors, auto-waiting, and debugging workflows.
10.4k · bundle
gonglingrui
Script Evaluator
从思想性、艺术性、观赏性三维度评估影视剧本并打分。适用于剧本开发质量评估、修改方向确定、项目立项前审查
349 · bundle
paramchordiya
Code Review Standards
Enforces a principal-engineer self-review checklist on every code block before output, covering correctness, performance, security, naming, and testability.
0
jrennie99-glitch
Agent Reviewer
Agent skill for reviewer - invoke with $agent-reviewer
0
drnabeelkhan
Tool Evaluator
Assesses new tools, technologies, and integration options for adoption, comparing vendors and recommending implementations.
2
alunadev
Requesting Code Review
Self-contained code review checklist — formatting, type safety, duplication, readable conditions, logic-out-of-UI, test coverage, comment quality. Adapts to whatever tooling the project actually has (skips Prettier/TypeScript/test-runner checks if unconfigured). Use before merging, after completing a feature, or any time you want a second pass on code quality. In Claude Code, prefer the official `code-review` skill or `pr-review-toolkit`/`feature-dev` code-reviewer agents instead — both are real, installed, and go deeper than this checklist. This copy exists for Codex/Cursor portability, where those agents aren't available.
3
danielpradilla
Launch Control Checklist
Build a launch checklist with owners, dates, dependencies, decision gates, and follow-through.
0
bobmatnyc
Code Review Standards
Severity-tagged code review checklist (CRITICAL/HIGH/MEDIUM/LOW) used by code-critic agent
71 · bundle
auto-skiller
Assess Quality
Evaluates execution outcomes against defined success criteria, scoring each criterion and producing a structured verdict with actionable feedback.
1 · bundle
jorcan
Agents
Evaluates execution transcripts and output files against a list of expectations, assigning pass/fail verdicts with cited evidence and critiquing the assertions themselves.
0 · bundle
samyakjhaveri
Eval Grader
Grades and classifies evaluation batch results, applying exclusions, diagnosing failure modes, computing pass rates, and generating summary tables for papers.
0
whd4
Executing Plans
Use when you have a written implementation plan to execute in a separate session with review checkpoints
0
muratcankoylan
Evaluation
Build evaluation frameworks for agent systems with deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, and outcome measurement.
16.9k · bundle
gonglingrui
Drama Evaluator
依据竖屏短剧评估标准,从核心爽点、故事类型等维度评估打分。适用于评估故事改编为竖屏短剧的潜力、分析市场竞争力
349 · bundle
builderio
Plan Arbiter
Compare, cross-review, and merge competing plans from multiple agents into one executable direction with a clear handoff.
3.4k · bundle
nvidia
Nemo Evaluator Plugin
Run evaluation tasks against a NeMo Platform server using the Evaluator plugin CLI and Python SDK.
2.2k · bundle
github
Phoenix Evals
Build and run evaluators for AI/LLM applications using Phoenix, covering error analysis, custom evaluators, experiments, and production monitoring.
36.2k · bundle
sdiamante13
Review Plan
Reviews development plans for gaps, hidden assumptions, critical misses, and XP violations, then presents prioritized findings for confirmation.
7 · bundle