Results for “adversarial-evaluation”
20 skillsMore results
A2
VS-Enhanced Theoretical Framework Architect with Critique & Visualization Full VS 5-Phase process: Modal theory avoidance, Long-tail exploration, differentiated framework presentation Absorbed A3 (Devil's Advocate) critique and A6 (Conceptual Framework Visualizer) capabilities Use when: building theoretical foundations, designing conceptual models, deriving hypotheses, critiquing frameworks, visualizing models Triggers: theoretical framework, 이론적 프레임워크, conceptual model, 개념적 모형, hypothesis derivation, critique, devil's advocate, 반론, visualization, diagram
1k
C2c Eval
Benchmarks language model agents on the C2C multi-agent negotiation task, reporting win rate across starting positions.
3
Santa Method
Multi-agent adversarial verification with convergence loop. Two independent review agents must both pass before output ships.
0
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
0
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2
Santa Method
Multi-agent adversarial verification with convergence loop. Two independent review agents must both pass before output ships.
1
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
Agent Self Evaluation
Rates an agent's own output on five axes — accuracy, completeness, clarity, actionability, conciseness — producing a structured scorecard with evidence and improvement suggestions.
226k · bundle
Concurrency Control
`analysis-agent`/`task-agent`/`review-agent`: primary-Skill-selected for races, locks, optimistic conflicts, or worker overlap; never task owner; skip without concurrency impact.
4 · bundle
Evaluation
Build evaluation frameworks for agent systems, covering rubric design, test set creation, and automated evaluation pipelines.
42.4k
Advanced Evaluation
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
55 · bundle
Agent Evaluation
Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis. Use when defining agent quality, comparing prompts or models, validating a release, measuring tool-use reliability, investigating regressions, or deciding whether an agent is ready for production.
159 · bundle
Teach Back Evaluator
The learner teaches the concept to the AI, which plays a curious novice peer and identifies gaps through authentic questions. Use when the learner wants to test their understanding — teaching forces a different kind of organisation than studying.
0
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
11
Santa Method
Runs a multi-agent adversarial verification loop where two independent reviewers must both pass before output ships, with a fix cycle for convergence.
1
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
3
Santa Method
Multi-agent adversarial verification with convergence loop. Two independent review agents must both pass before output ships.
0
Advanced Evaluation
Provides production-grade techniques for evaluating LLM outputs using LLMs as judges, covering direct scoring, pairwise comparison, bias mitigation, rubric generation, and confidence calibration.
16.9k · bundle
Agentic Engineering
Operate as an agentic engineer using eval-first execution, decomposition, and cost-aware model routing.
1