Results for “adversarial-evaluation”

51 skills
More results
dvy1987
Adversarial Hat
Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".
3 · bundle
lionelndong
Research Adversarial
Skeptical pushback on the research dossier before it feeds the outline. Asks whether claims are cited, whether surprising findings are actually surprising, whether we missed the strongest competitor angle. One revision pass on FAIL (BLOG_AGENT_RESEARCH_REVISION_BUDGET, default 1).
0
github
From The Other Side Anitta
Provides a rigorous thinking partner profile that challenges assumptions, calibrates claims to evidence, and improves decision quality under uncertainty.
36.2k
brycewang-stanford
A2
VS-Enhanced Theoretical Framework Architect with Critique & Visualization Full VS 5-Phase process: Modal theory avoidance, Long-tail exploration, differentiated framework presentation Absorbed A3 (Devil's Advocate) critique and A6 (Conceptual Framework Visualizer) capabilities Use when: building theoretical foundations, designing conceptual models, deriving hypotheses, critiquing frameworks, visualizing models Triggers: theoretical framework, 이론적 프레임워크, conceptual model, 개념적 모형, hypothesis derivation, critique, devil's advocate, 반론, visualization, diagram
1k
sirnosh
Bmad Ml Kayo
Adversarial reviewer that stress-tests claims and conclusions. Use when the user asks to talk to KAY/O, requests an adversarial review, or needs claims validated before publication.
0 · bundle
mukul975
Analyzing Campaign Attribution Evidence
Systematically evaluates evidence to determine which threat actor is responsible for a cyber operation using the Diamond Model and Analysis of Competing Hypotheses.
24.6k · bundle
lovits
Ultraqa
Adversarial dynamic e2e QA workflow - generate hostile scenarios, test, verify, fix, report, and clean up
0
qhjqhj00
C2c Eval
Benchmarks language model agents on the C2C multi-agent negotiation task, reporting win rate across starting positions.
3
lionelndong
Adversarial Quality Gate
Decide whether an article deserves publication through hard checks and skeptical comparison.
0
livelybug
Santa Method
Multi-agent adversarial verification with convergence loop. Two independent review agents must both pass before output ships.
0
curiositech
Alphago Deep Rl
Strategic patterns for solving intractable problems through cascading approximation, self-improvement, and heterogeneous evaluation from DeepMind's AlphaGo system
10 · bundle
jarbitechture
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
0
leandrobenjaminl
Judgment Day
Runs an adversarial code review with two blind judges analyzing the same code from opposing perspectives to find flaws before production.
0
seb1n
Agent Red Teaming
Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings. Use when defining red-team rules of engagement, assessing prompt injection or excessive agency, testing tool and identity boundaries, evaluating memory or cross-agent attacks, scoring a campaign, or verifying remediation in an approved environment.
159 · bundle
thedixitjain
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session. Use when the user runs /hub:eval or asks to score, compare, or pick a winner among completed AgentHub agents.
2
anantha-236
Santa Method
Multi-agent adversarial verification with convergence loop. Two independent review agents must both pass before output ships.
1
alirezarezvani
Self Eval
Honestly evaluate AI work quality using a two-axis scoring system with mandatory devil's advocate reasoning and cross-session anti-inflation detection.
20.4k
alirezarezvani
Eval
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
20.4k
shenmuxing
Proof Checker V2
Independent DeepSeek-backed adversarial proof-audit step for existing theorem, lemma, proposition, or proof artifacts in Markdown, LaTeX, or proof logs. Use when asked to check, audit, verify, red-team, or adversarially review a proof; when a completed proof task needs a correctness pass; or when a broader proof workflow dispatches an independent reviewer to find gaps, hidden assumptions, counterexamples, or unjustified steps.
2 · bundle
affaan-m
Benchmark Methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
lionelndong
Visuals Adversarial
Skeptical pushback on visual placement — both density and quality. Reads the annotated outline plus the visuals manifest and asks (a) whether the article hits the density target from editorial-principles-visuals.md, (b) whether each [VISUAL:...] earns its place, (c) whether sections without one would benefit. One revision pass on FAIL (BLOG_AGENT_VISUALS_REVISION_BUDGET, default 1).
0
samyakjhaveri
Critique Swarm
Launches four parallel adversarial review agents covering security, scope, fidelity, and test gaps, then synthesizes findings into a ranked verdict for plans or feature branches.
0
delorenj
Bmad Review Edge Case Hunter
Walk every branching path and boundary condition in content, report only unhandled edge cases. Orthogonal to adversarial review - method-driven not attitude-driven. Use when you need exhaustive edge-case analysis of code, specs, or diffs.
1 · bundle
pablolion
Bmad Review Edge Case Hunter
Walk every branching path and boundary condition in content, report only unhandled edge cases. Orthogonal to adversarial review - method-driven not attitude-driven. Use when you need exhaustive edge-case analysis of code, specs, or diffs.
12
mukul975
Designing Adversary Engagement With Mitre Engage
Plan, run, and measure adversary engagement operations using the MITRE Engage framework, covering the Engage Matrix, 10-Step Operational Process, and mapping Activities to ATT&CK techniques.
24.6k · bundle
salacoste
Bmad Review Edge Case Hunter
Walk every branching path and boundary condition in content, report only unhandled edge cases. Orthogonal to adversarial review - method-driven not attitude-driven. Use when you need exhaustive edge-case analysis of code, specs, or diffs.
1
affaan-m
Agent Self Evaluation
Rates an agent's own output on five axes — accuracy, completeness, clarity, actionability, conciseness — producing a structured scorecard with evidence and improvement suggestions.
226k · bundle
phuryn
Identify Assumptions Existing
Stress-test a feature idea for an existing product by surfacing risky assumptions across Value, Usability, Viability, and Feasibility using multi-perspective devil's advocate thinking.
22.6k
alirezarezvani
Roast
Pressure-test any business idea with a five-angle adversarial panel (Critic, Champion, Analyst, Investigator, Customer) that returns a single GO/RESHAPE/KILL verdict with the cheapest test to de-risk it.
20.4k · bundle
machenjie
Concurrency Control
`analysis-agent`/`task-agent`/`review-agent`: primary-Skill-selected for races, locks, optimistic conflicts, or worker overlap; never task owner; skip without concurrency impact.
4 · bundle
delorenj
Bmad Code Review
Adversarial code review using parallel review layers and structured triage. Use when the user says "run code review" or "review this code"
1 · bundle
antigravity
Evaluation
Build evaluation frameworks for agent systems, covering rubric design, test set creation, and automated evaluation pipelines.
42.4k
kursku
Advanced Evaluation
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
55 · bundle
lionelndong
Keyword Redteam
Layer 4 of the keyword research pipeline. Spawns a "skeptical SEO" adversarial sub-agent to argue against every survivor of Layers 1-3. Catches mechanical-classifier blind spots — wrong SERP intent, hidden link-graph gauntlets, AIO trajectory shifts, vanity-rank metrics. Same pattern as quality-check's adversarial draft read, applied to keyword selection.
0