Results for “adversarial-evaluation”

15 skills
More results
github
From The Other Side Anitta
Provides a rigorous thinking partner profile that challenges assumptions, calibrates claims to evidence, and improves decision quality under uncertainty.
36.2k
sirnosh
Bmad Ml Kayo
Adversarial reviewer that stress-tests claims and conclusions. Use when the user asks to talk to KAY/O, requests an adversarial review, or needs claims validated before publication.
0 · bundle
lionelndong
Adversarial Quality Gate
Decide whether an article deserves publication through hard checks and skeptical comparison.
0
curiositech
Alphago Deep Rl
Strategic patterns for solving intractable problems through cascading approximation, self-improvement, and heterogeneous evaluation from DeepMind's AlphaGo system
10 · bundle
leandrobenjaminl
Judgment Day
Runs an adversarial code review with two blind judges analyzing the same code from opposing perspectives to find flaws before production.
0
lionelndong
Visuals Adversarial
Skeptical pushback on visual placement — both density and quality. Reads the annotated outline plus the visuals manifest and asks (a) whether the article hits the density target from editorial-principles-visuals.md, (b) whether each [VISUAL:...] earns its place, (c) whether sections without one would benefit. One revision pass on FAIL (BLOG_AGENT_VISUALS_REVISION_BUDGET, default 1).
0
samyakjhaveri
Critique Swarm
Launches four parallel adversarial review agents covering security, scope, fidelity, and test gaps, then synthesizes findings into a ranked verdict for plans or feature branches.
0
alirezarezvani
Roast
Pressure-test any business idea with a five-angle adversarial panel (Critic, Champion, Analyst, Investigator, Customer) that returns a single GO/RESHAPE/KILL verdict with the cheapest test to de-risk it.
20.4k · bundle
delorenj
Bmad Code Review
Adversarial code review using parallel review layers and structured triage. Use when the user says "run code review" or "review this code"
1 · bundle
lionelndong
Outline Adversarial
Skeptical pushback on the outline before drafting starts. Asks whether the outline is MECE, BLUF, problem-agitate-solution; whether it covers the topic better than the SERP top-5; whether sections duplicate; whether each section earns its visual. One revision pass on FAIL (BLOG_AGENT_OUTLINE_REVISION_BUDGET, default 1).
0
pwdev-solucoes
Flow Verify
Adversarially verify an implementation against its objective, acceptance criteria, definition of done, and prohibitions using fresh evidence. Use for completion gates, release readiness, or independent verification after implementation and review.
2 · bundle
testdouble
Spike Consumer Adversarial
Incident Post-Mortem Builder
218
jiachen-t-wang
Llava Critic Learning To Evaluate Multimodal Models Arxiv 24
LLaVA-Critic: Learning to Evaluate Multimodal Models
6
dvy1987
Idea Evaluation
Score an unbuilt business idea on desirability, viability, feasibility, distribution wedge, why-now, founder-market-fit, market size, alternatives, defensibility, capital intensity, and regulatory/ethical risk — and return a GO / ITERATE / KILL verdict with kill criteria and a next kill test. Load when the user asks to evaluate a business idea, score a startup idea, screen an idea, decide whether to pursue this venture, do an idea review, or says "is this a good business idea", "should I build this", "evaluate this startup", "screen this idea", "go/no-go on this idea", "kill or pursue". Sub-skill of `venture-exploration`. Calls `fermi` for sizing, `assumption-mapping` for hidden beliefs, optional `pre-mortem` / `adversarial-hat` for high-stakes ideas. Does NOT evaluate built products — for that use `reality-check`.
3 · bundle