Results for “big-bench-hard”
49 skillsMore results
Whiteboard
Plan a chunk of work too big for one agent session by putting it on a shared whiteboard — a map of investigation tickets on GitHub Issues — working them until nothing is left to decide, then snapshotting the board into a handoff artifact. Tickets needing nobody are worked back-to-back; the session stops when the human is the blocker. Use only when the user explicitly invokes whiteboard or asks to draw, work, run, or snapshot a whiteboard/map — not for ordinary planning requests.
0 · bundle
Bench Read
Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.
0
Dialectic
Multi-phase dialectical stress-test for HIGH-STAKES decisions only (architecture choices, irreversible product calls, strategic bets). Heavy-cost skill — do NOT use for routine questions, brainstorming, or simple tradeoffs. User must explicitly invoke or describe a decision they call "high-stakes", "irreversible", or "needs stress-testing". Triggers: "stress test this decision", "dialectic on", "challenge this thesis", "should I really".
6 · bundle
Hard Negative Mixing For Contrastive Learning Arxiv 2010 010
Hard Negative Mixing for Contrastive Learning
6
Cost
Evaluates a containerized framework for deploying distributed big data workloads, measuring execution time and cloud cost scaling from four to eight nodes.
3
Muller Brockmann Grid Systems
Build editorial/magazine/report webpages on a GENUINE Müller-Brockmann modular grid (International Typographic Style), not a decorative one. Encodes the discipline — columns + modules + baseline, grotesque type, flush-left ragged-right, restrained black/white/red — AND the front-end engineering that makes the grid load-bearing: one CSS-variable source of truth, a grid-toggle overlay that lives in the SAME content box as the content, subgrid 'bands' so every element snaps to a column line, an 8px baseline lock, and runtime optical alignment that puts display type's ink (not its box) on the line. Ships a scaffold generator and a Puppeteer harness that proves 0px adherence. Use when: building any editorial, magazine, report, or longform page that must read as rigorously grid-aligned, Swiss, International Typographic Style, or 'Müller-Brockmann'. Triggers: magazine spread, grid system, Swiss design, editorial layout, show the grid, grid overlay toggle, baseline grid, modular grid, 12-column layout, 瑞士网格, 版面网格.
8 · bundle
Gpt Taste
Generates award-level UI with GSAP motion, bento grids, and Python-driven randomization for layout variety.
Web Proto Brutalist
Generates brutalist web prototypes with Swiss industrial-print aesthetics, using grotesque typography, oversized numbers, ASCII decorations, and hazard-red accents.
· bundle
Bald Eagle Strategy
BALD EAGLE v3.0 — XYZ Alpha Hunter (Hardened). Focused on 6 high-liquidity XYZ assets: CL, BRENTOIL, GOLD, SILVER, SP500, XYZ100. Conviction-scaled leverage (5-10x based on score). Wider DSL for macro assets. Maker-only execution. Scanner calls create_position internally. v3.0: focused assets, conviction-scaled leverage, XYZ-tuned DSL, no thesis exit.
1 · bundle
Hundred Million Offers
Create irresistible offers using the Value Equation, bonus stacking, risk-reversing guarantees, and ethical scarcity.
1.6k · bundle
Adversarial Hat
Put on the adversarial hat and systematically attack any document, plan, strategy, or idea to expose its weakest points before commitment. Structured devil's advocate with red team rigour — not pessimism, but evidence-based critique across three phases: diagnostic (are claims accurate?), creative (is the problem artificially constrained?), challenge (are solutions robust?). Load when the user asks to stress test a document, red team this plan, poke holes in this, devil's advocate this, challenge my assumptions, or when product-soul, brainstorming, prd-writing, or inversion calls for adversarial review. Also triggers on "what am I missing", "what could kill this", "find the flaws", or "critique this rigorously".
3 · bundle
Deck Xhs White
Generates a white magazine-style deck with rainbow bar, gradient text, macaron cards, and black pill for dual-use on Xiaohongshu and horizontal PPT.
· bundle
Business Modeling
Pick the right business-model canvas (Lean Canvas, Business Model Canvas, or Value Proposition Canvas) for the stage and fill it with specifics — one segment, one primary canvas, top-3 assumptions, no fluff in the moat or channel boxes. Load when the user asks to fill a business model canvas, lean canvas, value proposition canvas, model this business, map the business model, says "fill the BMC", "make a Lean Canvas", "Value Proposition Canvas for this", "model this idea", "what's the business model", "design the business model". Sub-skill of `venture-exploration`. Hard-bans "everyone" segments, generic channels ("SEO/social/content/ads"), and "unfair advantage = AI/data/network effects" with no concrete asset. Does NOT score viability — for that use `idea-evaluation`.
3 · bundle
Awesome Rebuttal
Install a local rebuttal workspace for academic papers with structured intake, reviewer analysis, and strategy planning to produce venue-compliant author responses.
298 · bundle
Branch Hygiene
Composite skill — one-pass cleanup of stale local branches, dead worktrees, merged branches, and abandoned remote PR branches. Chains `git fetch --prune` → `clean_gone` (kill [gone] branches) → worktree prune + offer-to-remove dead worktrees → list-and-delete branches merged to main and release → delete remote PR branches whose PRs merged >7 days ago. Use instead of running `clean_gone` alone — that only catches half the rot. Daily-friction composite; fires on "clean up branches", "branch hygiene", "stale worktrees", and on session start when local branch count > 30.
1 · bundle
Matryoshka Representation Learning Arxiv 2205 13147v4
Matryoshka Representation Learning
6
Grill Me
Runs a relentless interview session to sharpen a plan or design.
3
Openrlhf Training
Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.
10.4k · bundle
Bison Strategy
BISON v2.0 — Conviction Holder (Hardened). Top 10 assets by volume. All signals are score contributors — no hard gates. Scanner enters via create_position internally (Wolverine pattern). RatchetStop exits. Thesis exit REMOVED. v2.0: every hard gate converted to score contributor, ensureExecutionAsTaker=false, conviction-scaled margin 25-37%.
1 · bundle
Performance Engineer
Benchmark, load test, capacity plan, and cache with k6, JMeter, Locust, and pgbench. Use when the user says "slow", "performance", "load test", "stress test", "how many users can it handle", "capacity", "cache", "benchmark", "k6".
2
Training Llms Megatron
Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek.
1 · bundle
Hunch
Discover, bet on, track, and settle Hunch prediction markets in natural language with on-chain settlement on Base via x402.
1.2k · bundle
Context Compression
Manage and compress conversation context in long sessions. Detect when context is growing large, summarize completed work phases, archive old findings while preserving key decisions. Prevents context degradation.
3
Poster Hero
Generates visually striking vertical posters for social media sharing, with full-screen gradient backgrounds, bold typography, and decorative SVG elements.
· bundle
Training Compute Optimal Large Language Models Arxiv 2203 15
Training Compute-Optimal Large Language Models
6
Decision Advisor
Structure a hard business decision — frame the choice, score options against weighted criteria, run a pre-mortem and stress-test, and produce a recommendation + decision record for any active brand
0
Igce Builder Lh Tm
Labor-hour and time-and-materials cost buildup with fully burdened hourly rates (BLS OEWS + GSA CALC+ + per diem MCPs). USE WHEN the user asks to "build a T&M IGCE", "labor hour estimate", "fully burdened hourly rate", "LH IGCE", "time and materials pricing", or "burden multiplier stack". Exports JSON + XLSX to Studio. DO NOT USE FOR FFP wrap rates (`igce-builder-ffp`), cost-reimbursement fee caps (`igce-builder-cr`), or OT bids (`ot-prototype-strategist`).
0
Hot Seat
Put the user in the hot seat — one question at a time until their plan, decision, or idea actually holds up. Use whenever the user says "hot seat me", "grill me", "stress-test this", "poke holes in this", "am I missing anything", or drops a plan and wants it challenged before anyone acts on it — even if they never say the words. Also used as a sub-procedure by the whiteboard and connotation-cop skills.
0
Paw Cra Design Batch
Batch visual production workflow for campaigns. Accepts a content calendar or campaign brief, produces multiple platform-ready assets with organized bundle output and machine-readable manifest. Trigger for 'batch design', 'content calendar to assets', 'campaign batch', 'produce all campaign visuals', or 'bulk asset generation'.
85 · bundle
Perf Long Tasks
Long Tasks
18 · bundle
Bigquery Basics
Manage datasets, tables, and jobs in BigQuery. Run SQL queries, manage BigQuery resources, and perform basic data ingestion and analysis.
14.4k · bundle
Training Llms Megatron
Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies for maximum GPU efficiency.
10.4k · bundle
Muapi Giant Product Showcase
Creates a dramatic 'giant product' visual by compositing a product image into a scene where it appears building-sized next to a person, with an optional animation step.
3.7k
Hard Call
Hard Call
0
Training Llms Megatron
Trains large language models (2B-462B parameters) using NVIDIA Megatron-Core with advanced parallelism strategies. Use when training models >1B parameters, need maximum GPU efficiency (47% MFU on H100), or require tensor/pipeline/sequence/context/expert parallelism. Production-ready framework used for Nemotron, LLaMA, DeepSeek.
0 · bundle