Results for “proofpoint”
9 skillsexp-eval
实验判决门:Review LLM 独立评判实验结果 → 4 种判决路径 → 自动更新 claims confidence、ideas status、graph edges
77
genie-proof-prompts
Rewrite any prompt, instruction, task description, or spec into a "genie-proof" version — instructions so explicit, literal, and loophole-free that even a maliciously literal genie (or an LLM, contractor, or junior dev) could not misinterpret them. Use this skill whenever the user asks to genie-proof, tighten, harden, de-ambiguate, or "make bulletproof" a prompt or instruction; whenever they complain that an AI/model/person "didn't do what I meant," "took me too literally," or "found a loophole"; or whenever they hand over a vague prompt and ask to make it precise, explicit, unambiguous, or idiot-proof. Also trigger on phrases like "wish to a genie," "monkey's paw," "lawyer-proof this prompt," or "leave nothing to interpretation."
0
result-to-claim
Use when experiments complete to judge what claims the results support, what they don't, and what evidence is still missing. Codex MCP evaluates results against intended claims and routes to next action (pivot, supplement, or confirm). Use after experiments finish — before writing the paper or running ablations.
1k
key-moments
Rank a topic's user-flow branches by proof priority (value × risk × frequency) right after user-flow-map, ordering the branches, gating variation breadth, and promoting or pruning flows so state-model and ux-variations grow the tree in proof order — writes only existing flow-tree ordering fields, no schema change.
1 · bundle
playwright-expert
Write robust, maintainable end-to-end tests with Playwright using Page Object Model, proper selectors, auto-waiting, and debugging workflows.
10.4k · bundle
ai-regression-testing
Prevents AI-introduced regressions with sandbox-mode API testing, automated bug-check workflows, and patterns that catch blind spots where the same model writes and reviews code.
226k
unassisted-evidence-checkpoint
After scaffolded practice, run an unassisted check — a problem with no AI help. Separates what the learner can do with support from what they can do independently. Critical for preventing phantom attainment.
0
dflow-docs
Discover and use DFlow documentation, Agent CLI, Trading API, Metadata API, Proof KYC, prediction markets, and the hosted DFlow docs MCP. Use before implementing DFlow features or when field-level endpoint details are needed.
0 · bundle
squad
Computes the SQuAD metric using torchmetrics, given predictions and ground truth. Use when evaluating question-answering outputs with exact match and F1 scores.
3