Results for “heuristic-estimation”
50 skillsMore results
heuristic-evaluation
Conduct expert heuristic evaluations of digital interfaces using Nielsen's 10 usability heuristics and domain-specific criteria.
1.7k
ux-heuristics
Evaluate and improve interface usability using heuristic analysis based on Nielsen's 10 heuristics, Krug's laws, and severity ratings.
1.6k · bundle
hehan-skill
贺涵(都市剧虚构)认知与表达框架(压缩蒸馏):导师型男主、职场金句、情感争议 触发:我的前半生 等。虚构;忌情感操控教程
9 · bundle
astar
A* pathfinding skill for heuristics and optimization.
1.7k · bundle
bleurt
Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
3
statistical-analysis
Guides statistical hypothesis testing with assumption checks, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting for research data.
30.2k · bundle
design-ux
Run a heuristic evaluation of interactive UIs against Nielsen's 10 usability heuristics and interaction add-ons, scoring the rendered artifact and producing a prioritized fix list.
42.4k
hewei-skill
贺炜(体育解说)认知与表达框架(压缩蒸馏):文学比喻嵌入赛况、克制激情、终场金句… 触发:足球诗人解说 等。非煽动球迷对立
9 · bundle
helm-liang-2022
Holistic evaluation framework for language models measuring accuracy, calibration, robustness, and fairness
10 · bundle
menli
Evaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions.
3
self-eval
Honestly evaluate AI work quality using a two-axis scoring system with mandatory devil's advocate reasoning and cross-session anti-inflation detection.
20.4k
experiment-designer
Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions, including A/B testing, sample size estimation, and statistical interpretation.
20.4k · bundle
model-evaluation
Every metric encodes an opinion about which mistake hurts.
2
heath-no-fluff
heath-no-fluff
0
heretic
Runs directional ablation and refusal-direction analysis for open-weight models the user may modify; use to reduce benign over-refusal or measure refusal/KL trade-offs, not for training.
42 · bundle
evaluating-llms-harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
0 · bundle
dialectic
Multi-phase dialectical stress-test for HIGH-STAKES decisions only (architecture choices, irreversible product calls, strategic bets). Heavy-cost skill — do NOT use for routine questions, brainstorming, or simple tradeoffs. User must explicitly invoke or describe a decision they call "high-stakes", "irreversible", or "needs stress-testing". Triggers: "stress test this decision", "dialectic on", "challenge this thesis", "should I really".
6 · bundle
ux-heuristics
Evaluate and improve interface usability using heuristic analysis. Use when the user mentions "usability audit", "UX review", "users are confused", "heuristic evaluation", "form usability", "navigation problems", "Nielsen heuristics", "cognitive walkthrough", or "usability testing". Also trigger when reviewing a design for usability issues, improving form completion rates, or evaluating information architecture and navigation. Covers Nielsens 10 heuristics, severity ratings, and information architecture. For visual design fixes, see refactoring-ui. For conversion-focused audits, see cro-methodology.
28 · bundle
epic-hypothesis
Frame an epic as a testable hypothesis with target user, expected outcome, and validation method. Use when defining a major initiative before roadmap, discovery, or delivery planning.
5.6k · bundle
hare
Computes the HARE Score, an entity- and relation-centric metric for evaluating machine-generated histopathology reports against ground truth, using GatorTronS+SapBERT embeddings and relation F1.
3
ivx-cf-evaluation
Design and implement evaluation harnesses for models, agents, and code. Use when creating benchmarks, designing eval metrics, or comparing system outputs.
0 · bundle
eval
Evaluate LLM outputs systematically — benchmarks, automated metrics, human preference, and regression tracking
1 · bundle
accuracy
Evaluates an AI judge system's pairwise ranking accuracy on generated commit messages against a heuristic ground truth from five automatic text metrics, using the MCMD dataset.
3
interpret-results
Analyzes evaluation results by requiring a stated hypothesis before examining data, then compares expectations to actual result files to prevent post-hoc rationalization.
0
hmmsim
Use when you need to characterize score distributions of a profile HMM on random sequences, such as calibration checks, benchmarking, or filter-behavior experiments.
0 · bundle
impediment-prioritization
Ranks any list of impediments and their countermeasures using a value-stream scoring model (ROI, Cost to Implement, Ease of Deployment, Risk Factor) and a fixed prioritization formula.
36.2k · bundle
hanhan-skill
韩寒(作家车手)认知与表达框架(压缩蒸馏):反套路叙事、冷幽默、公共发言锋利… 触发:三重门赛车 等。不伪造赛事实;尊重他人
9 · bundle
pua-ja
日本語の生産性コーチングモード。明示的な依頼、反復失敗、受け身、検証不足、品質不満のときに、構造化トラブルシューティングと証拠ベースの完了確認を促す。
0
speculative-decoding
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques for 1.5-3.6× speedup without quality loss.
10.4k · bundle
sentaku
選択肢(A/B/C)の深掘り比較→淘汰→推奨で判断負担を下げ判断の質を上げるスキル。5段階(L1固定3点/L1.5案拡張Diverge・自動/L2評価軸マトリクス/L3複数LLM弁証論/L4過去判断照合)。 「比較して」「深掘りして」「メリデメ教えて」「お勧めは?」「徹底的に」「過去の判断と照合」「前にどう決めたっけ」「/sentaku」等で発火。teian(浅)の深掘り要求を受け取り、brainstorming(深:設計全体)と棲み分け。
0
ads-test
Design and evaluate paid-ad experiments with hypotheses, randomization, sample-size calculations, guardrails, and decision rules for A/B and split tests.
dior
Quantifies how sensitive a language model benchmark's reliability and ranking stability are to specific design choices, such as the selection of scenarios, subscenarios, examples, and few-shot prompts. Use when the user has predictions and gold and needs to compute DIoR.
3
holonic-alignment
当寻求个人行动或项目在更宏大系统中的意义,或需要将日常工作与长远价值连接时
11 · bundle
geyou-skill
葛优(演员)认知与表达框架(压缩蒸馏):蔫坏冷面、话少留白、小人物体面与大谎 触发:甲方乙方、活着 等。虚构创作谈
9 · bundle
aer-literature
Use when positioning a manuscript against the existing economics literature, building the antecedents map for the introduction, deciding what to cite, or verifying that every reference in the bibliography is real, correctly attributed, and cited to the published version. Apply at topic selection for the novelty scan and again before drafting the introduction.
1k · bundle