Results for “experiment-analysis”

79 skills
cloudthinker-ai
managing-eppo
Manage feature flags, experiments, and metrics in Eppo via its REST API, including flag configuration, experiment design, statistical analysis, and metric pipelines.
7
neuralblitz
analytical
Applies quantitative and qualitative analysis techniques, interprets experimental data, validates procedures, and selects appropriate methods with uncertainty quantification.
1
brycewang-stanford
c1
VS-Enhanced Quantitative Design Consultant with Materials & Sampling Enhanced VS 3-Phase process: Avoids obvious experimental designs, proposes context-optimal quantitative strategies Absorbed C4 (Experimental Materials Developer) and D1 (Sampling Strategy Advisor) capabilities Use when: selecting quantitative research design, planning experimental/survey methodology, power analysis, developing materials, sampling Triggers: RCT, quasi-experimental, experimental design, survey design, power analysis, sample size, factorial design, materials, stimuli, sampling strategy
1k
github
phoenix-evals
Build and run evaluators for AI/LLM applications using Phoenix, covering error analysis, custom evaluators, experiments, and production monitoring.
36.2k · bundle
nvidia
tao-analyze-changenet-rca
Performs deep root cause analysis on NVIDIA TAO Visual ChangeNet classification experiments, using image-evidence-driven investigation to diagnose model failures and produce actionable reports.
2.2k · bundle
ichichuang
research-paper-writing
End-to-end pipeline for writing ML/AI research papers — from experiment design through analysis, drafting, revision, and submission. Covers NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Integrates automated experiment monitoring, statistical analysis, iterative writing, and citation verification.
0 · bundle
More results
openai
jupyter-notebook
Create, scaffold, and edit Jupyter notebooks for experiments, exploratory analysis, or tutorials using bundled templates and a helper script.
23.3k · bundle
modbender
cro-chief-revenue-officer
Optimize conversion rates with funnel analysis, A/B testing, statistical significance, and compliance-safe experiments.
12 · bundle
galyarderlabs
signup-flow-cro
Analyzes signup and onboarding flows to identify friction, motivation gaps, guardrails, and experiment opportunities for improving conversion.
20
omer-metin
a-b-testing
The science of learning through controlled experimentation. A/B testing isn't about picking winners—it's about building a culture of validated learning and reducing the cost of being wrong. This skill covers experiment design, statistical rigor, feature flagging, analysis, and building experimentation into product development. The best experimenters know that every test, positive or negative, teaches something valuable. Use when "a/b test, experiment, hypothesis, statistical significance, sample size, feature flag, variant, control, treatment, p-value, conversion rate, test winner, split test, experimentation, testing, statistics, feature-flags, hypothesis, growth, optimization, learning, validation" mentioned.
128 · bundle
shenmuxing
analysis-plan
Turn a theory-first research idea into a proof-oriented research plan. Use after idea-creator-analysis, or when the user asks for theorem targets, assumptions, proof obligations, impossibility routes, analysis validation, or a non-experimental plan.
2 · bundle
brycewang-stanford
acl-experiments
Use when designing or auditing experiments for an ACL paper, covering tuned LLM baselines, multi-dataset and multilingual evaluation, statistical significance and variance, human evaluation with agreement reporting, contamination and prompt-sensitivity controls, ablations, and error-analysis expectations in NLP reviewing.
1k
shenmuxing
idea-discovery-analysis
Orchestrate a full theory-first idea discovery pipeline. Use when the user wants the analysis counterpart of idea-discovery: literature context, proof-oriented idea generation, novelty checking, DeepSeek critique, and an analysis plan instead of an experiment-first workflow.
2 · bundle
akillness
data-analysis
Guide through a structured data analysis workflow: define the question, validate data quality, select the appropriate analytical method, and produce decision-ready findings with caveats.
42 · bundle
lingxling
denario
Automates scientific research workflows from data analysis to publication, orchestrating multiple agents for hypothesis generation, methodology development, computational experiments, and LaTeX paper writing.
253 · bundle
shenmuxing
idea-creator-analysis
Generate and rank theory-first or proof-oriented research ideas. Use when the user wants non-experimental research ideas, theoretical methods, proof programs, theorem candidates, impossibility results, convergence/sample-complexity analyses, or "analysis" variants of idea creation.
2 · bundle
alirezarezvani
ab-test-setup
Design statistically valid A/B tests with hypothesis frameworks, sample size calculations, and analysis checklists.
20.4k · bundle
github
phoenix-cli
Debug LLM applications using the Phoenix CLI: fetch traces, analyze errors, structure trace review with open and axial coding, inspect datasets, review experiments, and query the GraphQL API.
36.2k · bundle
ahang1598
doubao-data-analysis
结构化业务数据分析:附件读取与口径核验、定向筛选、规则/阈值判定、指标异动归因、漏斗/留存/实验分析、经营复盘及可审计报告。当用户提供 Excel、CSV、PDF、图片或多份业务材料,要求查数、判异常、解释变化、比较方案或形成行动建议时使用。Use for evidence-grounded analysis of structured business data, including filtering, rule checks, reconciliation, diagnostics, experiments, and decision reports.
9 · bundle
alirezarezvani
experiment-designer
Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions, including A/B testing, sample size estimation, and statistical interpretation.
20.4k · bundle
agricidaniel
ads-test
Design and evaluate paid-ad experiments with hypotheses, randomization, sample-size calculations, guardrails, and decision rules for A/B and split tests.
k-dense-ai
statistical-analysis
Guides statistical hypothesis testing with assumption checks, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting for research data.
30.2k · bundle
dvy1987
experiment-readout
Analyse experiment results, run validity checks (SRM, exposure parity, data integrity, novelty/primacy), interpret causally, make a ship/iterate/kill decision against the pre-declared rule, and append to cumulative learnings. Forces honest readouts — strips significance claims from underpowered or peek-violating tests; never lets directional results masquerade as causal wins. Load when results exist, or when the user says "read out this experiment", "analyse the test", "did the test win", "interpret the results", "what did we learn", "ship or kill", or when the experimentation orchestrator routes here.
3 · bundle
brycewang-stanford
analyze-results
Analyze ML experiment results, compute statistics, generate comparison tables and insights. Use when user says "analyze results", "compare", or needs to interpret experimental data.
1k
nexu-io
experiment-readout
Transforms A/B test and product experiment data into actionable readouts with hypothesis, metrics, interpretation, and decision.
· bundle
github
arize-experiment
Creates, runs, and analyzes Arize experiments for evaluating and comparing model performance using the ax CLI.
36.2k · bundle
phuryn
ab-test-analysis
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
22.6k
zhouziyue233
results-analysis
Comprehensive results analysis for empirical research: generate publication-quality descriptive statistics and balance tables, interpret regression coefficients with economic magnitude and effect sizes, assess identification assumption diagnostics, and produce structured results memos. Use when asked to create summary statistics, Table 1, balance tests, interpret results, assess economic significance, or write results narratives.
7
qhjqhj00
ara-compiler
Compiles any research input — PDF papers, GitHub repositories, experiment logs, code directories, or raw notes — into a complete Agent-Native Research Artifact (ARA) with cognitive layer (claims, concepts, heuristics), physical layer (configs, code stubs), exploration graph, and grounded evidence. Use when ingesting a.
3 · bundle
lambenthan
exp-design
Claim-driven 实验设计:界定目标 claims → 设计实验块(baseline/validation/ablation/robustness)→ 构建执行顺序 → 可选 Review LLM review → 写入 wiki
77
michaelschecht
autoresearch
Autonomously runs iterative experiment loops to optimize code against a measurable metric. Use when the user wants to improve execution time, memory usage, test pass rate, or any numeric performance goal across repeated experiments — NOT for one-shot bug fixes or simple code review.
0
lambenthan
exp-eval
实验判决门:Review LLM 独立评判实验结果 → 4 种判决路径 → 自动更新 claims confidence、ideas status、graph edges
77
subvisual
research-synthesis
research-synthesis
0 · bundle
michaelschecht
ab-testing-statistics
Design and evaluate A/B tests with power, sample size, and robust metric interpretation. Use when: (1) planning controlled experiments, (2) reading p-values/effects, (3) sequential testing safeguards. NOT for: dark-pattern optimization.
0
peteedoo
ab-test-analysis
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations. Use when evaluating experiment results, checking if a test reached significance, interpreting split test data, or deciding whether to ship a variant.
0
dvy1987
experimentation
Orchestrator for the experimentation skill suite — turn assumptions and product questions into rigorous, well-instrumented experiments and decision-grade readouts. Routes through backlog → spec → runbook → readout based on user need and existing artefacts. Platform-agnostic with PostHog as the primary binding. Load when the user asks to design an experiment, A/B test something, set up an experiment, run a holdout, test a hypothesis, decide what to test next, read out experiment results, analyse a test, or says "should we A/B test this", "experiment on the landing page", "is this lift real", "ship or kill this test", "what should we test next", "build an experiment backlog", "test the pricing page", "validate this with an experiment".
3 · bundle