Results for “model-comparison”
76 skillspymc
Build, fit, validate, and compare Bayesian models using PyMC, including hierarchical models, MCMC sampling, variational inference, posterior predictive checks, and model comparison.
253 · bundle
model-benchmark
Benchmark LLM performance across tasks — latency, quality, cost comparison.
0
pymc-bayesian-modeling
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
1 · bundle
pymc-bayesian-modeling
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
0 · bundle
pymc-bayesian-modeling
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
5 · bundle
pymc
Bayesian modeling with PyMC. Build hierarchical models, MCMC (NUTS), variational inference, LOO/WAIC comparison, posterior checks, for probabilistic programming and inference.
3 · bundle
More results
llm-council
Run Fireworks-hosted open-weight model councils that compare responses and synthesize a final answer.
42.4k · bundle
pymc
Build, fit, validate, and compare Bayesian models using PyMC's modern API, including hierarchical models, MCMC sampling, variational inference, posterior predictive checks, and model comparison.
30.2k · bundle
free-llm
Query free LLM APIs from OpenRouter, Groq, Cerebras, Google AI, and Mistral, with commands to compare models and check status.
5
backtest-persistence
Save backtest results to SQLite database for comparison. Trigger when: (1) tracking backtest history, (2) comparing model performance, (3) querying best backtests.
3
marginaleffects
Manual for the marginaleffects R and Python package, and guide to the book "Model to Meaning". Use when users ask about predictions, comparisons, slopes, marginal effects, average treatment effects (ATE/ATT/CATE), hypothesis testing, contrasts, counterfactuals, risk ratios, odds ratios, causal inference with G-computation, or need help with marginaleffects functions like predictions(), comparisons(), slopes(), hypotheses(), datagrid(), avg_predictions(), avg_comparisons(), avg_slopes(), or plot functions.
1k · bundle
alterlab-pymc
Bayesian modeling and probabilistic programming with PyMC — hierarchical models, MCMC (NUTS) sampling, variational inference, LOO/WAIC model comparison, and posterior predictive checks. Use when fitting Bayesian or hierarchical models, estimating posteriors and credible intervals, running probabilistic inference, or comparing models with LOO/WAIC. Part of the AlterLab Academic Skills suite.
60 · bundle
huggingface-best
Queries Hugging Face benchmark leaderboards to find the best AI models for a task, filters by device constraints, and returns a ranked comparison table with scores.
10.8k
sui-knowledge
Answers questions about the Sui blockchain ecosystem, including concepts, tokenomics, validators, staking, and comparisons with other chains.
1 · bundle
benchmark-models
Cross-model benchmark for gstack skills. (gstack)
0
sui-knowledge
Answers questions about the Sui blockchain ecosystem, including concepts, tokenomics, validators, staking, and comparisons with other chains.
1 · bundle
template-comparison
Compares two or more dotnet new templates side by side to help users choose between them based on parameters, feature support, frameworks, and classifications.
4k
model-evaluation
Every metric encodes an opinion about which mistake hurts.
2
market-landscape-scan
Compare a product against alternatives by positioning, capabilities, gaps, and strategic moves.
0
visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
compare-options
Systematically evaluates alternatives against weighted criteria, builds a comparison matrix, and recommends the best option with documented rationale and trade-offs.
1 · bundle
model-monitoring
The layers trade timeliness against definitiveness.
2
model-evaluation
Evaluate model quality with task-appropriate metrics and systematic error analysis. Use when: (1) comparing models, (2) analyzing failures, (3) setting go/no-go thresholds. NOT for: production monitoring implementation.
0
sui-knowledge
Answers questions about the Sui blockchain ecosystem, covering concepts, tokenomics, validators, staking, and comparisons with other chains.
10 · bundle
model-version-protocol
Model-trader version compatibility protocol: Embed version metadata in checkpoints, validate at load time. Trigger when: (1) training and live trading versions diverge, (2) models fail to load, (3) action interpretation issues.
3
train-sentence-transformers
Train or fine-tune sentence-transformers models for retrieval, similarity, clustering, classification, and reranking, with support for bi-encoders, cross-encoders, and sparse encoders.
10.8k · bundle
aa-segment-performance-comparator
Compares the performance of two or more audience segments across key metrics side by side, using Adobe Analytics data to determine winners, losers, and spreads for each metric.
142 · bundle
cja-segment-performance-comparator
Compares 2–5 audience segments side by side across key metrics, identifying winners, losers, and significant differences to inform product decisions and personalization strategy.
142 · bundle
plan-arbiter
Compare, cross-review, and merge competing plans from multiple agents into one executable direction with a clear handoff.
3.4k · bundle
diff
Reconcile a converted or built web page against its source prototype using pixel/layout and structural content diffs to catch fidelity regressions.
142 · bundle
alphagbm-compare
Compares 2-5 stocks or options across GBM Five Pillars scores, options metrics, technicals, and valuations, highlighting winners per category and providing an overall recommendation.
1.2k
nocaps-novel-object-captioning-at-scale-arxiv-1812-08658v2
Nocaps: Novel Object Captioning at Scale
6
model-research
Research latest Claude models and update selection guide — run when new models drop or periodically to stay current
8 · bundle
model-selection
Recommend model families and validation strategy based on data, constraints, and objective. Use when: (1) choosing algorithms, (2) balancing bias/variance, (3) planning benchmark baselines. NOT for: final legal/compliance sign-off.
0
mdad
Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
3
ml-monitoring
Monitor a live model for data quality, input and prediction drift, performance decay, and fire retraining triggers.
0