Results for “weights-and-biases”
51 skillsazure-mgmt-weightsandbiases-dotnet
Manage Weights & Biases ML experiment tracking instances on Azure using the .NET SDK. Create, configure, list, update, and delete W&B instances with marketplace integration and SSO.
2.7k
weights-biases-run-monitor
Streams live training metrics, system stats, and gradients from active W&B runs, alerts on metric regressions, and posts summaries to Slack.
28
weights-and-biases
Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform
1 · bundle
More results
weights-and-biases
Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform
0 · bundle
weights-and-biases
Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B.
10.4k · bundle
model-evaluation
Every metric encodes an opinion about which mistake hurts.
2
bias-audit
Audits decisions and situations for operating psychological biases using Munger's 25 tendencies framework, producing a structured analysis with countermeasures.
6
pros-cons
Structure a decision into steelmanned, impact-weighted pros and cons - irreversible consequences flagged, the decisive unknown named, and a clear recommendation you can reject.
0
endo-sdm-aed
This skill guides clinicians in sharing decision‑making when selecting an antiepileptic drug (AED) by providing quantitative estimates of each drug’s expected weight effect. It is triggered when a clinician asks how to discuss weight‑change risks when choosing an AED or what information to provide on weight effects of specific agents such as valproate versus lamotrigine.
10
ads-budget
Plan and review paid-media budgets, bidding, pacing, marginal return, forecasts, CPA, ROAS, MER, LTV:CAC, constraints, and allocation across supported platforms.
feature-importance-analysis
"Feature importance" is ambiguous.
2
bbq-eval
Evaluates social bias in question-answering models using the BBQ benchmark, measuring accuracy and a bias score across ambiguous and disambiguated contexts to reveal reliance on stereotypes.
3
alphagbm-buffett-analysis
Scores any US stock ticker through Warren Buffett's four-lens framework (business simplicity, moat, management, valuation) and returns a weighted HOLDABLE/WATCHABLE/AVOID verdict.
1.2k
business-modeling
Pick the right business-model canvas (Lean Canvas, Business Model Canvas, or Value Proposition Canvas) for the stage and fill it with specifics — one segment, one primary canvas, top-3 assumptions, no fluff in the moat or channel boxes. Load when the user asks to fill a business model canvas, lean canvas, value proposition canvas, model this business, map the business model, says "fill the BMC", "make a Lean Canvas", "Value Proposition Canvas for this", "model this idea", "what's the business model", "design the business model". Sub-skill of `venture-exploration`. Hard-bans "everyone" segments, generic channels ("SEO/social/content/ads"), and "unfair advantage = AI/data/network effects" with no concrete asset. Does NOT score viability — for that use `idea-evaluation`.
3 · bundle
cab-eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
endo-sdm-antidepressant
Recommends a shared decision‑making process that provides patients with quantitative estimates of the expected weight effect of antidepressants to inform drug choice, also considering expected treatment length. Triggers include clinician questions such as ‘How should I discuss weight‑change risks with this patient starting an antidepressant?’ or ‘What tool can I use to show expected weight impact of sertraline vs bupropion?’.
10
eas
Validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity, using the AmbGIMT dataset and pairwise score comparisons.
3
weights-and-biases
Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform
0 · bundle
interpreting-culture-index
Interprets Culture Index survey data, behavioral profiles, and personality assessments from JSON or PDF. Supports individual profile interpretation, team composition analysis, burnout detection, hiring profiles, manager coaching, interview transcript analysis, and conflict mediation.
6k · bundle
alphagbm-marks-cycle
Provides a single 0-100 cycle score blending VIX, SPY IV Rank, Put/Call ratio, and valuation percentile to determine offense vs. defense posture, based on Howard Marks' market cycle framework.
1.2k
visuals-adversarial
Skeptical pushback on visual placement — both density and quality. Reads the annotated outline plus the visuals manifest and asks (a) whether the article hits the density target from editorial-principles-visuals.md, (b) whether each [VISUAL:...] earns its place, (c) whether sections without one would benefit. One revision pass on FAIL (BLOG_AGENT_VISUALS_REVISION_BUDGET, default 1).
0
risk-matrix
Identify and prioritize risks by impact and controllability. Use for risk management, project planning, and strategic decision support.
1 · bundle
hanhan-skill
韩寒(作家车手)认知与表达框架(压缩蒸馏):反套路叙事、冷幽默、公共发言锋利… 触发:三重门赛车 等。不伪造赛事实;尊重他人
9 · bundle
axiom
First-principles assumption auditor. Classifies each hidden assumption (fact / convention / belief / interest-driven), ranks by fragility × impact, and rebuilds conclusions from verified premises. Bilingual: auto-detects Chinese or English.
0 · bundle
heretic
Runs directional ablation and refusal-direction analysis for open-weight models the user may modify; use to reduce benign over-refusal or measure refusal/KL trade-offs, not for training.
42 · bundle
endo-sdm-antipsychotic
Recommends using weight‑neutral antipsychotic alternatives when possible and employing shared decision‑making that provides quantitative estimates of expected weight effect to guide drug choice. Triggered when a clinician asks, “How do I involve this patient in choosing an antipsychotic with minimal weight gain?” or “What resources show weight‑change projections for risperidone vs aripiprazole?”.
10
ptw-analysis
Price-to-win lens using GSA CALC+, BLS OEWS, and incumbent USASpending award patterns for a pursuit. Use when user asks for realism checks or competitive pricing posture before proposal — draft skill, not production-verified.
0
menli
Evaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions.
3
priority-decision-system
Prioritize product work with explicit criteria, scoring, tradeoffs, and decision rationale.
0
testing-strategies
Three shapes get argued about as if one were correct.
2
xusanduo-skill
许三多(军旅剧虚构)认知与表达框架(压缩蒸馏):钝感坚持、不抛弃不放弃、草根成长 触发:士兵突击 等。禁止军事违法教程
9 · bundle
cailan-skill
蔡澜(美食 / 生活家)认知与表达框架(压缩蒸馏):享乐主义正当化、旅行搭子叙事 触发:食神专栏 等。非过量饮酒医疗建议
9 · bundle
panel-data
Econometrics skill for panel data models. Activates when the user asks about: "panel data", "fixed effects", "random effects", "Hausman test", "within estimator", "between estimator", "two-way fixed effects", "clustered standard errors panel", "FE model", "RE model", "pooled OLS", "unobserved heterogeneity", "panel regression", "first difference estimator", "entity fixed effects", "time fixed effects", "面板数据", "固定效应", "随机效应", "豪斯曼检验", "双向固定效应", "面板回归", "个体效应", "时间效应", "一阶差分"
1k · bundle
portfolio-and-measurement
Improve existing content and close the learning loop without cannibalizing new-content work.
0
transaction-consistency
Use with analysis-agent or task-agent for task-local transaction, isolation, and conflict decisions. Do not use without a transaction decision or as task owner.
4 · bundle
eval-judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle