Results for “bias-score”

52 skills
More results
dvy1987
Eval Judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
qhjqhj00
Eas
Validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity, using the AmbGIMT dataset and pairwise score comparisons.
3
sethmblack
Bias Audit
Audits decisions and situations for operating psychological biases using Munger's 25 tendencies framework, producing a structured analysis with countermeasures.
6
ekatasingh1107
Lead Scorer
Score raw leads as HOT/WARM/COOL based on config-driven weights from agency.config.json
2 · bundle
trailofbits
Interpreting Culture Index
Interprets Culture Index survey data, behavioral profiles, and personality assessments from JSON or PDF. Supports individual profile interpretation, team composition analysis, burnout detection, hiring profiles, manager coaching, interview transcript analysis, and conflict mediation.
6k · bundle
rulebase-co
Cx Customer Health Score
Use to design or audit the support contribution to a customer health score, so the score predicts something instead of averaging weakly-related signals into a colour. Trigger for "build a customer health score", "add support signal to health scoring", "our health scores don't predict churn", red/amber/green account scoring, or a health score nobody trusts.
1
alphagbm
Alphagbm Fear Score
Calculates a per-ticker panic index (0-100) from six weighted signals including VIX, IV Rank, RSI-14, volume anomaly, put/call ratio, and consecutive down days, triggering Bull Put Spread entry signals at scores ≥60.
1.2k
rulebase-co
Cx Outsourcer Scorecard
Use to compare BPO sites, vendors or partner teams fairly, adjusting for the work mix each is given before concluding anything about performance. Trigger for "compare our BPO sites", "which vendor is performing best", "site A scores lower than site B", outsourcer QBR packs, partner MI reporting, or setting contractual quality targets with a vendor.
1
antigravity
UI Score
Score a UI file's design quality 0-100 against StyleSeed's design language with per-category breakdown, worst offenders, and prioritized fix list.
42.4k
alphagbm
Alphagbm Marks Cycle
Provides a single 0-100 cycle score blending VIX, SPY IV Rank, Put/Call ratio, and valuation percentile to determine offense vs. defense posture, based on Howard Marks' market cycle framework.
1.2k
lionelndong
Draft Score
Lightweight ContentShake AI self-check the /draft stage can call before saving. Returns just SEO + Quality scores (no full optimization) so the writer knows whether the draft is in winning territory before /quality-check runs. Fails soft when SEMRUSH_API_KEY is unset.
0
gtynnn060110-hash
Modify Skill
Update or correct an existing skill file based on judge feedback or improved understanding.
6 · bundle
rulebase-co
Cx Survey Design
Use to design or fix a customer support survey — CSAT, NPS, or CES — and to diagnose response bias. Trigger for "design a CSAT survey", "our CSAT doesn't match reality", "should we use NPS", survey response rate, non-response bias, survey timing or scale choice, comparing satisfaction across teams or channels, or interpreting a satisfaction trend.
1 · bundle
alphagbm
Alphagbm Options Score
Score and rank options contracts for any ticker using a multi-factor model covering liquidity, IV attractiveness, Greeks balance, and risk/reward. Returns scored option chains with the best contracts highlighted.
1.2k
snoodleboot-io
Model Evaluation
Every metric encodes an opinion about which mistake hurts.
2
alirezarezvani
Self Eval
Honestly evaluate AI work quality using a two-axis scoring system with mandatory devil's advocate reasoning and cross-session anti-inflation detection.
20.4k
qhjqhj00
Visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
tradermonty
Dual Axis Skill Reviewer
Review AI agent skills using a dual-axis method: deterministic code-based checks and LLM deep review, with weighted scoring and improvement recommendations.
2.3k · bundle
rulebase-co
Cx QA Appeal Process
Use to design or audit a QA dispute and appeal workflow with timeboxes, adjudication standards, and second-level consistency so appeals improve trust instead of rewriting scores without rules. Trigger for "QA appeal process", "agents disputing scores", "who adjudicates QA disputes", "overturn rate too high", second-level review standards, or calibration erosion from ad-hoc score changes.
1
k-dense-ai
Statistical Analysis
Guides statistical hypothesis testing with assumption checks, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting for research data.
30.2k · bundle
rulebase-co
Cx Effort Score
Use to measure customer effort from behavioural signals instead of CES surveys, and to audit whether a composite effort score is honest. Trigger for "customer effort score", "behavioural CES", effort without survey, repeat contacts and channel switches, transfers and reopens, or "our CES doesn't match operational data".
1
lionelndong
Quality Check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
x3allamerican
Csa Bsi Scoring
Use this skill when the user asks about FMCSA Compliance Safety Accountability (CSA) program scoring — the seven BASIC categories, BASIC Severity Indicator (BSI), peer percentiles, intervention thresholds, how violations age out, SMS Methodology, ISS (Inspection Selection System), what "alert" status means, Safety Measurement System mechanics, or DataQ disputes. Cite SMS Methodology v3.20.
1
qhjqhj00
Bleurt
Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
3
smith6jt-cop
Bimodal Score Diagnosis
Diagnosing and fixing bimodal matching score distributions in MaxFuse
3
qhjqhj00
Bss Eval
Evaluates speech language models on beyond-semantic speech attributes such as dialect comprehension, multi-turn context memory, emotion perception, age-aware response generation, and non-verbal cue handling, reporting accuracy and judge-based scores.
3
gonglingrui
Score Analyzer
Score Analyzer
349 · bundle
lucassantana-dev
Eval
Evaluate LLM outputs systematically — benchmarks, automated metrics, human preference, and regression tracking
1 · bundle
seb1n
Lead Scoring
Score and prioritize leads based on firmographic fit and behavioral engagement signals, producing ranked tiers for sales team focus. Use when the user requests lead scoring or provides relevant inputs for this workflow.
159
brycewang-stanford
B2
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias, grading evidence Triggers: quality appraisal, RoB, GRADE, Newcastle-Ottawa, risk of bias, methodological quality
1k
smith6jt-cop
Reward Function Hold Bias
Fix HOLD bias in RL reward function. Trigger when: (1) model learns to always HOLD, (2) trade rate is too low (<10%), (3) slippage penalty exceeds typical price moves.
3
lionelndong
Keyword Prioritization
Deterministically score, route, rank, and select at most one fully vetted Pleasur.ai Stage 01 blog-keyword candidate using product-fit-dominant business value, traffic opportunity, brand fit, DR-relative winnability, and a free-seeker penalty. Use only after BID and AIO evaluation are complete.
0 · bundle
phuryn
Ab Test Analysis
Analyze A/B test results with statistical significance, sample size validation, confidence intervals, and ship/extend/stop recommendations.
22.6k
intelli-verse-x
Ivx Ams Content Ops
Score and iteratively improve marketing content with an expert panel until 90+. Use for content quality gates, expert panel reviews, editorial scoring, or when another skill needs a content QA loop.
0 · bundle