Results for “scoring”
225 skillsEval Judge
Score LLM and agent outputs using LLM-as-judge techniques — direct scoring against rubrics or pairwise comparison between two outputs. Includes built-in bias mitigation for position bias, length bias, and self-enhancement bias. Load when the user asks to score an output, judge a response, evaluate against a rubric, compare two outputs, do direct scoring, run pairwise comparison, or says "rate this", "which response is better", "score this against the rubric", "judge this output", "LLM as judge this". Sub-skill of eval-output orchestrator.
3 · bundle
Experiment Designer
Design, prioritize, and evaluate product experiments with clear hypotheses and defensible decisions, including A/B testing, sample size estimation, and statistical interpretation.
20.4k · bundle
Faf Expert
Configure and optimize .faf files, MCP servers, and bi-directional sync for AI context across multiple platforms, with championship scoring to achieve 85%+ AI-readiness.
42.4k
Agentic Eval
Implement iterative evaluation and refinement loops for AI agent outputs, using self-critique, evaluator-optimizer patterns, and rubric-based scoring to improve quality.
36.2k
Aeon Autoresearch
Generates four improved variations of any installed skill, scores them against a weighted rubric, and applies the winning version while preserving the original.
1.2k · bundle
Prioritizing Vulnerabilities With Cvss Scoring
Calculate CVSS scores, interpret vector strings, and prioritize vulnerabilities using CVSS alongside EPSS and CISA KEV for effective risk-based remediation.
24.6k · bundle
Implementing Epss Score For Vulnerability Prioritization
Integrate FIRST's Exploit Prediction Scoring System (EPSS) API to prioritize vulnerability remediation based on real-world exploitation probability within 30 days.
24.6k · bundle
Kb Structure
Structures company knowledge bases for business plan automation, defining a 7-category schema, completeness scoring, and storage layouts for Notion or local Markdown.
0
Bizplan Writing
Guides writing Korean government funding program business plans (사업계획서) from an evaluator's perspective, covering structure, scoring emphasis, and writing principles.
0
Cab Eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
Business Growth Skills
4 business growth agent skills and plugins for Claude Code, Codex, Gemini CLI, Cursor, OpenClaw. Customer success (health scoring, churn), sales engineer (RFP), revenue operations (pipeline, GTM), contract & proposal writer. Python tools (stdlib-only).
3 · bundle
Ivx Ams Content Ops
Score and iteratively improve marketing content with an expert panel until 90+. Use for content quality gates, expert panel reviews, editorial scoring, or when another skill needs a content QA loop.
0 · bundle
QA Agent
The analyst that watches your other analysts. It reads the trail your reports and workflows leave behind, then surfaces scoring blind spots, CRM hygiene gaps, and workflow drift, and turns every miss into a training signal. Built for GTM teams running any stack of reports, customizable to your process. Trigger on "run QA", "weekly QA report", "system health check", "where are our scoring blind spots", "what should we coach on this week", "audit our workflow performance", or any system-level health or feedback-loop question.
0
Implementing Attack Surface Management
Builds an external attack surface management (EASM) program using Shodan, Censys, and ProjectDiscovery tools for asset discovery, subdomain enumeration, service fingerprinting, and exposure scoring.
24.6k · bundle
Detecting AI Model Prompt Injection Attacks
Detects prompt injection attacks targeting LLM-based applications using regex pattern matching, heuristic scoring, and DeBERTa transformer classification.
24.6k · bundle
Performing Cve Prioritization With Kev Catalog
Integrate the CISA Known Exploited Vulnerabilities catalog with EPSS and CVSS to prioritize CVE remediation based on real-world exploitation evidence.
24.6k · bundle
Universal Single Cell Annotator
Annotates single-cell RNA-seq data by scoring marker genes, transferring labels with CellTypist, or reasoning over cluster markers with an LLM.
567 · bundle
Devops
Audits deployment readiness across CI/CD, containers, monitoring, IaC, secrets, CDN, and DNS, scoring each area and identifying critical gaps, with optional auto-fix via chained sub-skills.
13
Gwas Prs
Calculate polygenic risk scores from direct-to-consumer genetic data using published scoring files from the PGS Catalog and contextualize results against population reference distributions.
17 · bundle
Product Lens
Validates product direction before building, runs product diagnostics, founder reviews, user journey audits, and feature prioritization, producing actionable briefs and recommendations.
1
Aya Eval
Evaluates open-ended generation quality of multilingual LLMs across brainstorming, planning, and long-form tasks, using AYA and DOLLY datasets with qualitative fluency and quality scoring.
3
Prospecting
Build qualified prospect lists for B2B SaaS, general B2B, or local small businesses by defining an ICP, sourcing candidates, qualifying with evidence, scoring, and outputting a lead sheet.
36.3k · bundle
Trustlayer Sybil Scanner
Detects fake reviews, Sybil rings, rating manipulation, and reputation laundering in ERC-8004 agent ratings across 20+ blockchains using the TrustLayer API.
1.2k · bundle
Performing Indicator Lifecycle Management
Tracks indicators of compromise from initial discovery through validation, enrichment, deployment, monitoring, and retirement to maintain a high-quality, actionable indicator database.
24.6k · bundle
Document
Audits documentation health by scanning for README, changelog, API docs, ADRs, runbooks, onboarding guides, and diagrams, scoring coverage against project maturity, and recommending sub-skills to fill gaps.
13
Self Review
Reviews an academic paper using the NeurIPS review form with three reviewer personas, ensemble scoring, and reflection refinement. Extracts text from PDF, runs structured review, and outputs actionable feedback.
3 · bundle
Advanced Evaluation
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", or mentions direct scoring, pairwise comparison, position bias, evaluation pipelines, or automated quality assessment.
55 · bundle
Detecting Insider Threat With Ueba
Detect insider threats by modeling normal user and entity behavior with Elasticsearch, computing anomaly scores, and correlating low-confidence indicators into high-confidence alerts.
24.6k · bundle
Alphagbm Buffett Analysis
Scores any US stock ticker through Warren Buffett's four-lens framework (business simplicity, moat, management, valuation) and returns a weighted HOLDABLE/WATCHABLE/AVOID verdict.
1.2k
Production Audit
Audits a codebase for production readiness using local evidence, scoring ship/block risk and listing concrete fixes without sending repo data to external services.
0
Company Research
Discovers and researches companies matching an ideal customer profile, using Browserbase Search API and a Plan→Research→Synthesize pattern to produce scored reports and CSV exports.
3.6k · bundle
Design UX
Run a heuristic evaluation of interactive UIs against Nielsen's 10 usability heuristics and interaction add-ons, scoring the rendered artifact and producing a prioritized fix list.
42.4k
Domain Driven Design
Model software around the business domain using bounded contexts, aggregates, and ubiquitous language, with scoring and diagnostic tools for evaluating domain model quality.
1.6k · bundle
Dx
Audits a project's developer experience by scoring devcontainer, git hooks, linting, build caching, environment setup, and release pipeline, then generates a DX health report with prioritized recommendations.
13
Threat Modeling
Threat modeling workflow for software systems: scope, data flow diagrams, STRIDE analysis, risk scoring, and turning mitigations into backlog and tests. Use when designing new features, reviewing architecture changes, handling sensitive data, or hardening auth/payment/multi-tenant flows.
71 · bundle
Multi Search
Parallel multi-source search combining Web, Scholar, Smart, and Tavily results with confidence scoring and AI synthesis. Best for comprehensive research requiring cross-source validation. Use when: the user needs web search, research, source discovery, or content extraction.
1 · bundle