Packs
4 packscurated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · pack
@owl-listener
Prototyping Testing
Prototyping and testing skills: wireframe specs, usability heuristics, heuristic evaluations, accessibility audits, A/B test design, and benchmark analysis.
8 skills · pack
@owl-listener
Visual Critique
Visual critique skills: hierarchy analysis, brand consistency checks against mood/voice/tokens, composition evaluation, and typography audits — with a /critique-screen command that compiles a prioritised fix list.
7 skills · pack
@alirezarezvani
Ra Qm Team
14 regulatory affairs & quality management skills for HealthTech/MedTech: ISO 13485 QMS, MDR 2017/745, FDA 510(k)/PMA, GDPR/DSGVO, ISO 27001 ISMS, CAPA management, risk management, clinical evaluation, SOC 2 compliance.
10 skills · pack
Results for “evaluation”
27 skillsheuristic-evaluation
Conduct expert heuristic evaluations of digital interfaces using Nielsen's 10 usability heuristics and domain-specific criteria.
1.7k
scholar-evaluation
Systematically evaluate scholarly work using the ScholarEval framework, providing structured assessment across research quality dimensions including problem formulation, methodology, analysis, and writing with quantitative scoring and actionable feedback.
30.2k · bundle
interpret-results
Analyzes evaluation results by requiring a stated hypothesis before examining data, then compares expectations to actual result files to prevent post-hoc rationalization.
0
mdr-745-specialist
Classify medical devices under EU MDR 2017/745, build technical documentation, plan clinical evaluations, and manage post-market surveillance and EUDAMED integration.
20.4k · bundle
research-analyst
Conducts thorough landscape research, competitive analysis, best practices evaluation, and evidence-based recommendations. Expert in market research and trend analysis.
10
tool-evaluator
Assesses new tools, technologies, and integration options for adoption, comparing vendors and recommending implementations.
2
More results
dev-research
Researches and recommends programming technologies, libraries, and architectural approaches with verified references from official documentation.
0
at-self-eval
Summarize a contributor's Git history, a provided work log, or both into a concise, review-friendly self-evaluation for quarterly, semi-annual, or promotion cycles.
167
cto-advisor
Provides technical leadership frameworks for architecture decisions, engineering team scaling, technology strategy, and technical debt assessment.
20.4k · bundle
mos
Evaluates the naturalness, speaker similarity, and real-time synthesis speed of a Mandarin speech cloning system across diverse practical application scenarios.
3
bmad-ml-viper
Adversarial robustness and ML safety specialist. Use when the user asks to talk to Viper, requests the adversary, or needs failure mode analysis, attack surface review, and robustness evaluation.
0 · bundle
score
Audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
3
cab-eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
market-segments
Identify and analyze 3-5 distinct customer segments with demographics, jobs-to-be-done, pain points, and product fit analysis for market opportunity evaluation.
22.6k
company-research
Create a comprehensive company research brief with executive quotes, product strategy, and organizational context for competitive analysis, partnership evaluation, interview preparation, or market entry decisions.
5.6k · bundle
market-sizing
Estimate market size using TAM, SAM, and SOM with top-down and bottom-up approaches for market opportunity assessment, investor pitches, or market entry evaluation.
22.6k
review
Routes quality reviews to specialized critic agents based on file type or flags, covering peer review, code review, and manuscript polish.
3 · bundle
research-router
Route research prompts to evidence, retrieval, scraping, data, market, or scientific research skills. Use when prompts mention deep research, current facts, citations, Exa, iterative retrieval, search-first, scraping, data pipelines, PubMed, USPTO, gget, literature review, or scholar evaluation.
0 · bundle
pytdc
Access AI-ready drug discovery datasets and benchmarks from Therapeutics Data Commons, covering ADME, toxicity, drug-target interactions, and molecular generation with standardized splits and evaluation metrics.
30.2k · bundle
ux-researcher
Conducts end-to-end user research, from study design through behavioral insight synthesis, using frameworks like Nielsen's Heuristics, SUS, and User Journey Mapping.
2
paper-review-sim
Simulates a NeurIPS/SC/ICSE-style peer review with five reviewer personas (HPC, ML, Stats, Reproducibility, Devil's Advocate) that verify every claim against actual result data before submission.
0
bis-eval
Benchmarks energy-function-based safe control algorithms on the BIS (Benchmark of Interactive Safety) dataset, scoring safety, efficiency, and hybrid performance in human-robot and robot co-working scenarios.
3
gov-program-knowledge
Provides domain knowledge of Korean government and private funding program announcement systems, evaluation criteria, and key program characteristics, including TIPS, Early-Stage Startup Package, Startup Growth Technology Development, and AI Voucher.
0
scientific-critical-thinking
Evaluate scientific claims and evidence quality by assessing experimental design, identifying biases and confounders, and applying evidence grading frameworks like GRADE and Cochrane Risk of Bias.
30.2k · bundle
b2
VS-Enhanced Evidence Quality Appraiser - Prevents Mode Collapse with context-adaptive quality assessment Enhanced VS 3-Phase process: Avoids automatic tool application, delivers research-specific evaluation strategies Use when: appraising study quality, assessing risk of bias, grading evidence Triggers: quality appraisal, RoB, GRADE, Newcastle-Ottawa, risk of bias, methodological quality
1k
fact-check
Verifies claims, articles, screenshots, and URLs through source-grounded analysis with an evidence ledger, source credibility evaluation, and manipulation detection. Supports quick checks, full fact-check cards, two-source comparisons, and prebunking in multiple languages and policy contexts.
74 · bundle
academic-paper-strategist
Systematic strategic planning framework for philosophy and interdisciplinary academic papers targeting preprint platforms (PhilArchive, arXiv, PhilSci-Archive). Use when users want to: (1) plan a paper on a specific topic, (2) identify research gaps and assess originality, (3) develop optimized paper outlines, (4) prepare for preprint submission, or (5) understand platform requirements and writing standards. Triggered by phrases like 'plan a paper on,' 'help me design a paper about,' 'identify research gaps in,' 'is this idea original,' or when users need structured research planning. The skill guides through three phases: Platform Analysis (identifying target venue and studying sample papers), Theoretical Framework (AI-driven literature search and gap identification), and Outline Optimization (structured design with reviewer-perspective self-assessment). Each phase includes quality evaluation standards and validation checkpoints. Output: optimized detailed outline ready for systematic writing (use with acade
1k · bundle