Results for “medical-benchmark”

51 skills
More results
dotnet
Microbenchmarking
Create, run, configure, and review BenchmarkDotNet microbenchmarks for .NET code, covering project setup, comparison strategies, and cost-aware execution.
4k · bundle
affaan-m
Benchmark Methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
qhjqhj00
Mdad
Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
3
levalencia
Doctorg
Evidence-based health research using tiered trusted sources with GRADE-inspired evidence ratings. Integrates Apple Health data for personalized context. Use when user asks health, nutrition, exercise, sleep, or wellness questions.
3 · bundle
aniruddhaadak80
Model Benchmark
Benchmark LLM performance across tasks — latency, quality, cost comparison.
0
intelli-verse-x
Ivx Cf Benchmarking
Design and execute performance benchmarks for models and systems. Use when measuring throughput, latency, or comparing system variants.
0 · bundle
nvidia
Digital Health Clinical Asr Eval
Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
2.2k · bundle
dotnet
Build Perf Diagnostics
Diagnose MSBuild build performance bottlenecks using binary log analysis, covering timeline analysis, performance summary interpretation, and seven common bottleneck categories.
4k
nvidia
Jetson LLM Benchmark
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
2.2k · bundle
rulebase-co
Cx Benchmark Methodology
Use to compare CX performance to a published or vendor benchmark without fooling yourself — scope mismatch, survivor bias, and definition mismatch usually make external benchmarks incomparable, and internal baselines often beat them. Trigger for "how do we compare to industry", "is our CSAT good", benchmark slide for the board, vendor benchmark report, "are we above average", outsourcing RFP benchmarks, or when someone cites a round-number industry standard.
1
alphagbm
Alphagbm Health Check
Audits a research knowledge base for stale profiles, thesis drift, and orphan pages, returning a 0-100 health score with actionable recommendations.
1.2k
livelybug
Benchmark Models
Cross-model benchmark for gstack skills. (gstack)
0
affaan-m
Benchmark
Measure performance baselines, detect regressions before and after PRs, and compare stack alternatives using browser, API, and build benchmarks.
226k
tradermonty
Market Breadth Analyzer
Quantifies market breadth health using public CSV data, generating a 0-100 composite score across six components to assess market participation and rally breadth.
2.3k · bundle
alirezarezvani
Qms Audit Expert
Provides ISO 13485 internal audit methodology for medical device quality management systems, covering audit planning, execution, nonconformity classification, and external audit preparation.
20.4k · bundle
tianhao909
Evaluating Llms Harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
1 · bundle
anantha-236
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
1
construct-ai-primary
Performance Benchmarking
Use when evaluating, measuring, or comparing the performance of systems, functions, or services. This skill provides a framework for establishing baselines, measuring performance, and validating that changes meet performance requirements.
0
lionelndong
Quality Check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
brycewang-stanford
Nejm Fit
Use this first, before any writing, to stress-test whether a clinical study clears NEJM's bar — practice-changing clinical impact, methodological rigor, and generalizability. Decides NEJM vs Lancet/JAMA vs a specialty journal.
1k
ahang1598
Doubao Medical Report
必须在用户需要医学报告解读时使用。包括:用户上传体检报告、检验报告、检查单、化验单、血常规/尿常规/生化/肝肾功能/血脂血糖等检验检查图片、照片、截图、PDF、文档、表格或文件;用户只发报告图片/附件且没有文字说明;用户说“帮我看看”“看下这个报告”“这个结果正常吗”“有什么问题”“报告怎么解读”;用户表达体检报告解读、医院报告解读、影像/超声/CT/MRI/内镜/病理报告解读等需求。用于梳理报告内容,解释异常指标和检查发现,识别需要关注的风险信号,并给出就医沟通、复查随访、观察监测和生活方式管理建议。
9 · bundle
qcmuu
Evaluating Llms Harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
0 · bundle
mhassan0000
Benchmark
Measures performance baselines, detects regressions before and after PRs, and compares stack alternatives.
1
rajanthar
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
0
alirezarezvani
Mdr 745 Specialist
Classify medical devices under EU MDR 2017/745, build technical documentation, plan clinical evaluations, and manage post-market surveillance and EUDAMED integration.
20.4k · bundle
chen-yu-hao
Pytdc
Therapeutics Data Commons. AI-ready drug discovery datasets (ADME, toxicity, DTI), benchmarks, scaffold splits, molecular oracles, for therapeutic ML and pharmacological prediction.
5 · bundle
antigravity
AI Analyzer
Integrates multi-dimensional health data to detect anomalies, predict risks (hypertension, diabetes, cardiovascular), and generate personalized recommendations and interactive HTML reports.
42.4k
pwdev-solucoes
Performance Engineer
Benchmark, load test, capacity plan, and cache with k6, JMeter, Locust, and pgbench. Use when the user says "slow", "performance", "load test", "stress test", "how many users can it handle", "capacity", "cache", "benchmark", "k6".
2
danielpradilla
Product Health Diagnostic
Analyze product health across acquisition, activation, engagement, retention, quality, and monetization.
0
rulebase-co
Cx Customer Health Score
Use to design or audit the support contribution to a customer health score, so the score predicts something instead of averaging weakly-related signals into a colour. Trigger for "build a customer health score", "add support signal to health scoring", "our health scores don't predict churn", red/amber/green account scoring, or a health score nobody trusts.
1
manu14357
AI Analyzer
AI驱动的综合健康分析系统,整合多维度健康数据、识别异常模式、预测健康风险、提供个性化建议。支持智能问答和AI健康报告生成。
16
alterlab-ieu
Alterlab Pytdc
Loads Therapeutics Data Commons (TDC, PyTDC) AI-ready drug-discovery datasets and benchmarks — ADME, toxicity, drug-target interaction (DTI), scaffold splits, and molecular oracles for therapeutic ML and pharmacological prediction. Use when fetching a standardized benchmark dataset, applying scaffold or cold-split evaluation, or sourcing labeled molecules for ADMET, toxicity, or DTI modeling. Sources data, splits, and oracles only — defer molecular featurization (ECFP/fingerprints), model training, and transformers to a molecular-ML skill (e.g. deepchem). Part of the AlterLab Academic Skills suite.
60 · bundle
majiayu000
Epic Consulting
Provides consulting workflows for Epic EMR configuration, including ClinDoc, Bridges interfaces, HL7 messaging, provider matching, and orderset builds.
567 · bundle
ichichuang
Evaluating Llms Harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
0 · bundle
nous-hermeshub
AI Analyzer
AI驱动的综合健康分析系统,整合多维度健康数据、识别异常模式、预测健康风险、提供个性化建议。支持智能问答和AI健康报告生成。
1