Results for “micro-benchmark”

51 skills
More results
qhjqhj00
Mdad
Quantifies the minimum accuracy gap needed between two models for a sampled micro-benchmark to reliably preserve their ranking, using the MDAD metric from Yauney et al. (2025).
3
anantha-236
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
1
affaan-m
Benchmark
Measure performance baselines, detect regressions before and after PRs, and compare stack alternatives using browser, API, and build benchmarks.
226k
mhassan0000
Benchmark
Measures performance baselines, detects regressions before and after PRs, and compares stack alternatives.
1
yanacuti1121
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
2
rajanthar
Benchmark
Use this skill to measure performance baselines, detect regressions before/after PRs, and compare stack alternatives.
0
affaan-m
Benchmark Methodology
Scores competitors across nine weighted dimensions with explicit 1–5 rubrics and a tension plot, producing comparable profile cards for competitive analysis.
226k
kk20300113-png
Benchmark
Performance regression detection using the browse daemon. Establishes baselines for page load times, Core Web Vitals, and resource sizes. Compares before/after on every PR. Tracks performance trends over time. Use when: "performance", "benchmark", "page speed", "lighthouse", "web vitals", "bundle size", "load time". (gstack) Voice triggers (speech-to-text aliases): "speed test", "check performance".
0
construct-ai-primary
Performance Benchmarking
Use when evaluating, measuring, or comparing the performance of systems, functions, or services. This skill provides a framework for establishing baselines, measuring performance, and validating that changes meet performance requirements.
0
tianhao909
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
1 · bundle
livelybug
Benchmark
Performance regression detection using the browse daemon. (gstack)
0
qcmuu
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
0 · bundle
pwdev-solucoes
Performance Engineer
Benchmark, load test, capacity plan, and cache with k6, JMeter, Locust, and pgbench. Use when the user says "slow", "performance", "load test", "stress test", "how many users can it handle", "capacity", "cache", "benchmark", "k6".
2
inference-sh
Competitor Teardown
Run structured competitive analysis with feature matrices, SWOT, pricing comparison, review mining, and positioning maps using the inference.sh CLI.
584
lionelndong
Quality Check
Benchmark-relative quality gate. Scores the draft against the research dossier's beat spec (depth, consensus coverage, evidence) plus AI-tell and voice signals, runs an adversarial read armed with the SERP benchmark, and emits the verdict that gates the pipeline.
0 · bundle
fivebucksventures
Content Performance Analyst
Analyze organic content performance for any active brand — your own published posts (engagement by topic, format, persona, angle, hook archetype, Direction) plus competitor content benchmarking — and produce a Performance Brief that feeds the social calendar.
0
aniruddhaadak80
Model Benchmark
Benchmark LLM performance across tasks — latency, quality, cost comparison.
0
mhassan0000
Repo Scan
Audits source code across C++, Android, iOS, and Web to classify files, detect embedded third-party libraries, and produce four-level verdicts with interactive HTML reports.
1
intelli-verse-x
Ivx Cf Benchmarking
Design and execute performance benchmarks for models and systems. Use when measuring throughput, latency, or comparing system variants.
0 · bundle
vimalinx
Hmmsim
Use when you need to characterize score distributions of a profile HMM on random sequences, such as calibration checks, benchmarking, or filter-behavior experiments.
0 · bundle
ekatasingh1107
Social Listener
Monitor social platforms and forums for brand mentions, category discussions, and competitor activity
2 · bundle
dontbesilent2025
Dbs Benchmark
Helps find and analyze competitors to imitate using a five-filter method, focusing on profitability and feasibility while eliminating personal bias.
brycewang-stanford
Macro Briefing
Macroeconomic monitor. Supports UK, US, Euro area, and Australia. Pulls GDP, inflation, employment, wages, rates, trade, housing, and fiscal data. Each country follows its central bank's reporting structure. Produces a single clean briefing with a traffic-light assessment and user-selectable sub-components.
1k · bundle
trailofbits
Trailmark Summary
Runs a Trailmark summary analysis on a codebase to auto-detect languages, count entry points, and list dependencies.
6k · bundle
browserbase
Competitor Analysis
Discovers competitors via search API, deeply researches each with a multi-lane pattern, and compiles an HTML report with overview, per-competitor deep dives, feature/pricing matrix, and mentions feed.
3.6k · bundle
rulebase-co
Cx Benchmark Methodology
Use to compare CX performance to a published or vendor benchmark without fooling yourself — scope mismatch, survivor bias, and definition mismatch usually make external benchmarks incomparable, and internal baselines often beat them. Trigger for "how do we compare to industry", "is our CSAT good", benchmark slide for the board, vendor benchmark report, "are we above average", outsourcing RFP benchmarks, or when someone cites a round-number industry standard.
1
mukul975
Benchmarking Kubernetes With Kube Bench
Run CIS Kubernetes Benchmark checks and remediate findings with kube-bench.
24.6k · bundle
qhjqhj00
Cost
Evaluates a containerized framework for deploying distributed big data workloads, measuring execution time and cloud cost scaling from four to eight nodes.
3
fukukei23
Sentaku
選択肢(A/B/C)の深掘り比較→淘汰→推奨で判断負担を下げ判断の質を上げるスキル。5段階(L1固定3点/L1.5案拡張Diverge・自動/L2評価軸マトリクス/L3複数LLM弁証論/L4過去判断照合)。 「比較して」「深掘りして」「メリデメ教えて」「お勧めは?」「徹底的に」「過去の判断と照合」「前にどう決めたっけ」「/sentaku」等で発火。teian(浅)の深掘り要求を受け取り、brainstorming(深:設計全体)と棲み分け。
0
livelybug
Benchmark Models
Cross-model benchmark for gstack skills. (gstack)
0
lionelndong
Content Gap Analysis
Layer 1b of the keyword research pipeline. Finds keyword opportunities by comparing the brand's blog against competitors AND by expanding seeds + modifiers via Semrush (phrase_fullsearch / phrase_related). Auto-discovers competitors via domain_organic_organic when none are provided, derives the keyword gap via domain_domains, tags every row with `gap_mode`, and outputs a candidate-keyword CSV ready for downstream BID/AIO vetting.
0
orchestra-research
Nemo Evaluator Sdk
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution on local Docker, Slurm HPC, or cloud platforms.
10.4k · bundle
lionelndong
Draft Score
Lightweight ContentShake AI self-check the /draft stage can call before saving. Returns just SEO + Quality scores (no full optimization) so the writer knows whether the draft is in winning territory before /quality-check runs. Fails soft when SEMRUSH_API_KEY is unset.
0
rulebase-co
Cx Outsourcer Scorecard
Use to compare BPO sites, vendors or partner teams fairly, adjusting for the work mix each is given before concluding anything about performance. Trigger for "compare our BPO sites", "which vendor is performing best", "site A scores lower than site B", outsourcer QBR packs, partner MI reporting, or setting contractual quality targets with a vendor.
1
qhjqhj00
Posh
Evaluates automated metrics and vision-language models on identifying granular errors in detailed image descriptions and ranking paired descriptions against human judgments, using macro F1, pairwise accuracy, Spearman rank ρ, and Kendall's τ.
3