Packs

4 packs

Results for “bench”

212 skills
More results
nvidia
jetson-llm-benchmark
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
2.2k · bundle
rulebase-co
cx-benchmark-methodology
Use to compare CX performance to a published or vendor benchmark without fooling yourself — scope mismatch, survivor bias, and definition mismatch usually make external benchmarks incomparable, and internal baselines often beat them. Trigger for "how do we compare to industry", "is our CSAT good", benchmark slide for the board, vendor benchmark report, "are we above average", outsourcing RFP benchmarks, or when someone cites a round-number industry standard.
1
hoangnguyen0403
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.
542
k-dense-ai
benchling-integration
Integrate with Benchling's Python SDK and REST API to manage registry entities, inventory, ELN entries, workflows, and Data Warehouse queries for life sciences R&D automation.
30.2k · bundle
jiachen-t-wang
grit-general-robust-image-task-benchmark-arxiv-2306-14818v2
Grit: General Robust Image Task Benchmark
6
mukul975
hardening-windows-endpoint-with-cis-benchmark
Hardens Windows endpoints using CIS Benchmark recommendations to reduce attack surface, enforce security baselines, and meet compliance requirements.
24.6k · bundle
deanpeters
finance-metrics-quickref
Look up SaaS finance metrics, formulas, and benchmarks fast. Use when you need a quick metric definition, formula, or benchmark during analysis.
5.6k
composiohq
benchmark-email-automation
Automate Benchmark Email marketing tasks through Composio's Rube MCP integration, including contact management, campaign creation, and email sending.
66.9k
mukul975
auditing-cloud-with-cis-benchmarks
Conduct cloud security audits using CIS benchmarks for AWS, Azure, and GCP, including automated assessments, remediation, and continuous compliance monitoring.
24.6k · bundle
mukul975
performing-docker-bench-security-assessment
Audits Docker host and daemon configuration against the CIS Docker Benchmark, generating compliance reports with pass/fail/warn results and remediation steps.
24.6k · bundle
jiachen-t-wang
dataperf-benchmarks-for-data-centric-ai-development-arxiv-22
DataPerf: Benchmarks for Data-Centric AI Development
6
qhjqhj00
score
Audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
3
jiachen-t-wang
nlvr2-a-visual-reasoning-benchmark-for-natural-language-arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
mukul975
hardening-linux-endpoint-with-cis-benchmark
Hardens Linux endpoints using CIS Benchmark recommendations for Ubuntu, RHEL, and CentOS to reduce attack surface, enforce security baselines, and meet compliance requirements.
24.6k · bundle
eryajf
bench-read
Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.
0
whd4
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
0
danstrem2
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
2
ssrjkk
ab-setup
Setup with Apache Bench. installation.
2 · bundle
majiayu000
dd
Clones disks, benchmarks I/O, and converts files with progress monitoring options.
567 · bundle
pwdev-solucoes
performance-engineer
Benchmark, load test, capacity plan, and cache with k6, JMeter, Locust, and pgbench. Use when the user says "slow", "performance", "load test", "stress test", "how many users can it handle", "capacity", "cache", "benchmark", "k6".
2
thedotmack
standup
Facilitates a read-only standup across git worktrees, branches, or PRs to compare changes and produce one consolidation plan.
· bundle
bankrbot
hunch
Discover, bet on, track, and settle Hunch prediction markets in natural language with on-chain settlement on Base via x402.
1.2k · bundle
tools-only
010-api-0769441a
Reference for Benchling REST API v2 endpoints, covering authentication, pagination, error handling, and CRUD operations for sequences, entities, containers, and notebook entries.
7 · bundle
jiachen-t-wang
mmbench-is-your-multi-modal-model-an-all-around-player-arxiv
MMBench: Is Your Multi-modal Model an All-around Player?
6
github
react18-batching-patterns
Diagnose and fix automatic batching regressions in React 18 class components with a decision tree for setState behavior changes.
36.2k · bundle
getsentry
create-branch
Create a git branch following Sentry naming conventions, automatically deriving the branch type and description from current work context.
845
ssrjkk
ab-ci
CI with Apache Bench. CI integration.
2 · bundle
antigravity
inngest
Build serverless background jobs, event-driven workflows, and durable execution with Inngest, without managing queues or workers.
42.4k
sirnosh
bmad-ml-research-party
Run multi-agent research discourse session. Use when the user requests to "start a research party" or "run a journal club".
0 · bundle
inference-sh
agent-browser
Control a headless browser to navigate pages, click elements, fill forms, take screenshots, record video, and execute JavaScript using Playwright and inference.sh.
584 · bundle