Packs
4 packs@jiachen-t-wang
Curation Bench
Curation Bench from Jiachen-T-Wang/curation-bench-pro.
99 skills · pack
@gtynnn060110-hash
Environment
Environment from gtynnn060110-hash/continual-skill-bench-final.
7 skills · pack
curated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · pack
@owl-listener
Prototyping Testing
Prototyping and testing skills: wireframe specs, usability heuristics, heuristic evaluations, accessibility audits, A/B test design, and benchmark analysis.
8 skills · pack
Results for “bench”
212 skillsbenchmarking-kubernetes-with-kube-bench
Run CIS Kubernetes Benchmark checks and remediate findings with kube-bench.
24.6k · bundle
performing-kubernetes-cis-benchmark-with-kube-bench
Audit Kubernetes cluster security posture against CIS benchmarks using kube-bench with automated checks for control plane, worker nodes, and RBAC.
24.6k · bundle
bench-automation
Automate Bench operations through Composio's Bench toolkit via Rube MCP, with tool discovery and connection management.
66.9k
benchmark-models
Cross-model benchmark for gstack skills. (gstack)
0
results
Documents benchmark results for a multi-agent system and provides commands to run benchmark suites.
0 · bundle
model-benchmark
Benchmark LLM performance across tasks — latency, quality, cost comparison.
0
More results
jetson-llm-benchmark
Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.
2.2k · bundle
cx-benchmark-methodology
Use to compare CX performance to a published or vendor benchmark without fooling yourself — scope mismatch, survivor bias, and definition mismatch usually make external benchmarks incomparable, and internal baselines often beat them. Trigger for "how do we compare to industry", "is our CSAT good", benchmark slide for the board, vendor benchmark report, "are we above average", outsourcing RFP benchmarks, or when someone cites a round-number industry standard.
1
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.
542
benchling-integration
Integrate with Benchling's Python SDK and REST API to manage registry entities, inventory, ELN entries, workflows, and Data Warehouse queries for life sciences R&D automation.
30.2k · bundle
grit-general-robust-image-task-benchmark-arxiv-2306-14818v2
Grit: General Robust Image Task Benchmark
6
hardening-windows-endpoint-with-cis-benchmark
Hardens Windows endpoints using CIS Benchmark recommendations to reduce attack surface, enforce security baselines, and meet compliance requirements.
24.6k · bundle
finance-metrics-quickref
Look up SaaS finance metrics, formulas, and benchmarks fast. Use when you need a quick metric definition, formula, or benchmark during analysis.
5.6k
benchmark-email-automation
Automate Benchmark Email marketing tasks through Composio's Rube MCP integration, including contact management, campaign creation, and email sending.
66.9k
auditing-cloud-with-cis-benchmarks
Conduct cloud security audits using CIS benchmarks for AWS, Azure, and GCP, including automated assessments, remediation, and continuous compliance monitoring.
24.6k · bundle
performing-docker-bench-security-assessment
Audits Docker host and daemon configuration against the CIS Docker Benchmark, generating compliance reports with pass/fail/warn results and remediation steps.
24.6k · bundle
dataperf-benchmarks-for-data-centric-ai-development-arxiv-22
DataPerf: Benchmarks for Data-Centric AI Development
6
score
Audits medical LLM benchmarks across five lifecycle phases using 46 medically tailored criteria to assess clinical relevance, data integrity, safety-critical capabilities, validity, and governance.
3
nlvr2-a-visual-reasoning-benchmark-for-natural-language-arxi
NLVR2: A Visual Reasoning Benchmark for Natural Language
6
hardening-linux-endpoint-with-cis-benchmark
Hardens Linux endpoints using CIS Benchmark recommendations for Ubuntu, RHEL, and CentOS to reduce attack surface, enforce security baselines, and meet compliance requirements.
24.6k · bundle
bench-read
Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.
0
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
0
agent-evaluation
Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.
2
ab-setup
Setup with Apache Bench. installation.
2 · bundle
dd
Clones disks, benchmarks I/O, and converts files with progress monitoring options.
567 · bundle
performance-engineer
Benchmark, load test, capacity plan, and cache with k6, JMeter, Locust, and pgbench. Use when the user says "slow", "performance", "load test", "stress test", "how many users can it handle", "capacity", "cache", "benchmark", "k6".
2
standup
Facilitates a read-only standup across git worktrees, branches, or PRs to compare changes and produce one consolidation plan.
· bundle
hunch
Discover, bet on, track, and settle Hunch prediction markets in natural language with on-chain settlement on Base via x402.
1.2k · bundle
010-api-0769441a
Reference for Benchling REST API v2 endpoints, covering authentication, pagination, error handling, and CRUD operations for sequences, entities, containers, and notebook entries.
7 · bundle
mmbench-is-your-multi-modal-model-an-all-around-player-arxiv
MMBench: Is Your Multi-modal Model an All-around Player?
6
react18-batching-patterns
Diagnose and fix automatic batching regressions in React 18 class components with a decision tree for setState behavior changes.
36.2k · bundle
create-branch
Create a git branch following Sentry naming conventions, automatically deriving the branch type and description from current work context.
845
ab-ci
CI with Apache Bench. CI integration.
2 · bundle
inngest
Build serverless background jobs, event-driven workflows, and durable execution with Inngest, without managing queues or workers.
42.4k
bmad-ml-research-party
Run multi-agent research discourse session. Use when the user requests to "start a research party" or "run a journal club".
0 · bundle
agent-browser
Control a headless browser to navigate pages, click elements, fill forms, take screenshots, record video, and execute JavaScript using Playwright and inference.sh.
584 · bundle