Skills
All skills
Official skills
Leaderboard
Saved
Categories
Coding & Dev Tools
9777
AI & ML
5020
DevOps & Infra
2868
Integrations & APIs
2327
Productivity
2192
Security
1976
All categories →
Plugins
Docs
menu-rounded
Skills
Categories
Plugins
Docs
My skills
Saved
light-dark-mode
Light
Dark
System
Sign in
Profile
My skills
Saved
Collections
Edit profile
Submit a skill
Sign out
Plugins
1 plugin
curated
Run Agent Evaluation
Sets up evaluation framework, runs benchmarks, and produces comparative analysis of agent performance.
9 skills · plugin
Results for “benchmarks”
2 skills
diegosouzapw
LLM
Routes prompts to any LLM model across multiple providers via CLI tools or APIs, with auto-discovery of new models and benchmark data.
54
·
bundle
More results
orchestra-research
Evolving AI Agents
Optimize AI agents through automated evolution cycles using LLM-driven mutation of prompts, skills, and memory against measurable benchmarks.
10.4k
·
bundle