Plugins
3 pluginscurated
ML Model Lifecycle
Train, evaluate, and deploy a production ML system with monitoring.
10 skills · plugin
@owl-listener
Visual Critique
Visual critique skills: hierarchy analysis, brand consistency checks against mood/voice/tokens, composition evaluation, and typography audits — with a /critique-screen command that compiles a prioritised fix list.
7 skills · plugin
@alirezarezvani
Agenthub
Multi-agent collaboration — spawn N parallel subagents that compete on code optimization, content drafts, research approaches, or any task that benefits from diverse solutions. 7 slash commands (/hub:init, /hub:spawn, /hub:status, /hub:eval, /hub:merge, /hub:board, /hub:run), agent templates, DAG-based orchestration, LLM judge mode, message board coordination.
8 skills · plugin
Results for “l-eval”
463 skillsScikit Learn
Build and evaluate machine learning models using scikit-learn for classification, regression, clustering, dimensionality reduction, and preprocessing.
30.2k · bundle
Polars
Process tabular data with Polars' expression API, lazy evaluation, and parallel execution for fast in-memory analysis and pandas migration.
0 · bundle
Architecture Designer
Design high-level system architecture, create Architecture Decision Records (ADRs), evaluate technology trade-offs, and plan for scalability.
10.4k · bundle
Harness Engineering
Designs autonomous agent harnesses with locked evaluators, editable surfaces, durable logging, novelty gates, pruning, rollback, and human approval boundaries.
16.9k
Phoenix Observability
Self-hosted observability platform for LLM applications, providing tracing, evaluation, datasets, experiments, and real-time monitoring to debug and improve AI systems.
3 · bundle
Arize Dataset
Manage Arize datasets and examples using the ax CLI: create, list, get, export, and append datasets for evaluation and experimentation.
36.2k · bundle
Arbor
Run autonomous optimization loops that iteratively improve artifacts against evaluators using hypothesis tree refinement, without overfitting.
30.2k · bundle
Polars
Process in-memory tabular data with a fast, expression-based DataFrame library that supports lazy evaluation, parallel execution, and Apache Arrow semantics.
3
Sue
Evaluates whether a lawsuit is worth pursuing, explains the litigation process from filing to resolution, and guides case preparation, settlement negotiations, and small claims alternatives.
2
Langfuse
LLM observability with Langfuse — tracing, evals, prompt management, cost tracking
2
Langfuse
Expert in Langfuse - the open-source LLM observability platform. Covers tracing, prompt management, evaluation, datasets, and integration with LangChain, LlamaIndex, and OpenAI. Essential for debugging, monitoring, and improving LLM applications in production. Use when "langfuse, llm observability, llm tracing, prompt management, llm evaluation, monitor llm, debug llm, langfuse, observability, tracing, llm-monitoring, evaluation, prompt-management, debugging, analytics" mentioned.
128 · bundle
Phoenix Observability
Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
1 · bundle
Phoenix Observability
Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
0 · bundle
Ndcg 10
Evaluates how well internal model representations (hidden states) predict token-level information importance in summarization tasks, using NDCG@10 and Spearman's rank correlation.
3
Prompt Engineer
Designs and optimizes prompts for LLM-powered applications, covering system prompt architecture, context management, output formatting, and evaluation.
0
RAG Quality
Evaluate retrieval quality from the local RAG index
1 · bundle
Agent Designer
Design multi-agent system architectures, generate tool schemas for Anthropic and OpenAI formats, and evaluate execution logs for cost, latency, and failure bottlenecks.
20.4k · bundle
Technical Job Search
Helps software engineers with discrete job search tasks: job description analysis, CV tailoring, cover letter writing, offer evaluation, and follow-up emails.
36.2k
Polars
Process data with high-performance DataFrames using Polars' expression-based API, lazy evaluation, and parallel execution for ETL, analytics, and pandas migration.
30.2k · bundle
Ttsds
Evaluates text-to-speech systems by measuring distributional distance between synthetic and real speech across five factors, producing a scalar score without subjective MOS ratings.
3
Polars
Process in-memory datasets with Polars' expression API, lazy evaluation, and parallel execution, including pandas migration patterns and I/O for CSV, Parquet, and JSON.
5
Cab Eval
Benchmarks LLM bias by scoring responses to automatically generated open-ended questions across sensitive attributes, producing a composite fitness score from 0 to 5.
3
Setup
Set up a new autoresearch experiment interactively. Collects domain, target file, eval command, metric, direction, and evaluator. Use when the user runs /ar:setup or asks to start optimizing a file with the autoresearch loop.
11
Retro Learn
Convert delivery findings into skill, eval, workflow, and documentation improvements.
542
MCP Builder
Guides the creation of high-quality MCP servers that let LLMs interact with external services through well-designed tools, covering planning, implementation, testing, and evaluation.
253 · bundle
Arbor
Runs an autonomous optimization loop that iteratively improves an artifact against an objective and evaluator using Hypothesis Tree Refinement, with subagent executors in isolated git worktrees.
253 · bundle
Polars
Provides a fast in-memory DataFrame library for datasets that fit in RAM, with lazy evaluation, parallel execution, and an Apache Arrow backend for ETL pipelines and analytics.
42.4k
Abc Eval
Benchmarks large language models on symbolic music understanding and instruction following using text-based ABC notation, covering syntax parsing, error detection, segment-level reasoning, and sequence-level musical analysis.
3
Digital Health Clinical Asr Eval
Score a clinical ASR manifest against a chosen NIM, produce a five-section KER leaderboard, and route the user via a post-eval decision tree.
2.2k · bundle
Workload Manager Basics
Validate enterprise workloads against Google Cloud best practices using public client libraries and the REST API to manage evaluations, rules, scanned resources, and validation results.
14.4k · bundle
RAG Architect
Designs and implements production-grade RAG systems by chunking documents, generating embeddings, configuring vector stores, building hybrid search pipelines, applying reranking, and evaluating retrieval quality.
10.4k · bundle
Acquisition Channel Advisor
Evaluate acquisition channels using unit economics, customer quality, and scalability to decide whether to scale, test, or kill a growth channel.
5.6k · bundle
Visor
Evaluates text-to-image models on spatial relationship accuracy using the VISOR metric, separating object detection from spatial correctness to reveal biases like object priority and merging.
3
Developer Eval Driven Development
Build and improve AI or probabilistic software through evaluation-driven development. Use for LLM applications, agents, prompts, RAG, tool use, classifiers, model migrations, quality regressions, golden datasets, LLM-as-judge rubrics, benchmarks, or requests to add evals and measurable release gates. Pair with TDD for deterministic code; do not use as the primary guide for ordinary unit testing without model behavior.
1 · bundle
Auc
Evaluates machine learning classifiers on their ability to distinguish signal from background in particle physics simulations, measuring how well algorithms rank signal events above background ones using the AUC metric.
3
Bleurt
Evaluates the correlation between automatic text generation scores and human quality ratings, including robustness to domain and quality drift, using metrics like Kendall's Tau and Pearson correlation.
3