Benchmark E2e

Benchmarks end-to-end ML execution quality across multiple modes (no-plugin/manual, plugin-driven, and AutoGluon-backed AutoML). Automatically identifies exactly one dataset scenario (hard-fraud, hard-attrition, or xhard-churn) and runs the benchmark against that single scenario. Use when asked to compare E2E workflows, measure agent reliability/cost/speed, or recommend which skills should be used for the detected scenario.

lawwu 6b3bd64 8 files · 45.9 KB Updated

File contents

lawwu/agentic-ml-plugin/tree/main/plugins/agentic-ml/skills/benchmark-e2e commit 6b3bd64566

Frequently asked questions

npx skillmds@latest add lawwu/benchmark-e2e