LLM Agent Benchmarking

当需要测试、基准评测 LLM 智能体(行为测试、能力评估、可靠性指标、回归与生产监控)时使用;做统计化多次运行评估、行为契约/对抗测试、回归门禁与数据泄漏检测,产出含通过率、置信区间、违规项与上线建议的评测报告;不适用于模型训练评估(loss/perplexity)、公平性偏见测试或纯 UX 测试。触发词:智能体评测、benchmark、对抗测试、回归测试、数据泄漏

findscripter c36444a 6.4 KB Updated

File contents

findscripter/everything-skills/tree/main/04-ai/llm-agent-benchmarking commit c36444aaab

Frequently asked questions

npx skillmds@latest add findscripter/llm-agent-benchmarking