Adaptive LLM Testing Eval

This evaluation probes the effectiveness of diversity-based adaptive test selection strategies for black-box LLM applications. It measures how quickly and reliably different prioritization methods detect failures in prompt templates compared to random baselines, while also assessing the diversity of generated outputs. Use when the user wants to benchmark on BBH & P3 Prompt Templates, or asks about evaluating this task. Reports APFD.

qhjqhj00 ab7b72b 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/adaptive-llm-testing-eval commit ab7b72b2e5

Frequently asked questions

npx skillmds add qhjqhj00/adaptive-llm-testing-eval