Black Box LLM Granularity Eval

This evaluation probes a black-box LLM's ability to generate high-cardinality, continuous probability scores for binary classification tasks. It measures how well different prompting and post-processing methods improve operational granularity (control over precision-recall operating points) while maintaining predictive performance. Use when the user wants to benchmark on 11 binary classification datasets (combined into a joint dataset for one experiment), or asks about evaluating this task. Reports PRAUC.

qhjqhj00 672515d 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/black-box-llm-granularity-eval commit 672515d3a8

Frequently asked questions

npx skillmds add qhjqhj00/black-box-llm-granularity-eval