Routenlp Eval

Evaluates a closed-loop LLM routing system's ability to dynamically select between a four-tier model portfolio based on task difficulty, balancing inference cost, response quality, and latency. The benchmark probes how well a router can escalate queries to more capable models only when necessary, while using distillation and conformal cascading to maintain performance at lower cost tiers. Use when the user wants to benchmark on EDGAR (NER), EDGAR (Summarization), BANKING77* (Intent Classification), BANKING77* (Response Generation), CUAD* (Clause Extraction), CUAD* (Risk Assessment), or asks about evaluating this task. Reports Quality Ratio.

qhjqhj00 faf66c5 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/routenlp-eval commit faf66c5f10

Frequently asked questions

npx skillmds add qhjqhj00/routenlp-eval