Teleqna Eval

Evaluates large language models' domain-specific knowledge in telecommunications, covering general terminology, research concepts, and complex technical standards. It also benchmarks model performance against human telecom professionals under strict no-search conditions. Use when the user wants to benchmark on TeleQnA, or asks about evaluating this task. Reports accuracy.

qhjqhj00 f5eb978 2.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/teleqna-eval commit f5eb978ddc

Frequently asked questions

npx skillmds add qhjqhj00/teleqna-eval