Vicuna Benchmark Eval

This benchmark assesses general language model capabilities and safety across diverse tasks like Fermi problems, roleplay, and coding. It evaluates how well models balance helpfulness, accuracy, and safety on non-safety-specific queries. Use when the user wants to benchmark on Vicuna_Benchmark, or asks about evaluating this task. Reports Net Win Rate.

qhjqhj00 e97f759 2.8 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/vicuna-benchmark-eval commit e97f759def

Frequently asked questions

npx skillmds add qhjqhj00/vicuna-benchmark-eval