Openassistant Human Eval Eval

Evaluates the quality of aligned language models on multi-turn conversational prompts by measuring how often their generated responses are preferred over supervised fine-tuning targets by human annotators. Use when the user wants to benchmark on OpenAssistant test set, or asks about evaluating this task. Reports winrate.

qhjqhj00 a7debb4 3.1 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/openassistant-human-eval-eval commit a7debb4093

Frequently asked questions

npx skillmds add qhjqhj00/openassistant-human-eval-eval