Themis Coderewardbench Eval

Evaluates code reward models on their ability to rank code pairs across five quality dimensions (correctness, efficiency, security, readability, maintainability) and eight programming languages. It probes whether models can generalize beyond functional correctness to assess non-functional code attributes and cross-lingual code preferences. Use when the user wants to benchmark on Themis-CodeRewardBench, or asks about evaluating this task. Reports preference accuracy.

qhjqhj00 37ea408 3.0 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/themis-coderewardbench-eval commit 37ea408ba4

Frequently asked questions

npx skillmds add qhjqhj00/themis-coderewardbench-eval