Armor Eval

This benchmark meta-evaluates objective music evaluation (OE) metrics by measuring how well their similarity scores and classification outputs align with human subjective judgments. It probes whether automated algorithms can reliably capture human perception of musical quality and distinguish human-composed from AI-generated music across diverse genres and generative models. Use when the user wants to benchmark on Armor, or asks about evaluating this task. Reports correlation coefficient.

qhjqhj00 4cd81f5 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/armor-eval commit 4cd81f5c69

Frequently asked questions

npx skillmds add qhjqhj00/armor-eval