Designbench Eval

Evaluates multimodal large language models (MLLMs) on front-end web development tasks, including code generation, editing, and repair across multiple frameworks (React, Vue, Angular, HTML/CSS). It probes capabilities in visual-to-code translation, framework-specific syntax handling, code localization, and component reuse. Use when the user wants to benchmark on DesignBench, or asks about evaluating this task. Reports Compilation Success Rate (CSR).

qhjqhj00 f63d67f 4.4 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/designbench-eval commit f63d67fcc5

Frequently asked questions

npx skillmds add qhjqhj00/designbench-eval