Webcompass Eval

This benchmark evaluates code language models on multimodal web coding tasks, including generating, editing, and repairing web applications from text, image, or video inputs. It probes functional correctness, UI consistency, and interactive behavior using an automated agent-based pipeline and checklist-guided LLM judges. Use when the user wants to benchmark on WebCompass, or asks about evaluating this task. Reports Overall Score.

qhjqhj00 a446485 4.9 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/webcompass-eval commit a44648551f

Frequently asked questions

npx skillmds add qhjqhj00/webcompass-eval