Frontalk Eval

Probes a model's ability to generate and iteratively refine front-end code through multi-turn conversational instructions, handling both textual and visual feedback. It specifically measures functional correctness, user experience design quality, and the model's tendency to overwrite prior implementations in long-context interactions. Use when the user wants to benchmark on FronTalk, or asks about evaluating this task. Reports pass rate (PR), usability (UX).

qhjqhj00 438da94 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/frontalk-eval commit 438da9414c

Frequently asked questions

npx skillmds add qhjqhj00/frontalk-eval