Coding Agent Eval Scoring

Use when scoring a coding model or coding agent against a rubric from its real output - PRs, branches, diffs, generated docs - and the score has to survive being challenged. Runs objective gates before any human judgment, requires the evaluator to build and run the code with the exact commands the task instructions specified, and runs the same prompts blind through a control model so a deduction can be attributed to the model rather than to a broken instruction. Triggers on "모델 평가", "성적서", "채점 기준", "vibe eval", "코딩 모델 비교", "에이전트 벤치마크", "PR 기준으로 평가", "대조군", scoring an AI coding tool trial or vendor bake-off.

HojinJava Updated

File contents

HojinJava/claude_skill/tree/main/coding-agent-eval-scoring commit b4246a99f5

Frequently asked questions

npx skillmds@latest add hojinjava/coding-agent-eval-scoring