Review LLM Annotations And Improve Prompt

Calculate agreement between human ground truth and machine labels for a text LLM judge metric, then analyze transcripts and reviewer notes to propose an improved metric prompt. One metric at a time.

coval-ai Updated

File contents

coval-ai/coval-external-skills/tree/main/skills/human-review/review-llm-annotations-and-improve-prompt commit f04f6c6b2d

Frequently asked questions

npx skillmds@latest add coval-ai/review-llm-annotations-and-improve-prompt