Proofrag

Evaluate a RAG or LLM app. Use when the user wants to test, score, benchmark, or catch regressions in a retrieval/RAG/LLM system, generate an evaluation/golden dataset from their docs, measure hallucination/groundedness/correctness, or gate CI on answer quality. Generates a golden set from the user's own corpus, runs LLM-as-judge plus retrieval metrics, and produces a shareable HTML scorecard.

unshDee 885cc77 7.4 KB Updated

File contents

unshDee/proofrag/tree/main/skills/proofrag commit 885cc776d6

Frequently asked questions

npx skillmds@latest add unshdee/proofrag