Care RAG Eval

Evaluates whether LLMs using retrieval-augmented generation can faithfully apply clinical guidelines (specifically Written Exposure Therapy) to answer questions. It probes context fidelity, reasoning complexity, and question type to measure if models actually ground their inferences in retrieved evidence rather than relying on parametric knowledge or guessing. Use when the user wants to benchmark on CARE-RAG WET Guidelines, or asks about evaluating this task. Reports Inference Score.

qhjqhj00 f5d046c 3.5 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/care-rag-eval commit f5d046c58d

Frequently asked questions

npx skillmds add qhjqhj00/care-rag-eval