Relaxed Perplexity

This protocol evaluates the reliability, consistency, and inter-correlation of various open-ended and close-ended evaluation metrics on healthcare LLM outputs. It specifically probes how well metrics capture factual coherence and content quality while being robust to output rephrasing and sampling variations. Use when the user has predictions and gold and needs to compute Relaxed Perplexity.

qhjqhj00 fc0d2f8 4.2 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/relaxed-perplexity commit fc0d2f8722

Frequently asked questions

npx skillmds add qhjqhj00/relaxed-perplexity