Langfuse Evaluation

Designs and runs LLM evaluation with Langfuse — the strategy and workflow layer for scoring quality, building datasets, and running experiments. Use whenever the user is evaluating LLM output quality with Langfuse: "evaluate my LLM app", "which eval method should I use", "set up LLM-as-a-judge", "create a dataset / run an experiment", "score my traces", "offline vs online evaluation", "test prompt changes before deploying", "build a regression test set", or interpreting experiment results. Owns eval STRATEGY and the datasets/experiments/scores workflow; defers judge calibration and CI/CD experiment code to the vendored `langfuse` skill, and exact SDK code to live docs.

jbaham2 aef49ea 11 files · 42.1 KB Updated

File contents

jbaham2/claude-langfuse-plugin/tree/main/skills/langfuse-evaluation commit aef49ea9c0

Frequently asked questions

npx skillmds@latest add jbaham2/langfuse-evaluation