Bench Dataset Evaluator

Reference documentation for the BenchDatasetEvaluator operator. Covers the constructor, two comparison modes (match/semantic), and pipeline usage. Use when: comparing predicted answers against ground truth answers in benchmark evaluation.

opendcai 81f6b80 4 files · 17.3 KB Updated

File contents

opendcai/dataflow-webui/tree/main/skills/canonical/core_text/eval/bench-dataset-evaluator commit 81f6b80ba6

Frequently asked questions

npx skillmds@latest add opendcai/bench-dataset-evaluator