Trustbench Eval

Evaluates the ability of autonomous LLM agents to perform safe, domain-specific actions in real-time by measuring how effectively a trust verification framework reduces harmful actions while maintaining task completion. It probes epistemic calibration, runtime safety intervention, and domain-specific verification reliability. Use when the user wants to benchmark on MedQA, FinQA, TruthfulQA, or asks about evaluating this task. Reports harmful_actions.

qhjqhj00 ed8cb9e 3.6 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/trustbench-eval commit ed8cb9e618

Frequently asked questions

npx skillmds add qhjqhj00/trustbench-eval