Quadsentinel Safety Eval

This evaluation probes the ability of multi-agent guardrail systems to enforce machine-checkable safety policies over agent trajectories in real-time. It measures how effectively a system detects and blocks unsafe actions while minimizing false positives on enterprise web-agent and malicious behavior benchmarks. Use when the user wants to benchmark on ST-WebAgentBench, AgentHarm, or asks about evaluating this task. Reports Accuracy.

qhjqhj00 e51ea30 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/quadsentinel-safety-eval commit e51ea3067b

Frequently asked questions

npx skillmds add qhjqhj00/quadsentinel-safety-eval