Yufeng Xguard Eval

Evaluates the safety classification capabilities of guardrail models across multiple dimensions, including prompt/response safety detection, multilingual robustness, adversarial jailbreak resilience, and safe content completion. It also tests the model's ability to dynamically adapt to new moderation policies without retraining. Use when the user wants to benchmark on Aegis / Aegis2.0, WildGuard, StrongReject, SEval2.0, E-commerce Benchmark, Adaptive Policy Scope Benchmark, or asks about evaluating this task. Reports F1 score.

qhjqhj00 6924cf4 3.3 KB Updated 3 repo stars

File contents

qhjqhj00/research-skills-pool/tree/main/skill-factory/output/yufeng-xguard-eval commit 6924cf47af

Frequently asked questions

npx skillmds add qhjqhj00/yufeng-xguard-eval