Human Agent Trust Exploit Detection

Detect social engineering, deceptive responses, false assurances, or prompts that induce unsafe user actions.

Tencent Updated

File contents

Tencent/AI-Infra-Guard/tree/main/agent-scan/agent_scan/prompt/skills/human-agent-trust-exploit-detection commit 142228a393

Frequently asked questions

npx skillmds@latest add tencent-ai-infra-guard/human-agent-trust-exploit-detection