Model Safety

Review AI features for misuse, prompt injection, unsafe tool use, data exposure, unreliable outputs, oversight, and operational controls.

egeetas 6b8d303 2 files · 1.3 KB Updated

File contents

Model Safety

Apply to AI features that process untrusted content, retrieve private data, call tools, make consequential recommendations, or act on external systems.

Review Areas

  • Threat actors, misuse cases, prohibited outcomes, affected users, and consequence severity.
  • Prompt injection and instruction/data separation across uploads, web content, retrieval, connectors, and tool outputs.
  • Least-privilege tools, allowlists, argument validation, sandboxing, rate limits, approvals, and reversible actions.
  • Data minimization, tenant isolation, retention, logging, redaction, model-provider boundaries, and secret handling.
  • Hallucination, uncertainty, citations, abstention, validation, human review, and appeal or recovery paths.
  • Adversarial evaluations, monitoring, incident response, and staged rollout.

Do not represent prompt wording alone as a security boundary. Recommend layered technical controls and state residual risk.

egeetas/skills/tree/main/skills/model-safety commit 6b8d3037b3

Frequently asked questions

npx skillmds@latest add egeetas/model-safety