Constitutional AI

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system. Use when this capability is needed.

tomevault-io 26a6044 2 files · 8.7 KB Updated

File contents

tomevault-io/skills-registry/tree/main/davila7--claude-code-templates--safety-alignment-constitutional-ai commit 26a6044782

Frequently asked questions

npx skillmds@latest add tomevault-io/constitutional-ai