Constitutional AI

Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment, reducing harmful outputs without human labels. Powers Claude's safety system.

synthetic-sciences 600017c 7.2 KB Updated

File contents

synthetic-sciences/openscience/tree/main/backend/cli/skills/llm-tools/constitutional-ai commit 600017c319

Frequently asked questions

npx skillmds@latest add synthetic-sciences/constitutional-ai