Interpreting Controlling Behavior Constitutions

Learn and apply natural-language constitutions that map prompt edits to predictable model behavior changes. Use atomic concept edits (ACEs) to systematically probe, interpret, and steer LLM or text-to-image model outputs. Trigger phrases: 'learn a constitution for this model', 'find what prompt changes affect output', 'build an ACE constitution', 'probe model behavior with edits', 'steer model output systematically', 'interpret model sensitivity to prompt changes'.

ndpvt-web cb0fded 14.6 KB Updated

File contents

ndpvt-web/arxiv-claude-skills/tree/main/skills/interpreting-controlling-behavior-constitutions commit cb0fded650

Frequently asked questions

npx skillmds@latest add ndpvt-web/interpreting-controlling-behavior-constitutions