constitutional-ai

orchestra-research/constitutional-ai · Agent Skill

by Orchestra Research · single

Published · Last updated


Train AI models to be harmless through self-critique and AI feedback using a set of constitutional principles, without requiring human labels for harmful outputs.

SKILL.md

Related

  1. fine-tuning-with-trl · orchestra-research bundle
    Fine-tune and align language models using reinforcement learning with TRL, including SFT, DPO, PPO, GRPO, and reward model training.
    10.4k
    repo stars
  2. openrlhf-training · orchestra-research bundle
    Train large language models (7B-70B+) with RLHF using PPO, GRPO, DPO, and other algorithms, accelerated by Ray and vLLM for distributed multi-GPU setups.
    10.4k
    repo stars
  3. huggingface-llm-trainer · huggingface bundle
    Train or fine-tune language and vision models using TRL or Unsloth on Hugging Face Jobs cloud infrastructure, with support for SFT, DPO, GRPO, and reward modeling, plus GGUF conversion for local deployment.
    10.8k
    repo stars
  4. trl-training · huggingface
    Train and fine-tune transformer language models using TRL (Transformers Reinforcement Learning) with support for SFT, DPO, GRPO, KTO, RLOO, and reward model training via CLI commands.
    10.8k
    repo stars
  5. gemma-trainer · google-gemma bundle
    Fine-tune Gemma models locally using QLoRA, Unsloth, or TRL for SFT, DPO, and reward modeling, with dataset preparation and conversion to GGUF or LiteRT-LM.
    3
    installs
  6. nv-reason-cxr · nvidia bundle
    Runs chest X-ray reasoning smoke tests using the NV-Reason-CXR-3B model via local inference or a public Hugging Face Space API.
    2.2k
    repo stars

Frequently asked questions

How do I install the constitutional-ai skill?

Run npx skillmds add orchestra-research/constitutional-ai in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the constitutional-ai skill do?

Train AI models to be harmless through self-critique and AI feedback using a set of constitutional principles, without requiring human labels for harmful outputs. It is listed under AI & ML, Coding & Dev Tools, Research & Search, Model Training & Fine-tuning, Prompt Engineering on SkillMD.

Is constitutional-ai safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with constitutional-ai?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is constitutional-ai free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published constitutional-ai?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.