Anthropic's method for training harmless AI through self-improvement. Two-phase approach - supervised learning with self-critique/revision, then RLAIF (RL from AI Feedback). Use for safety alignment,
airesearch_skills/07-safety-alignment/constitutional-ai/SKILL.md(main)