Alignment

AI Safety Scores Mostly Measure How Big the Model Is When Aligned Agents Form an Unaligned System Fine-Tuning Unlocks What Alignment Was Hiding, Not What You Taught It