Ai-Safety
AI Safety Scores Mostly Measure How Big the Model Is
Your AI's Safety Filters Speak English. Biology Isn't a Language They Understand.
When Aligned Agents Form an Unaligned System