Ai-Safety

AI Safety Scores Mostly Measure How Big the Model Is Your AI's Safety Filters Speak English. Biology Isn't a Language They Understand. When Aligned Agents Form an Unaligned System