All posts

Safety 6 posts

Every post filed under Safety, newest first.

Claude Code auto mode is the default now. Humans caught 14% of dangerous commands.
4 min read

Claude Code auto mode is the default…

Anthropic made auto mode the default on Pro, Max, and Team plans after a 1,053-tester study. The classifier blocked 89% of dangerous shell commands while manual approval fatigue dropped human catches to 5%.

All five major LLMs show pro-female hiring bias on Japanese resumes
4 min read

All five major LLMs show pro-female hiring…

A 43,200-call study on rirekisho-format resumes finds significant pro-female bias across Claude, GPT-4o, DeepSeek, Gemini, and Llama. Prompt fixes failed. Name removal helped but broke GPT-4o safety filters 42% of the time.