
9 min read
The three-layer security stack I use when…
Wiped databases, mass-deleted inboxes, leaked tokens. Prompts did not stop any of it. Here is the infrastructure, runtime, and network defense-in-depth stack teams are shipping instead.

Wiped databases, mass-deleted inboxes, leaked tokens. Prompts did not stop any of it. Here is the infrastructure, runtime, and network defense-in-depth stack teams are shipping instead.

GPT-Red is an internal automated red-teaming model trained with self-play RL. It hit 84% attack success versus 13% for human testers, discovered fake chain-of-thought injections, and helped cut GPT-5.6 Sol failures on the hardest benchmark by 6x.

SkillSpector joins a crowded field of agent skill scanners with 70+ vulnerability patterns and optional LLM analysis. New research shows packed and obfuscated skills still bypass most static tools. Scan first, sandbox second.