Blog 385 posts
Notes from the trenches on AI engineering, LLM apps, and the full-stack work that holds it all together.
Cursor 3.6 shipped auto-review on May 29, 2026 with a classifier subagent, sandbox layer, and roughly 84% fewer approval prompts. Here is how to configure it without treating convenience as a security boundary.
Alibaba's Tongyi Lab shipped Qwen-VLA, a unified policy that handles manipulation, navigation, and trajectory prediction across 11 robot embodiments via prompt conditioning. Here is why that pattern matters beyond robotics.
xAI opened grok-build-0.1 on the public API at $1 input and $2 output per million tokens. Here is how it compares for agent loops, cached context, and when I would route to it.
grok-build-0.1 hit the xAI API on May 29, 2026 at $1 per million input tokens with a 256K context window and native tool use. Here is how it fits next to Composer and frontier tiers in a real harness.
Direct Corpus Interaction lets agents search raw files with rg and grep. GrepSeek trains a 9B model to do it at scale, with a hybrid semantic-plus-terminal stack for production.
Meta alignment lead Summer Yue told her OpenClaw agent to suggest inbox cleanup, not execute it. Context compaction erased that constraint. Here's what operators should copy from the incident.
Mysterium VPN found over 12 million IPs serving public .env files with API keys and DB passwords. Local agents that read plaintext secrets multiply that risk. Here is how I vault credentials for production agents.
Every Claude-built landing page does not have to look like purple-gradient SaaS slop. Anthropic's frontend-design skill forces a token system, aesthetic risk, and an anti-default review pass before any HTML ships.