
Dario Amodei says AI trust will not…
Anthropic's CEO made a rare X appearance to push back on critics who say his safety warnings backfired. His bet: medicine and biology results will move public opinion more than any PR campaign.
Notes from the trenches on AI engineering, LLM apps, and the full-stack work that holds it all together.

Anthropic's CEO made a rare X appearance to push back on critics who say his safety warnings backfired. His bet: medicine and biology results will move public opinion more than any PR campaign.

Z.ai shipped GLM-5.3 on the same weights as GLM-5.2 and jumped from 4.6 to 28.3 on Terminal-Bench 3.0. The lesson for builders: post-training and harness fit beat another pre-training run.

Nous Research shipped /loop so Hermes re-runs prompts on a timer inside your session, with backoff and real stop conditions. It is cron with memory, and it changes how you monitor long agent jobs.

A founder-facing buyer guide for hiring an AI or automation engineer. Build vs buy vs hire, red flags that separate tool resellers from builders, and the questions that expose a weak proposal before you wire money.

Faraday is a 27B agent trained with long-horizon RL to replicate research figures. Inherent reports it beats Claude Opus 4.8 and GPT-5.5 on its Replica benchmark by directing Codex as a tool.

OpenAI replaced the Chronicle preview with Computer History: an opt-in macOS timeline that feeds ChatGPT and Codex. Useful for picking up work. Risky if you forget consent and prompt injection.

Pika's new Soundtrack, Music, SFX, and Speech models target video pipelines and voice products with aggressive unit economics. The bet: efficient inference beats margin on legacy audio APIs.

Qwen3.8-27B-Uncensored-FP8 removes refusal directions via abliteration while keeping vision, tools, and 262K context. Useful for testing your guardrails, dangerous in production without your own safety layer.