
GLM-5.3 got 50% better at coding without…
Z.ai shipped GLM-5.3 on the same weights as GLM-5.2 and jumped from 4.6 to 28.3 on Terminal-Bench 3.0. The lesson for builders: post-training and harness fit beat another pre-training run.

Z.ai shipped GLM-5.3 on the same weights as GLM-5.2 and jumped from 4.6 to 28.3 on Terminal-Bench 3.0. The lesson for builders: post-training and harness fit beat another pre-training run.

Theo Browne's T3 Code GUI added Grok via ACP and X or SuperGrok OAuth. No API key, same subscription you already pay for, alongside Claude Code and Codex in one dashboard.

Tensorlake ran 30 hard agentic tasks with DeepSeek V4 Flash wired through four harnesses. Pi won on pass rate and cost per success. Claude Code was fastest but burned 741k tokens per task.

Z.ai's GLM-5.2 open-weight flagship targets million-token coding trajectories with MIT weights, High and Max reasoning modes, and Terminal-Bench scores that jump from 62.0 to 81.0 versus GLM-5.1.

Developers are shipping browser Minecraft clones with Claude Fable 5 in 20 to 40 minutes for roughly $12 to $30. The interesting part is not the game. It is the systems design the model held in one context.

Cohere's first open-weight coding model activates 3B of 30B parameters per token, ships under Apache 2.0, and targets terminal agents. Here is when I would run it locally instead of a frontier API.

Composer 2.5 is now a third-party model in Grok Build's /model menu. Here's what that means for terminal agents, model routing, and who should switch.

Mellum2 is JetBrains' open 12B MoE with 2.5B active parameters per token, 131K context, and Apache 2.0 weights. Here is when it beats bigger dense models for routing, RAG, and agent sub-calls.

Cognition closed a $1B Series D at a $26B valuation with $492M run-rate revenue. The clearest proof point is internal: 89% of Cognition's committed code now comes from Devin. Here's what that means if you ship software for a living.

Qwen3 8B at Q4_K_M fits in about 5 GB of VRAM and hits roughly 20–50 tok/s on consumer GPUs, including older cards. Here is how to think about local agentic coding without the Mac Mini hype.