
Cerebras CS-4: 750 PFLOPS, 4,465 t/s on…
Cerebras unveiled CS-4 on August 18: three WSE-3 Turbo wafers, 750 PFLOPS, and up to 30x GPU inference speed. For agent workflows, the headline is wall-clock time per loop, not another leaderboard point.

Cerebras unveiled CS-4 on August 18: three WSE-3 Turbo wafers, 750 PFLOPS, and up to 30x GPU inference speed. For agent workflows, the headline is wall-clock time per loop, not another leaderboard point.

Agent traces reuse huge prefixes turn after turn. Mooncake Store gives vLLM a distributed KV pool: 3.8x throughput, 46x lower TTFT, and near-linear scale on GB200 clusters in vLLM’s report.

Moonshot's open Kimi K3 is a 2.8T MoE with 1M context. LatentMoE, hybrid KDA attention, and native MXFP4 QAT are what make the file size survivable, not just the benchmark scores.

Unsloth shipped dynamic NVFP4 Qwen3.6 checkpoints on July 10 that beat NVIDIA's own NVFP4 on throughput while holding benchmark accuracy. Here is the backend choice that makes or breaks the speedup.

Nemotron-Labs-TwoTower splits a 30B Nemotron backbone into a frozen context tower and a trainable diffusion denoiser, hitting 2.42x generation throughput with open weights on Hugging Face.

DeepSeek's DSpark speculative decoding framework adds a semi-autoregressive drafter and confidence-scheduled verification to V4 serving. Per-user generation runs 60-85% faster at matched throughput, lossless and open source.

Jalapeño is OpenAI's first custom Intelligence Processor, co-built with Broadcom in nine months for LLM inference. Early benchmarks show 1.5x to 3.6x better latency per watt than today's GPU racks, with volume deployment starting late 2026.

LMCache is an open Apache-2.0 KV cache layer for vLLM and SGLang that offloads and reuses prefixes across queries and engines. Reports cite up to 15x throughput and 3–10x TTFT wins on agentic workloads.