All posts

Inference 8 posts

Every post filed under Inference, newest first.

vLLM + Mooncake share KV cache across nodes so agents stop recomputing prefixes
5 min read

vLLM + Mooncake share KV cache across…

Agent traces reuse huge prefixes turn after turn. Mooncake Store gives vLLM a distributed KV pool: 3.8x throughput, 46x lower TTFT, and near-linear scale on GB200 clusters in vLLM’s report.

NVIDIA's TwoTower model writes text in parallel and keeps 98.7% of AR quality
3 min read

NVIDIA's TwoTower model writes text in parallel…

Nemotron-Labs-TwoTower splits a 30B Nemotron backbone into a frozen context tower and a trainable diffusion denoiser, hitting 2.42x generation throughput with open weights on Hugging Face.