vLLM + Mooncake share KV cache across nodes so agents stop recomputing prefixes

Agent traces reuse huge prefixes turn after turn. Mooncake Store gives vLLM a distributed KV pool: 3.8x throughput, 46x lower TTFT, and near-linear scale on GB200 clusters in vLLM’s report.

SaifullahSaifullah
5 min read
vLLM + Mooncake share KV cache across nodes so agents stop recomputing prefixes

Chatbots tolerate a cold prefill. Agents do not. A coding agent on turn 30 is dragging an 80K-token backpack of system prompt, tools, memory, and prior steps, then adding a few hundred fresh tokens from the latest tool result.

If every vLLM replica keeps its own KV cache, two things break you: local capacity fills up, and the router lands the next turn on a different node that has never seen the prefix. You recompute work you already paid for.

vLLM’s integration with Mooncake Store treats KV as a cluster-wide pool. Shared prefixes become hits across instances instead of lonely GPU memory.

Source: Serving Agentic Workloads at Scale with vLLM x Mooncake.

What agentic traces look like

vLLM analyzed Codex / GPT-style traces on SWE-bench Pro (610 traces, median ~33 turns). By turn 30, context often sits near 80K tokens. Average input:output was about 131:1. Most of each turn is prefix you already computed.

ObservationApprox. figure
Median turns per trace33
Context growth~12K to ~80K
Input:output ratio~131:1
Local-only cache hit rate (baseline story)~1.7%
With Mooncake Store~92% hits on their Codex bench

That is why “just add GPUs” fails without a shared cache story. You scale prefill pain linearly.

vLLM and Mooncake hero diagram for distributed KV cache serving

How the distributed pool fits together

Mooncake already powered prefill-decode disaggregation in vLLM via point-to-point KV transfer. Mooncake Store adds a distributed store:

  • Mooncake master tracks KV-block metadata and client health
  • Clients on GPU nodes manage DRAM / SSD tiers and move blocks over RDMA
  • vLLM’s MooncakeStoreConnector plugs into the existing KVConnector interface
  • Scheduler hashes prompt blocks, asks the store what already exists, and plans around hits
  • Workers register GPU KV for GPUDirect-style transfers where available

Prefill instances write into the pool. Later turns (even on other nodes) can reload matching prefixes instead of recomputing them.

Overall design of vLLM instances sharing a Mooncake Store KV pool

The headline numbers (and how to read them)

On realistic Codex agentic traces (1P1D, 12 GB200 GPUs), vLLM reports:

MetricWith Mooncake Store vs baseline
Throughput3.8x higher
P50 TTFT46x lower
End-to-end latency8.6x lower
Cache hit rate1.7% to 92.2%

Scaling under round-robin routing (the mean case for cross-node misses) stayed near-linear out to 60 GB200 GPUs with >95% hit rate in their plots.

Those numbers are cluster-class. The transferable lesson for smaller teams is still the hit-rate jump: once prefixes are shared, TTFT stops being a tax on every tool turn.

Scaling chart for vLLM with Mooncake Store across GB200 GPU counts
Comparison chart of Mooncake Store versus alternative KV transfer paths

Your mileage depends on RDMA fabric, model size, quantization of KV, and whether your router actually moves sessions. The architectural point still stands: agent serving is a distributed systems problem, not a single-replica demo.

Minimal mental model for builders

You do not need a 60-GPU cluster to use the idea:

  1. Single-node offload: MooncakeStoreConnector with kv_both to spill KV to CPU/SSD and stretch effective cache.
  2. Multi-instance prefix share: same store + fixed PYTHONHASHSEED so block hashes match across processes.
  3. Prefill-decode disagg: combine MooncakeConnector (P2P) with MooncakeStoreConnector (pool) via MultiConnector.

Docs:

MOONCAKE_CONFIG_PATH=mooncake_config.json \ vllm serve meta-llama/Llama-3.1-8B-Instruct \ --kv-transfer-config '{"kv_connector":"MooncakeStoreConnector","kv_role":"kv_both"}'

Prompt engineering becomes systems engineering

Once you have a shared KV pool, small prompt habits compound:

  • Keep the system prompt and tool schemas byte-stable across turns
  • Put volatile tool output and user deltas at the end
  • Avoid rewriting the entire agent “constitution” mid-session
  • Namespace caches with cache_prefix when multiple apps share one Mooncake master

Hash mismatches are silent killers. vLLM’s docs stress a fixed PYTHONHASHSEED across processes that share the store. Miss that and you will stare at a beautiful cluster that never hits cache.

HabitCache effect
Stable system + toolsLong shared prefix hits
Rewriting tools every turnForced misses
Sticky sessions onlyHelps, but fails on failover
Shared store + consistent hashesHits survive routing

When Mooncake is overkill

You might not need this on day one:

  • Single replica, short chats, tiny system prompts
  • Batch offline jobs where TTFT is irrelevant
  • Prototypes on one GPU with local prefix caching already hitting 90%+

You probably do need a plan when:

  • Agents run dozens of turns with 50K-200K contexts
  • You horizontally scale vLLM and the load balancer is not sticky
  • Prefill dominates your GPU bill

Why this belongs in Applied AI shipping

Everyone is building agents. Fewer people are budgeting for prefix reuse. If your agent loops 30 times on a fat system prompt plus tool schemas, inference cost is mostly prefill unless the cache hits.

The product lesson: design prompts and skills for stable, hashable prefixes. Put volatile tool output at the end. Keep system instructions and tool definitions identical across turns so a distributed KV pool can do its job.

I help teams ship agent workflows that survive contact with production latency and cost. If that is your bottleneck, book a free discovery call.

Share this post

Related posts