Chatbots tolerate a cold prefill. Agents do not. A coding agent on turn 30 is dragging an 80K-token backpack of system prompt, tools, memory, and prior steps, then adding a few hundred fresh tokens from the latest tool result.
If every vLLM replica keeps its own KV cache, two things break you: local capacity fills up, and the router lands the next turn on a different node that has never seen the prefix. You recompute work you already paid for.
vLLM’s integration with Mooncake Store treats KV as a cluster-wide pool. Shared prefixes become hits across instances instead of lonely GPU memory.
Source: Serving Agentic Workloads at Scale with vLLM x Mooncake.
What agentic traces look like
vLLM analyzed Codex / GPT-style traces on SWE-bench Pro (610 traces, median ~33 turns). By turn 30, context often sits near 80K tokens. Average input:output was about 131:1. Most of each turn is prefix you already computed.
| Observation | Approx. figure |
|---|---|
| Median turns per trace | 33 |
| Context growth | ~12K to ~80K |
| Input:output ratio | ~131:1 |
| Local-only cache hit rate (baseline story) | ~1.7% |
| With Mooncake Store | ~92% hits on their Codex bench |
That is why “just add GPUs” fails without a shared cache story. You scale prefill pain linearly.

How the distributed pool fits together
Mooncake already powered prefill-decode disaggregation in vLLM via point-to-point KV transfer. Mooncake Store adds a distributed store:
- Mooncake master tracks KV-block metadata and client health
- Clients on GPU nodes manage DRAM / SSD tiers and move blocks over RDMA
- vLLM’s
MooncakeStoreConnectorplugs into the existing KVConnector interface - Scheduler hashes prompt blocks, asks the store what already exists, and plans around hits
- Workers register GPU KV for GPUDirect-style transfers where available
Prefill instances write into the pool. Later turns (even on other nodes) can reload matching prefixes instead of recomputing them.

The headline numbers (and how to read them)
On realistic Codex agentic traces (1P1D, 12 GB200 GPUs), vLLM reports:
| Metric | With Mooncake Store vs baseline |
|---|---|
| Throughput | 3.8x higher |
| P50 TTFT | 46x lower |
| End-to-end latency | 8.6x lower |
| Cache hit rate | 1.7% to 92.2% |
Scaling under round-robin routing (the mean case for cross-node misses) stayed near-linear out to 60 GB200 GPUs with >95% hit rate in their plots.
Those numbers are cluster-class. The transferable lesson for smaller teams is still the hit-rate jump: once prefixes are shared, TTFT stops being a tax on every tool turn.


Your mileage depends on RDMA fabric, model size, quantization of KV, and whether your router actually moves sessions. The architectural point still stands: agent serving is a distributed systems problem, not a single-replica demo.
Minimal mental model for builders
You do not need a 60-GPU cluster to use the idea:
- Single-node offload:
MooncakeStoreConnectorwithkv_bothto spill KV to CPU/SSD and stretch effective cache. - Multi-instance prefix share: same store + fixed
PYTHONHASHSEEDso block hashes match across processes. - Prefill-decode disagg: combine
MooncakeConnector(P2P) withMooncakeStoreConnector(pool) viaMultiConnector.
Docs:
MOONCAKE_CONFIG_PATH=mooncake_config.json \ vllm serve meta-llama/Llama-3.1-8B-Instruct \ --kv-transfer-config '{"kv_connector":"MooncakeStoreConnector","kv_role":"kv_both"}'
Prompt engineering becomes systems engineering
Once you have a shared KV pool, small prompt habits compound:
- Keep the system prompt and tool schemas byte-stable across turns
- Put volatile tool output and user deltas at the end
- Avoid rewriting the entire agent “constitution” mid-session
- Namespace caches with
cache_prefixwhen multiple apps share one Mooncake master
Hash mismatches are silent killers. vLLM’s docs stress a fixed PYTHONHASHSEED across processes that share the store. Miss that and you will stare at a beautiful cluster that never hits cache.
| Habit | Cache effect |
|---|---|
| Stable system + tools | Long shared prefix hits |
| Rewriting tools every turn | Forced misses |
| Sticky sessions only | Helps, but fails on failover |
| Shared store + consistent hashes | Hits survive routing |
When Mooncake is overkill
You might not need this on day one:
- Single replica, short chats, tiny system prompts
- Batch offline jobs where TTFT is irrelevant
- Prototypes on one GPU with local prefix caching already hitting 90%+
You probably do need a plan when:
- Agents run dozens of turns with 50K-200K contexts
- You horizontally scale vLLM and the load balancer is not sticky
- Prefill dominates your GPU bill
Why this belongs in Applied AI shipping
Everyone is building agents. Fewer people are budgeting for prefix reuse. If your agent loops 30 times on a fat system prompt plus tool schemas, inference cost is mostly prefill unless the cache hits.
The product lesson: design prompts and skills for stable, hashable prefixes. Put volatile tool output at the end. Keep system instructions and tool definitions identical across turns so a distributed KV pool can do its job.
I help teams ship agent workflows that survive contact with production latency and cost. If that is your bottleneck, book a free discovery call.
