Conference
Artifacts
  • Available
  • Functional
  • Reproduced

TENURE: Residency Leases for Tiered LLM KV Caches

SoCC 2026 — ACM Symposium on Cloud Computing

Amir Noohi, Antonio Barbalace

PDFSoonSlidesSoonCodeSoonDOISoon
0 citations0 downloads

Abstract

In long-context LLM serving, the key–value (KV) cache outgrows GPU memory, so production systems keep it across DRAM, CXL memory and NVMe disk, and a tiering policy decides which blocks stay in fast memory. Deployed policies predict future reads from past ones, which fits the KV cache poorly. The blocks a node needs change with the router's mix of requests, so such a policy notices each change only after requests have waited on disk, although the engine already knows which blocks the queued requests' prefills will read. We measure 17 state-of-the-art tiering policies inside vLLM on a four-tier host over 10 workloads near capacity. No single policy has the best median time to first token (TTFT) on every workload, and at the same request rate the best and the worst policy differ by 1.12–6.17× in median TTFT within a workload, although waiting on KV loads makes up at most 74% of TTFT.

Tenure takes a counted lease on a request's blocks when the request enters the queue, which keeps them off disk until it finishes while letting them move among the fast tiers. Because KV blocks never change, dropping one to disk moves zero bytes, and a block read from CXL moves up to DRAM only when a queued request will read it, so a promotion never pushes down a block about to be read. Unleased blocks follow the classic order that has cost least on recent reads. Tenure has the best median TTFT on 6 of the 10 workloads, and on its worst it is within 1.30× of the best, while every baseline is at least 1.98× from the best on some workload.

BibTeX

@inproceedings{noohi2026tenure,
  author    = {Amir Noohi and Antonio Barbalace},
  title     = {{TENURE: Residency Leases for Tiered LLM KV Caches}},
  booktitle = {ACM Symposium on Cloud Computing (SoCC)},
  year      = {2026},
}