Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
addClaim

DKV: Anchor + Low-Rank Differential KV-Cache Compression for Scalable Long-Context Inference

Authors: Chimurkar, Om;

DKV: Anchor + Low-Rank Differential KV-Cache Compression for Scalable Long-Context Inference

Abstract

The key-value (KV) cache is the memory wall of long-context transformer inference: it growslinearly with sequence length, and on commodity hardware it, not the model weights, is whatfirst exhausts memory. DKV is a KV-cache compression runtime that keeps recent tokens exact and compresses older ones.Each 256-token block is reduced to an anchor token (kept exact), a rank-32 joint K|V truncated-SVDdelta, and a budget of exact residual tokens - the rows the low-rank basis reconstructs worst,which are precisely the distinctive tokens (digits, names, codes) that verbatim recall depends on.At decode time a fused buffer routes to the top-K relevant blocks, scores the query in low-rankspace without ever decompressing K, attends residuals and a dense recency window exactly, andmerges the halves with a flash-style log-sum-exp reduction. Prefill uses training-freeblock-sparse attention, so its cost grows sub-quadratically. Implemented entirely in MLX and measured on Qwen2.5-1.5B (int4) on an Apple M3 with 8.6 GB ofunified memory, against two dense full-KV baselines: a memory-optimized mlx_lm engine sharingDKV's exact int4 weights (the controlled comparison), and a standard PyTorch engine on theunquantized fp16 checkpoint (a practitioner-default reference, not a weight-matched control). DKV holds a bounded KV state 1.44x-2.25x smaller per block than a dense cache, recovers a buriedpasscode exactly at every context from 4k to 64k, and reaches 64k - where the default PyTorchfull-KV configuration runs out of memory at 16k, and which even the memory-optimized dense cachereaches only by prefilling 1.72x slower. The cost is per-token decode throughput wherever a densecache still fits; this trade-off is reported in full rather than hidden. A CUDA/Triton engine isimplemented and CPU-verified, but no GPU numbers are collected and no GPU performance claims aremade. All numbers are measured on the described host; none are estimated.

Keywords

Apple Silicon, low-rank compression, MLX, transformer inference, Efficient Inference, CUDA, Memory Compression, KV cache, long-context inference, sparse attention, LLM systems

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!