
The key-value (KV) cache is the memory wall of long-context transformer inference: it growslinearly with sequence length, and on commodity hardware it, not the model weights, is whatfirst exhausts memory. DKV is a KV-cache compression runtime that keeps recent tokens exact and compresses older ones.Each 256-token block is reduced to an anchor token (kept exact), a rank-32 joint K|V truncated-SVDdelta, and a budget of exact residual tokens - the rows the low-rank basis reconstructs worst,which are precisely the distinctive tokens (digits, names, codes) that verbatim recall depends on.At decode time a fused buffer routes to the top-K relevant blocks, scores the query in low-rankspace without ever decompressing K, attends residuals and a dense recency window exactly, andmerges the halves with a flash-style log-sum-exp reduction. Prefill uses training-freeblock-sparse attention, so its cost grows sub-quadratically. Implemented entirely in MLX and measured on Qwen2.5-1.5B (int4) on an Apple M3 with 8.6 GB ofunified memory, against two dense full-KV baselines: a memory-optimized mlx_lm engine sharingDKV's exact int4 weights (the controlled comparison), and a standard PyTorch engine on theunquantized fp16 checkpoint (a practitioner-default reference, not a weight-matched control). DKV holds a bounded KV state 1.44x-2.25x smaller per block than a dense cache, recovers a buriedpasscode exactly at every context from 4k to 64k, and reaches 64k - where the default PyTorchfull-KV configuration runs out of memory at 16k, and which even the memory-optimized dense cachereaches only by prefilling 1.72x slower. The cost is per-token decode throughput wherever a densecache still fits; this trade-off is reported in full rather than hidden. A CUDA/Triton engine isimplemented and CPU-verified, but no GPU numbers are collected and no GPU performance claims aremade. All numbers are measured on the described host; none are estimated.
Apple Silicon, low-rank compression, MLX, transformer inference, Efficient Inference, CUDA, Memory Compression, KV cache, long-context inference, sparse attention, LLM systems
Apple Silicon, low-rank compression, MLX, transformer inference, Efficient Inference, CUDA, Memory Compression, KV cache, long-context inference, sparse attention, LLM systems
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
