Low-Rank Key Value Attention: Reducing KV Cache Memory and Maintaining Head Diversity

In autoregressive decoding, each token requires repeatedly reading the KV cache from memory, and this cost scales linearly with sequence length, layers, and head count. This post introduces Low-Rank Key-Value (LRKV) attention, a drop-in modification to multi-head attention that reduces KV cache size by 45–53% vs standard MHA, while achieving lower test loss across model scales (128M → 6.3B), faster convergence in training steps, and stronger downstream performance after supervised midtraining.

James O'Neill9 April 2026
Low-Rank Key Value Attention: Reducing KV Cache Memory and Maintaining Head Diversity
Comparison of attention mechanisms. Standard MHA uses HHH independent K/V projections per head (high KV cache). MQA/GQA share K/V (low cache, reduced head-specific detail). LRKV combines a shared full-rank projection with head-specific low-rank residuals, achieving cache cost 2L(dh+Hr) while preserving head diversity.
Cross-scale pretraining curves (128M, 1.2B, 2.5B, 6.3B). Test cross-entropy loss vs training tokens (and compute). LRKV is consistently competitive, achieving the lowest test loss at multiple scales.
LRKV achieves superior training efficiency alongside best performance (2.5B scale).Memory vs Performance (left): Test BPB versus KV cache percentage for all methods. LRKV achieves optimal trade-off with lowest BPB at 48.4% cache usage (2.5B scale). Training Efficiency Advantage (right): LRKV reaches each baseline’s final test loss, quantifying training compute savings. LRKV reaches all baselines’ performance earlier.
LRKV rank ablations: final test loss vs cache size, and training curves by rank. LRKV dominates the memory/performance tradeoff space across ranks.
512M model trained at 8k context. LRKV widens the advantage over baselines in long-context settings, where KV cache pressure is highest.
Pairwise gauge-invariant head similarity. LRKV’s structure is nearly indistinguishable from MHA (off-diagonal similarity remains low), consistent with preserved specialization.
LRKV preserves head diversity across scales. Gauge-invariant effective rank shows LRKV with sufficient rank matches Standard MHA at 128M, while at 2.5B LRKV achieves 98.3% vs 98.9% for MHA using 48.4% of KV cache.