ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
Liu, Xiang, Tang, Zhenheng, Dong, Peijie · arXiv · 2025