AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity
EMNLP 2025
Sparse AttentionEfficient LLMs
Summary
AnchorAttention accelerates long-context LLM prefill with dynamically selected, stripe-shaped attention regions. It uses an anchor derived from shared attention patterns to identify important positions, then computes attention through discrete key-value loading. This finer granularity improves sparsity while retaining useful attention information. At a context length of 128k tokens, the method reports a 1.44× speedup over earlier leading approaches, together with higher recall.
BibTeX
@inproceedings{zhang-etal-2025-anchorattention,
title = {AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity},
author = {Zhang, Yu and Guo, Dong and Wu, Fang and Zhu, Guoliang and Ding, Dian and Zhang, Yiming},
booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing},
year = {2025},
pages = {8537--8549},
doi = {10.18653/v1/2025.emnlp-main.430},
url = {https://aclanthology.org/2025.emnlp-main.430/}
}