Publications

2025

AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity

Yu Zhang, Dong Guo, Fang Wu, Guoliang Zhu, Dian Ding, Yiming Zhang

EMNLP 2025

Sparse AttentionEfficient LLMs
Summary

AnchorAttention accelerates long-context LLM prefill with dynamically selected, stripe-shaped attention regions. It uses an anchor derived from shared attention patterns to identify important positions, then computes attention through discrete key-value loading. This finer granularity improves sparsity while retaining useful attention information. At a context length of 128k tokens, the method reports a 1.44× speedup over earlier leading approaches, together with higher recall.

BibTeX
@inproceedings{zhang-etal-2025-anchorattention,
  title = {AnchorAttention: Difference-Aware Sparse Attention with Stripe Granularity},
  author = {Zhang, Yu and Guo, Dong and Wu, Fang and Zhu, Guoliang and Ding, Dian and Zhang, Yiming},
  booktitle = {Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing},
  year = {2025},
  pages = {8537--8549},
  doi = {10.18653/v1/2025.emnlp-main.430},
  url = {https://aclanthology.org/2025.emnlp-main.430/}
}