Sliding-Window Attention Challenges Linear Attention
A new arXiv preprint argues that Sliding Window Attention with sinks outperforms linear-attention variants on long-context reasoning benchmarks. On Needle-in-a-Haystack and BABILong, SWA reportedly delivers 2–10 times higher performance without post-training.
Reddit r/MachineLearning · 16d ago

























