
New Softmax-free Attention Model with Structural Sparsity Released
A new 354M parameter attention model has been released, featuring a softmax-free architecture and custom Triton kernels. It leverages structural sparsity and tile-skipping to significantly reduce VRAM usage for long-context tasks.






