π¦Reddit r/LocalLLaMAβ’Stalecollected in 2h
SFM Beats Transformers on Long Sequences
#non-transformer#long-sequence#state-slots#benchmarkstate-flow-machinestate-flow-machinesfmascend-910huawei
π‘Non-transformer holds 62% acc where LMs crash at long seqsβnew arch!
β‘ 30-Second TL;DR
What Changed
62% acc at 4x length (40 ops) vs 2-3% for transformers
Why It Matters
Challenges transformer dominance for long-sequence stateful tasks like process simulation. Could inspire efficient on-device architectures beyond attention limits.
What To Do Next
Replicate SFM benchmark on Ascend NPU to test long-seq generalization.
Who should care:Researchers & Academics
Key Points
- β’62% acc at 4x length (40 ops) vs 2-3% for transformers
- β’Explicit addressable state slots updated via learned gates
- β’961K params; benchmark extrapolates to 320 ops
- β’Inspired by DeltaNet, Mamba but more explicit
π°
Weekly AI Recap
Read this week's curated digest of top AI events β
πRelated Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: Reddit r/LocalLLaMA β
This is a summary, not the original. Read the source, or get the weekly briefing.
Weekly AI briefing
One email a week. Unsubscribe anytime.