
Native Factorized Weights: Optimizing Transformer Rank for Generalization
The author introduces Native Factorized Weights (NFW), a method that initializes transformer layers as low-rank matrices to improve training efficiency and generalization. The research identifies a corpus-determined optimal rank that prevents memorization while outperforming dense models.




