
社群對 124B Gemma 模型的期待
討論社群對於 Google Gemma 模型推出更大規模 124B 參數版本的期待。這反映了市場對高參數開源權重模型的持續需求。
Tag: #model-scaling25 results

討論社群對於 Google Gemma 模型推出更大規模 124B 參數版本的期待。這反映了市場對高參數開源權重模型的持續需求。

Apple ML extends μP for hyperparameter transfer across model sizes, modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal base hyperparameters search at small scales for large model transfer.

Apple extends μP for hyperparameter transfer across modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal hypers from small to large models.

Extends hyperparameter transfer from small to large models across modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling, building on μP.

Apple ML extends μP parameterisations with Complete(d) Parameterisation for hyperparameter transfer. Covers scaling across modules, width, depth, batch size, and duration. Enables optimal hyperparameter search on small models for transfer to large-scale ones.