Search

Tag: #model-scaling25 results

Hyperparameter Transfer Across All Scaling Axes

Hyperparameter Transfer Across All Scaling Axes

Apple ML extends μP for hyperparameter transfer across model sizes, modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal base hyperparameters search at small scales for large model transfer.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Hyperparam Transfer Across All Scales

Hyperparam Transfer Across All Scales

Apple extends μP for hyperparameter transfer across modules, width, depth, batch, and duration. Introduces Complete(d) Parameterisation unifying width-depth scaling. Enables optimal hypers from small to large models.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Complete Hyperparameter Transfer for Scaling

Complete Hyperparameter Transfer for Scaling

Apple ML extends μP parameterisations with Complete(d) Parameterisation for hyperparameter transfer. Covers scaling across modules, width, depth, batch size, and duration. Enables optimal hyperparameter search on small models for transfer to large-scale ones.

Apple Machine LearningOfficialFeb 13#research#apple-ml#mu-p
Page 3 of 3