
Scaling Laws for Data-Constrained Mixture Training
Apple Machine Learning studies how to mix scarce target-domain data with abundant generic data during language-model pretraining. Across more than 2,000 training runs, the research examines the trade-off between insufficient target exposure, repeated examples, diminishing returns, and overfitting.





