
Rethinking Single Extractor for LLM Pretraining
Apple ML challenges the use of single fixed HTML extractors for web-scale LLM pretraining datasets, citing suboptimal coverage of diverse web content. Different extractors produce substantially varying pages after filtering, despite similar performance on language tasks. The work explores better extraction practices for improved data utilization.







