🍎Apple Machine Learning•較早收集於 18h
超越單一提取器:重新思考 LLM 預訓練的 HTML 轉文字提取

💡Unlock better LLM training data: single extractors fail coverage—fix your pipeline now.
⚡ 30-Second TL;DR
有什麼變化
單一固定提取器限制 LLM 資料集對多樣網路內容的涵蓋。
為什麼重要
這可能提升 LLM 預訓練資料品質,產生更穩健且涵蓋更廣網路知識的模型。它促使 AI 研究重新評估資料管道,減少不良提取造成的偏差。
下一步行動
Test multiple extractors like Trafilatura and boilerpy3 on your LLM web dataset pipeline.
誰應關注:Researchers & Academics
關鍵要點
- •單一固定提取器限制 LLM 資料集對多樣網路內容的涵蓋。
- •不同提取器導致過濾後存活頁面差異巨大。
- •語言模型表現相似掩蓋了資料品質與利用率的差異。
- •主張重新思考預處理以更好利用網路資料。
🧠 深度解析
背景與延伸:來自公開資料,非原文內容。引用 8 個來源。
🔑 增強重點摘要
- •Apple's research identifies that data filtering techniques significantly impact which web pages survive preprocessing pipelines, with model-based filtering approaches retaining more informative content than aggressive heuristic rules—a finding that directly addresses extractor limitations by showing filtering strategy is equally critical to extraction method selection[4].
- •Apple has expanded multilingual support and incorporated substantially more high-quality mathematical and programming content in their 2025 foundation model updates, indicating that extractor diversity becomes increasingly important as pretraining datasets target specialized domains beyond general text[4].
- •The Foundation Models framework now provides developers direct access to create production-quality generative AI features, suggesting that standardized extraction practices are becoming infrastructure-level concerns rather than isolated research problems[4].
🔮 前景展望AI analysis grounded in cited sources
Extractor diversity will become a standardized benchmark metric for evaluating LLM pretraining dataset quality
As Apple's research demonstrates substantial variation in surviving pages across different extractors despite similar downstream model performance, the field will likely adopt multi-extractor evaluation as a prerequisite for dataset transparency and reproducibility.
Model-based filtering will replace heuristic-only approaches in production LLM pretraining pipelines
Apple's 2025 updates show that incorporating model-informed signals into data filtering pipelines retains more informative content while maintaining quality, establishing a precedent that adaptive filtering outperforms fixed rule-based extraction.
⏳ 時間線
2023-06
Apple releases AXLearn, an open-source framework for training foundation models with high efficiency and scalability across heterogeneous hardware[1]
2025-06
Apple presents updated foundation models with expanded multilingual support, improved data filtering pipelines, and incorporation of mathematical and programming content[4]
2026-02
Apple Machine Learning publishes research on data-quality filtering for LLM pretraining, addressing extractor and filtering methodology challenges[6]
📎 來源 (8)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
- machinelearning.apple.com — Introducing Apple Foundation Models
- developer.apple.com — 360
- developer.apple.com — Models
- machinelearning.apple.com — Apple Foundation Models 2025 Updates
- developer.apple.com — Machine Learning
- machinelearning.apple.com — Research
- machinelearning.apple.com — Icml 2025
- paperdigest.org — Iclr 2026 Papers Highlights
📰
AI 週報
閱讀本週精選 AI 大事摘要 →
👉相關動態
AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: Apple Machine Learning ↗
每週 AI 簡報
每週一封,可隨時退訂。