面壁智能开源 Ultra-FineWeb-L1,提供 1T+ tokens 高质量英文网页数据,约 11.4 亿篇文档,来自最新 Common Crawl 快照(覆盖至 CC-MAIN-2025-51)。该数据集属于 UltraData 生态,作为 L1 过滤层,经文本提取、语言过滤、启发式过滤、敏感字段替换、去重和定制清洗等流程处理,旨在提升数据质量。数据集、论文和分类器均已公开。
🔥 Ultra-FineWeb-L1 is here —— 1T+ tokens of high-quality web data are now available!
High-quality data is the foundation of powerful LLMs. Ultra-FineWeb-L1 is now open-sourced as part of the UltraData ecosystem, providing a carefully processed English web corpus from Common Crawl.
Built for Better Data Quality
1️⃣ Fresh Web Data: Collected from recent Common Crawl snapshots, covering data up to CC-MAIN-2025-51, with 1T+ tokens and ~1.14B documents.
2️⃣ Refined Processing Pipeline: Combines text extraction, language filtering, heuristic filtering, sensitive-field replacement, deduplication, and customized cleaning to improve data quality.
⚡ Highlights:
- 1T+ tokens of filtered English web data
- Enhanced cleaning for noisy web content, encoding issues, and abnormal documents
- Designed as the L1 filtered layer in the UltraData management framework
- Aligned with the latest Ultra-FineWeb updates https://t.co/pXDfgulD38
🔗 Resources
🤗 Dataset: https://t.co/3aR5qtgw53
📄 Paper: https://t.co/Kg9LLUqZgB
🧩 Classifier: https://t.co/sP5yLlL68n
来源:@OpenBMB · x.com