--- license: apache-2.0 --- # Ultra-FineWeb-Classifier
π Technical Report | π€ Ultra-FineWeb | π¦ UltraData Collection | π UltraData
English | δΈζ
## π Introduction ***Ultra-FineWeb-Classifier*** is a lightweight bilingual quality classifier for selecting high-quality documents from large-scale English and Chinese web corpora. It is developed through the efficient verification-based filtering pipeline proposed in the [Ultra-FineWeb technical report](https://arxiv.org/abs/2505.05427) and implemented with [fastText](https://fasttext.cc/) to provide efficient, low-cost inference at web scale. The repository provides separate English and Chinese classifier weights. Applying these classifiers to FineWeb and Chinese FineWeb produces [Ultra-FineWeb](https://huggingface.co/datasets/openbmb/Ultra-FineWeb), a higher-quality web pre-training dataset containing approximately **1T English tokens** and **120B Chinese tokens**. Ultra-FineWeb serves as a core pre-training web dataset for the [MiniCPM4 Series](https://huggingface.co/collections/openbmb/minicpm4) and [MiniCPM5 Series](https://huggingface.co/collections/openbmb/minicpm5). - [Ultra-FineWeb-L1](https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1): **L1 filtered data** after basic cleaning, heuristic filtering, sensitive-field replacement, and deduplication. - [Ultra-FineWeb](https://huggingface.co/datasets/openbmb/Ultra-FineWeb): **L2 selected data** containing approximately **1T English tokens** and **120B Chinese tokens**. - [Ultra-FineWeb-Classifier](https://huggingface.co/openbmb/Ultra-FineWeb-classifier): lightweight English and Chinese quality classifiers for filtering web corpora. (**Current classifier**) - [Ultra-FineWeb-L3](https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3): **L3 refined data** built via Q&A pair generation and multi-style rewriting, containing **400B+ English tokens** and **200B+ Chinese tokens**. ## π’ What's New - **[2026.08.20]** The [***Ultra-FineWeb-L1***](https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L1) dataset is released, together with the L2 selected subset produced by the **Ultra-FineWeb-Classifier**. πππ - **[2026.05.28]** The [***Ultra-FineWeb-L3***](https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3) dataset is released, containing **400B+ English tokens** and **200B+ Chinese tokens**. - **[2026.05.25]** [***MiniCPM5-1B***](https://huggingface.co/openbmb/MiniCPM5-1B) is released! Ultra-FineWeb serves as its core pre-training web dataset. - **[2026.02.08]** The [***UltraData***](https://ultradata.openbmb.cn/) platform is now live, introducing the [L0-L4 tiered data management framework](https://arxiv.org/abs/2602.09003). - **[2025.06.16]** The **Ultra-FineWeb-classifier** is now available on Hugging Face: [openbmb/Ultra-FineWeb-classifier](https://huggingface.co/openbmb/Ultra-FineWeb-classifier). πππ - **[2025.06.06]** **Ultra-FineWeb-en** and **Ultra-FineWeb-zh** are now available on Hugging Face, released alongside the [MiniCPM4 Series](https://huggingface.co/collections/openbmb/minicpm4) models. - **[2025.05.15]** **Ultra-FineWeb** tops the Hugging Face Datasets Trending list, reaching the #1 spot! βοΈβοΈβοΈ - **[2025.05.09]** The **Ultra-FineWeb** technical report is available on [arXiv](https://arxiv.org/abs/2505.05427). π₯π₯π₯ ## π‘ Highlights > **Abstract:** Data quality has become a key factor in enhancing model performance with the rapid development of large language models (LLMs). Model-driven data filtering has increasingly become a primary approach for acquiring high-quality data. However, it still faces two main challenges: (1) the lack of an efficient data verification strategy makes it difficult to provide timely feedback on data quality; and (2) the selection of seed data for training classifiers lacks clear criteria and relies heavily on human expertise, introducing a degree of subjectivity. To address the first challenge, we introduce an efficient verification strategy that enables rapid evaluation of the impact of data on LLM training with minimal computational cost. To tackle the second challenge, we build upon the assumption that high-quality seed data is beneficial for LLM training, and by integrating the proposed verification strategy, we optimize the selection of positive and negative samples and propose an efficient data filtering pipeline. This pipeline not only improves filtering efficiency, classifier quality, and robustness, but also significantly reduces experimental and inference costs. In addition, to efficiently filter high-quality data, we employ a lightweight classifier based on *fastText*, and successfully apply the filtering pipeline to two widely-used pre-training corpora, *FineWeb* and *Chinese FineWeb* datasets, resulting in the creation of the higher-quality ***Ultra-FineWeb*** dataset. ***Ultra-FineWeb*** contains approximately 1 trillion (T) English tokens and 120 billion (B) Chinese tokens. Empirical results demonstrate that the LLMs trained on Ultra-FineWeb exhibit significant performance improvements across multiple benchmark tasks, validating the effectiveness of our pipeline in enhancing both data quality and training efficiency.