SogouT-16: Empowering Information Retrieval with a Billion-Page Chinese Web Corpus
SogouT-16: A New Web Corpus to Embrace IR Research
SogouT-16 is a massive, open-access Chinese Web corpus containing 1.17 billion pages from 2.03 million domains, sampled from the Sogou search engine. It establishes a new SOTA benchmark for Chinese Information Retrieval (IR) research, specifically designed for the NTCIR-13 "We Want Web" ad-hoc retrieval task.
TL;DR
SogouT-16 is the largest free public Chinese Web collection designed to revitalize Information Retrieval (IR) research. Spanning 1.17 billion pages, it introduces a refined sampling strategy balance between PageRank (quality) and site diversity, providing a critical infrastructure for current benchmarks like NTCIR-13.
The Scalability Bottleneck in IR Research
In the era of massive web growth, academic researchers often face a "data poverty" paradox. While the Web contains billions of pages, existing datasets like WT10g or SogouT-12 represent a web that no longer exists—one that is smaller, less diverse, and lacks the structural complexity of the modern Chinese internet.
The authors identify two primary pain points:
- The Scale-Feasibility Gap: Datasets like CWP200T are massive (200TB) but practically impossible for university labs to process without industrial-scale clusters.
- The Language Bias: Global benchmarks like ClueWeb12 intentionally filter out non-English pages, leaving Chinese IR research in a state of stagnation.
Methodology: Quality vs. Diversity
The core innovation of SogouT-16 lies in its sampling algorithm. How do you select 1 billion pages from a commercial search engine index while ensuring they aren't all from the same few "mega-sites"?
The authors proposed a probability-based selection:
- PageRank Weighting: Pages with higher popularity (Importance) are prioritized.
- Site Penalization: As more pages are picked from a single domain (e.g., a news portal), the chance of picking the next page from that site decreases logarithmically.
This ensures the corpus isn't just a mirror of a few giant domains but a representative sample of the 2.03 million domains available.
Figure 1: The sampling algorithm ensures a stable proportion of high-quality pages across diverse domains.
Technical Specs & Data Cleaning
Raw web data is notoriously "dirty." The SogouT-16 pipeline includes:
- Spam & Malware Filtering: Removal of phishing and virus-infected pages.
- Language Identification: Strict filtering to ensure a Chinese-centric corpus.
- Encoding Normalization: All files converted to UTF-8, solving a major headache for Chinese text processing where ASCII, GBK, and UTF-8 often conflict.
| Dataset | #Pages | Language | Latest Crawl |
|---|---|---|---|
| SogouT-16 | 1.17B | CHN | 2016 |
| SogouT-12 | 0.13B | CHN | 2012 |
| ClueWeb12 | 0.73B | ENG | 2012 |
| CWP200T | 7.00B | CHN | 2015 |
Comparison table highlights that SogouT-16 is the most balanced large-scale Chinese corpus for general IR use.
Beyond Just HTML: A Research Ecosystem
SogouT-16 isn't just a dump of HTML. To lower the barrier to entry, the authors released several auxiliary resources:
- SogouT-16-B (Category B): A 1.5TB subset for researchers who can't handle the full 81TB (uncompressed).
- Link Structure Graph: Crucial for studying Web topology and developing new ranking algorithms.
- Query Logs: Anonymized real-world search behavior, a rarity in public datasets.
- Word Embeddings: Pre-trained vectors to jumpstart NLP tasks.
Figure 2: Analysis showing the convergence of sample ratios as site size grows, validating the diversity of the corpus.
Critical Analysis & Future Impact
While SogouT-16 provides a robust foundation, its primary limitation is the static nature of the crawl (2016). In the rapidly evolving mobile web and social media landscape, static corpora struggle to capture dynamic content. However, for ad-hoc retrieval research and the development of state-of-the-art Chinese ranking models, SogouT-16 remains a gold standard.
By providing free online retrieval services (via Apache Solr), the authors have effectively reduced the "entry cost" for IR research, allowing researchers to query the billion-page index without needing a private supercomputer. This work is a significant step toward making large-scale web research accessible to the global academic community.
