SogouT-16: Empowering Information Retrieval with a Billion-Page Chinese Web Corpus

SogouT-16: A New Web Corpus to Embrace IR Research

2017-07-28
Cheng Luo, Yukun Zheng, Yiqun Liu, Xiaochuan Wang, Jingfang Xu, Min Zhang, Shaoping Ma, Shaoping Ma
Summary
Problem
Method
Results
Takeaways
Abstract

SogouT-16 is a massive, open-access Chinese Web corpus containing 1.17 billion pages from 2.03 million domains, sampled from the Sogou search engine. It establishes a new SOTA benchmark for Chinese Information Retrieval (IR) research, specifically designed for the NTCIR-13 "We Want Web" ad-hoc retrieval task.

TL;DR

SogouT-16 is the largest free public Chinese Web collection designed to revitalize Information Retrieval (IR) research. Spanning 1.17 billion pages, it introduces a refined sampling strategy balance between PageRank (quality) and site diversity, providing a critical infrastructure for current benchmarks like NTCIR-13.

The Scalability Bottleneck in IR Research

In the era of massive web growth, academic researchers often face a "data poverty" paradox. While the Web contains billions of pages, existing datasets like WT10g or SogouT-12 represent a web that no longer exists—one that is smaller, less diverse, and lacks the structural complexity of the modern Chinese internet.

The authors identify two primary pain points:

  1. The Scale-Feasibility Gap: Datasets like CWP200T are massive (200TB) but practically impossible for university labs to process without industrial-scale clusters.
  2. The Language Bias: Global benchmarks like ClueWeb12 intentionally filter out non-English pages, leaving Chinese IR research in a state of stagnation.

Methodology: Quality vs. Diversity

The core innovation of SogouT-16 lies in its sampling algorithm. How do you select 1 billion pages from a commercial search engine index while ensuring they aren't all from the same few "mega-sites"?

The authors proposed a probability-based selection:

  • PageRank Weighting: Pages with higher popularity (Importance) are prioritized.
  • Site Penalization: As more pages are picked from a single domain (e.g., a news portal), the chance of picking the next page from that site decreases logarithmically.

This ensures the corpus isn't just a mirror of a few giant domains but a representative sample of the 2.03 million domains available.

SogouT-16 Sampling Logic Figure 1: The sampling algorithm ensures a stable proportion of high-quality pages across diverse domains.


Technical Specs & Data Cleaning

Raw web data is notoriously "dirty." The SogouT-16 pipeline includes:

  • Spam & Malware Filtering: Removal of phishing and virus-infected pages.
  • Language Identification: Strict filtering to ensure a Chinese-centric corpus.
  • Encoding Normalization: All files converted to UTF-8, solving a major headache for Chinese text processing where ASCII, GBK, and UTF-8 often conflict.
Dataset#PagesLanguageLatest Crawl
SogouT-161.17BCHN2016
SogouT-120.13BCHN2012
ClueWeb120.73BENG2012
CWP200T7.00BCHN2015

Comparison table highlights that SogouT-16 is the most balanced large-scale Chinese corpus for general IR use.


Beyond Just HTML: A Research Ecosystem

SogouT-16 isn't just a dump of HTML. To lower the barrier to entry, the authors released several auxiliary resources:

  1. SogouT-16-B (Category B): A 1.5TB subset for researchers who can't handle the full 81TB (uncompressed).
  2. Link Structure Graph: Crucial for studying Web topology and developing new ranking algorithms.
  3. Query Logs: Anonymized real-world search behavior, a rarity in public datasets.
  4. Word Embeddings: Pre-trained vectors to jumpstart NLP tasks.

Sample Ratio Analysis Figure 2: Analysis showing the convergence of sample ratios as site size grows, validating the diversity of the corpus.

Critical Analysis & Future Impact

While SogouT-16 provides a robust foundation, its primary limitation is the static nature of the crawl (2016). In the rapidly evolving mobile web and social media landscape, static corpora struggle to capture dynamic content. However, for ad-hoc retrieval research and the development of state-of-the-art Chinese ranking models, SogouT-16 remains a gold standard.

By providing free online retrieval services (via Apache Solr), the authors have effectively reduced the "entry cost" for IR research, allowing researchers to query the billion-page index without needing a private supercomputer. This work is a significant step toward making large-scale web research accessible to the global academic community.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize SogouT-16 or more recent Chinese web corpora for training Large Language Models or IR benchmarking.
  • What are the original theories behind the Cranfield framework for IR evaluation, and how did it influence the design of the NTCIR-13 "We Want Web" (WWW) task?
  • Explore how the PageRank-based sampling methodology used in SogouT-16 has been adapted for creating diverse datasets in Multimodal or Cross-lingual information retrieval.
Contents
SogouT-16: Empowering Information Retrieval with a Billion-Page Chinese Web Corpus
1. TL;DR
2. The Scalability Bottleneck in IR Research
3. Methodology: Quality vs. Diversity
4. Technical Specs & Data Cleaning
5. Beyond Just HTML: A Research Ecosystem
6. Critical Analysis & Future Impact