Harvesting the Digital Steppe: Collecting Minor Languages of Russia from Social Networks
Languages of Russia: Using Social Networks to Collect Texts
The paper introduces a method for identifying and collecting linguistic corpora for minor languages of Russia from social networks, specifically VKontakte. By leveraging manually selected "lexical markers" and the Yandex.XML API, the authors automated the discovery of community pages to build natural, informal datasets for low-resource languages.
Executive Summary
TL;DR: Researchers from the Higher School of Economics (Moscow) have developed a robust pipeline to scrape and categorize texts in the minor languages of Russia from VKontakte. By identifying "Lexical Markers"—unique linguistic fingerprints—the team bypasses the scarcity of official digitized texts to build natural, informal datasets.
Positioning: This work is a crucial "infrastructure-enabling" study. It focuses on the data bottleneck for low-resourced languages, prioritizing the collection of "everyday written language" over the formal, often localized translations found on Wikipedia.
Motivation: The Failure of Wikipedia for Minor Languages
For many minor languages, Wikipedia is the only game in town. However, the authors argue it is often a poor linguistic source. A previous study on Bashkir and Tatar showed that the most "frequent" words on their respective Wikipedias were technical terms like "river" and "basin," rather than natural functional words.
Furthermore, the Internet presents a unique challenge: the diacritic problem. Speakers of minor languages often use Russian layouts, replacing specific characters (e.g., "ң" with "н"). This "Everyday Written Language" (EWL) makes traditional dictionary-based crawling ineffective.
Methodology: The Lexical Marker Pipeline
The heart of the project is a two-step discovery process designed to handle the complexity of 97 different national languages.
1. Lexical Markers
Instead of full dictionaries, the team manually curates a small set of Lexical Markers. To be a marker, a word must:
- Be Unique: Only occur in that specific language (verified via search engines).
- Be Frequent: Usually function words (pronouns, conjunctions).
- Be Cyrillic: Valid even when special diacritics are omitted.
2. Search and API Extraction
Once markers are identified, the workflow proceeds as follows:
- Yandex.XML Querying: Automated search queries find URLs containing these markers within the
vk.comdomain. - Community Focus: The authors specifically target "Community" pages rather than personal profiles, as communities provide a more concentrated linguistic environment.
- VK API Integration: Using the VKontakte API, the system downloads not just the text, but the social metadata: user age, city, and post comments.
Note: The table above illustrates the significant scaling achieved by switching from manual to automated discovery.
Experiments & Preliminary Results
The results demonstrate that automated discovery massively outscales manual efforts. For the Udmurt language, manual lists contained 72 communities, while the automated script identified 769 potential sources.
The "Suro-pojo" Challenge
The researchers encountered a major hurdle: Code-switching. In social networks, speakers often mix Russian and their native tongue in a single sentence (e.g., "Steve Jobs - honorary Udmurt"). This hybrid language, known as suro-pojo (mix), makes automatic language identification extremely difficult for standard tools.
Deep Insight & Conclusion
Takeaway
The value of this research lies in its sociolinguistic realism. By targeting social networks, the researchers are capturing how these languages are actually used by younger generations, including the adaptations they make to Russian keyboards and the heavy influence of the Russian language.
Limitations & Future Work
The primary bottleneck remains the API limitations (3 requests per second) and the difficulty of clean language separation. However, the data collected provides a foundation for training modern NLP models (like Transformers) that can eventually understand these nuanced, multi-lingual social interactions.
The authors plan to share these structured datasets with the global research community to encourage the development of better linguistic tools for Russia's diverse ethnic landscape.
