Crowdsourcing the CRIS: Solving the Data Quality Crisis in Research Systems
Crowdsourcing Opportunities for Research Information Systems
This paper explores integrating crowdsourcing into Current Research Information Systems (CRIS), specifically the Russian "Elibrary," to solve data quality issues like citation mismatches. It proposes a model where users validate suspected citation links, demonstrating that while intrinsic motivation is helpful, a points-based incentive system is necessary for large-scale data cleansing.
TL;DR
Scientific impact is often measured by data that is fundamentally broken. This paper investigates how to fix the massive backlogs in Research Information Systems (CRIS) by turning researchers into "micro-moderators." By shifting citation linkage from manual staff review to a crowdsourced, incentive-driven model, systems like Russia's Elibrary can achieve 60-80% increases in recorded citation counts and significantly more accurate institutional rankings.
The Problem: The "Moderation Bottleneck"
In many academic databases, data quality is a paid luxury. Organizations pay for the right to correct their employees' metadata. Even then, every change must be manually approved by a central moderator. This creates a "Moderation Bottleneck" where:
- Verifying a single citation can take 6 months.
- Reporting dates pass before corrections are live, leading to distorted bibliometric indicators.
- For the Russian Elibrary system, estimates suggest that 7.7% of publications and 27.6% of citations currently require correction.
Methodology: A Distributed Validation Engine
The author suggests that citation linkage—identifying if a typo-ridden reference actually refers to a specific paper—is a task humans do intuitively but machines struggle with.
The Model
- Analyzer: An algorithm flags potential matches between references and publications.
- List of Tasks: These "suspected links" are queued for human review.
- The 1-for-3 Rule: To submit one correction of their own, a user must validate three pairs from the queue.
- Consensus Control: A "majority rule" (typically 3 votes) is required to finalize a deletion or addition to the database.
(Note: Users act as filters for the automated suggestions, ensuring high precision.)
Experimental Insights: The Power of Accuracy
To prove the value of clean data, the author looked at the Central Economics and Mathematics Institute (CEMI RAS). Within one year of manual correction efforts:
- Citations for 2010 increased by 80.2%.
- The institute jumped from 120th to 73rd place in the national h-index rankings.
The Motivation Gap (Simulation Results)
The paper uses a Monte-Carlo simulation to ask: Can intrinsic motivation (the desire to look good on paper) sustain this?
- Finding: While 4,000 active users could clear the backlog, the workload per person is too high (~111 checks per task).
- Solution: The author proposes an External Motivation system. Users earn "points" for moderation which can be exchanged for access to premium services (e.g., uploading full-text papers) without involving cash transactions.
(The simulation shows the distribution of tasks vs. user capacity, highlighting where internal motivation peaks and extra incentives must take over.)
Critical Analysis & Conclusion
The brilliance of this approach is its Business Neutrality. The author argues that CRIS operators won't lose money because organizations will still pay for bulk management and "prestige" features. Meanwhile, the database gets cleaned for free by the community.
Limitations:
- The paper assumes users are honest; it doesn't deeply explore "malicious optimization" where users might collude to upvote fake citations.
- The technical difficulty of retrofitting legacy CRIS platforms with these real-time voting modules is acknowledged but not solved.
The Takeaway: Data quality is the bedrock of science policy. By gamifying the "boring" task of citation checking, CRIS platforms can transform from stagnant archives into living, community-curated ecosystems.
