Resolving the Academic Identity Crisis: A BSNMF Approach to Institutional Data Governance
Data Quality Management in Institutional Research Output Data Center
This paper introduces a comprehensive data quality management framework for developing institutional research output data centers. By utilizing a hybrid approach of Levenshtein text distance for department matching and Bayesian Symmetric Non-negative Matrix Factorization (BSNMF) for author name disambiguation, the authors achieve a 10.07% accuracy improvement in aggregating multi-affiliation scholar data.
TL;DR
Managing a university's research output is often a data science nightmare characterized by "same name, different person" and "same person, ten different affiliations." This paper presents an end-to-end framework for institutional data quality management that leverages the Levenshtein distance for fuzzy matching and Bayesian Symmetric Non-negative Matrix Factorization (BSNMF) to detect communities within co-author networks. The result? A 10% boost in data accuracy and a cleaner, more authoritative academic record.
Background: The Silica-fication of Research Data
In the era of "Double World-Class" initiatives, universities need unified data centers to support talent evaluation and subject development. However, data from sources like Web of Science or CNKI often lacks uniform standards. A researcher named "Yang Liu" might be listed under a State Key Lab in one paper and a School of Medicine in another. Without a robust Author Name Disambiguation (AND) strategy, the institutional "Big Data" becomes "Bad Data."
Methodology: The Three-Tiered Defense
The authors propose an eight-module pipeline, but the heavy lifting happens in the Matching and Fitting phases.
1. Fuzzy Matching with Levenshtein Distance
To handle the "non-standard abbreviation" problem (e.g., "Sch Agr & Biol" vs. "Coll Agr & Biol Sci"), the system uses an optimized Levenshtein algorithm. Unlike standard edit distance, the authors introduced domain-specific weights—for example, reducing distance penalties for predictable variations like "Affiliated Hospital 1" vs. "Affiliated Hospital 3" to ensure they are correctly distinguished or grouped.
2. Community Detection via BSNMF
The "Fitting" module treats the academic world as a graph. If two "Yang Liu" nodes share a dense network of co-authors, they are likely the same person. The paper employs BSNMF, a matrix learning method that captures the underlying community structure (modularity) of the co-author network.
Figure 1: The proposed conceptual framework for quality-controlled data integration.
Experiments: Performance at Scale
The study tested five community detection algorithms on a massive co-author network of 17,908 nodes derived from Shanghai Jiao Tong University's (SJTU) SCI output.
Comparative SOTA Analysis:
Compared to Greedy Community Detection (GCD) and Louvain (BGLL), the BSNMF method achieved a superior Modularity score of 0.7564, indicating it found more cohesive and meaningful academic clusters.
Table 2: Comparison of community detection methods. BSNMF achieves the highest Modularity.
Real-world Case: The "Yang Liu" Disambiguation
Through BSNMF, the system successfully identified that "Yang Liu" appearing across five different labels (e.g., "Stem Cell Res Ctr", "Renji Hosp") was actually a single entity in the School of Medicine, while a different "Yang Liu" belonged to the School of Life Science.
Critical Analysis & Future Outlook
The framework's strength lies in its iterative feedback loop (Data Center -> Faculty Claim -> Supervised Learning). By allowing scholars to "claim" or "reject" papers, the system continuously refines its feature set for the "Learning" module.
Limitations:
- Scalability: While matrix factorization is robust, the computational cost of BSNMF on global-scale networks (millions of nodes) remains a challenge compared to simpler heuristics.
- Cold Start: For new faculty members without a co-authorship history, the network-based "Fitting" module is less effective.
Conclusion: This research proves that institutional data quality isn't just an IT problem—it's a machine learning problem. By combining text similarity with latent community analysis, universities can finally bridge the gap between fragmented administrative records and a true "Big Data" academic ecosystem.
