Mapping the Peruvian Scientific Landscape: A Text Mining Approach to CV Analysis
Text Mining over Curriculum Vitae of Peruvian Professionals using Official Scientific Site DINA
This paper presents a Text Mining and exploratory data analysis of 25,000 professional profiles from Peru's National Directory of Researchers (DINA). By applying Natural Language Processing (NLP) to unstructured CV data, the study identifies key trends in academic backgrounds, scientific production, and linguistic competencies among Peruvian researchers.
TL;DR
This study applies Data Mining and NLP techniques to 25,000 curriculum vitae from Peru's official researcher database (DINA). It reveals a research community dominated by national university graduates with strong Master's level training but highlights a critical need for PhD growth and the preservation of indigenous languages like Quechua and Aymara.
Context: Harnessing Unstructured Academic Data
In the last decade, the Peruvian government, through Concytec, has significantly increased investment in Science and Technology. However, raw data on professionals—stored in the National Directory of Researchers (DINA)—remains largely unstructured. This paper addresses the "knowledge gap" by transforming these digital CVs into a structured overview of the nation's intellectual capital.
Methodology: The NLP Pipeline
The authors followed a standard yet robust Text Mining workflow to process the CV data:
- Scope Selection: Targeting the DINA platform as the primary source of truth.
- Preprocessing: Essential cleaning including the removal of "custom stopwords" such as experiencia and inicio which appear frequently in CVs but lack analytical value.
- Visualization: Using Word Clouds and distribution tables to interpret the "Inclination" of research lines and language skills.
Figure 1: High-frequency terms show a strong prevalence of "Universidad Nacional" and "Bachiller" within the Peruvian context.
Key Insights: Degrees, Research, and Languages
1. Academic Prowess and Institutional Hubs
The majority of professionals in the dataset were educated in national universities, which are free of charge in Peru. This distinguishes Peru from neighbors like Chile.
- Degree Distribution: 77% hold a Master’s degree, but only 28% have reached the PhD level.
- Scientific Production: The analysis of term frequencies like "Scopus", "ORCID", and "Elsevier" suggests a community increasingly aligned with international indexing standards.
2. The Language Barrier and Cultural Heritage
A unique aspect of this study is the granular analysis of language proficiency. While English is vital for global science (64% reading proficiency), the study looks inward at Peru's linguistic heritage:
- English: Most researchers possess intermediate to advanced reading skills, necessary for consuming global literature.
- Ancient Languages: Only a small fraction (around 4% for Quechua and <1% for Aymara) remain active in the scientific community. The authors warn that without intervention, these languages may face extinction within the professional sphere.
Table 1: Distribution of English levels across reading, speaking, and writing skills.
Critical Analysis & Conclusion
This paper serves as a foundational "snapshot" of Peru's scientific readiness. The shift from manual human resource selection to automated Text Mining allows for a macro-level understanding of an entire country's workforce.
Limitations: The study currently only analyzes 25,000 registers. Future work must scale this to the entire DINA database. Additionally, moving beyond simple frequency analysis toward Knowledge Graphs (as mentioned in the literature review) would allow for a deeper understanding of collaboration networks between institutions.
The Takeaway: For Peru to become a "sustainable nation through Science," it must leverage its national university graduates, bridge the PhD gap, and utilize its multilingual capabilities to foster international collaborations.
