Mapping the Peruvian Scientific Landscape: A Text Mining Approach to CV Analysis

Text Mining over Curriculum Vitae of Peruvian Professionals using Official Scientific Site DINA

2020-12-01
Josimar Edinson Chire Saire, Honorio Apaza
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a Text Mining and exploratory data analysis of 25,000 professional profiles from Peru's National Directory of Researchers (DINA). By applying Natural Language Processing (NLP) to unstructured CV data, the study identifies key trends in academic backgrounds, scientific production, and linguistic competencies among Peruvian researchers.

TL;DR

This study applies Data Mining and NLP techniques to 25,000 curriculum vitae from Peru's official researcher database (DINA). It reveals a research community dominated by national university graduates with strong Master's level training but highlights a critical need for PhD growth and the preservation of indigenous languages like Quechua and Aymara.

Context: Harnessing Unstructured Academic Data

In the last decade, the Peruvian government, through Concytec, has significantly increased investment in Science and Technology. However, raw data on professionals—stored in the National Directory of Researchers (DINA)—remains largely unstructured. This paper addresses the "knowledge gap" by transforming these digital CVs into a structured overview of the nation's intellectual capital.

Methodology: The NLP Pipeline

The authors followed a standard yet robust Text Mining workflow to process the CV data:

  1. Scope Selection: Targeting the DINA platform as the primary source of truth.
  2. Preprocessing: Essential cleaning including the removal of "custom stopwords" such as experiencia and inicio which appear frequently in CVs but lack analytical value.
  3. Visualization: Using Word Clouds and distribution tables to interpret the "Inclination" of research lines and language skills.

Word Cloud of Academical Information Figure 1: High-frequency terms show a strong prevalence of "Universidad Nacional" and "Bachiller" within the Peruvian context.

Key Insights: Degrees, Research, and Languages

1. Academic Prowess and Institutional Hubs

The majority of professionals in the dataset were educated in national universities, which are free of charge in Peru. This distinguishes Peru from neighbors like Chile.

  • Degree Distribution: 77% hold a Master’s degree, but only 28% have reached the PhD level.
  • Scientific Production: The analysis of term frequencies like "Scopus", "ORCID", and "Elsevier" suggests a community increasingly aligned with international indexing standards.

2. The Language Barrier and Cultural Heritage

A unique aspect of this study is the granular analysis of language proficiency. While English is vital for global science (64% reading proficiency), the study looks inward at Peru's linguistic heritage:

  • English: Most researchers possess intermediate to advanced reading skills, necessary for consuming global literature.
  • Ancient Languages: Only a small fraction (around 4% for Quechua and <1% for Aymara) remain active in the scientific community. The authors warn that without intervention, these languages may face extinction within the professional sphere.

English Proficiency Distribution Table 1: Distribution of English levels across reading, speaking, and writing skills.

Critical Analysis & Conclusion

This paper serves as a foundational "snapshot" of Peru's scientific readiness. The shift from manual human resource selection to automated Text Mining allows for a macro-level understanding of an entire country's workforce.

Limitations: The study currently only analyzes 25,000 registers. Future work must scale this to the entire DINA database. Additionally, moving beyond simple frequency analysis toward Knowledge Graphs (as mentioned in the literature review) would allow for a deeper understanding of collaboration networks between institutions.

The Takeaway: For Peru to become a "sustainable nation through Science," it must leverage its national university graduates, bridge the PhD gap, and utilize its multilingual capabilities to foster international collaborations.

Find Similar Papers

Try Our Examples

  • Look for recent papers utilizing Knowledge Graphs or Ontologies to automate the ranking and selection of academic CVs in Latin American contexts.
  • Which paper first established the "Lattes" platform's data analysis methodology, and how does the current DINA study adapt those techniques for the Peruvian scientific ecosystem?
  • Evaluate how Natural Language Processing is being used to preserve endangered indigenous languages like Quechua and Aymara in computational linguistics research.
Contents
Mapping the Peruvian Scientific Landscape: A Text Mining Approach to CV Analysis
1. TL;DR
2. Context: Harnessing Unstructured Academic Data
3. Methodology: The NLP Pipeline
4. Key Insights: Degrees, Research, and Languages
4.1. 1. Academic Prowess and Institutional Hubs
4.2. 2. The Language Barrier and Cultural Heritage
5. Critical Analysis & Conclusion