TSA Methodology: Decoding Faculty Sentiments through Opinion Mining
Data Mining and Opinion Mining: A Tool in Educational Context
This paper introduces the TSA (Text Analysis-Sentiment Analysis) methodology, a specialized data mining framework designed to process and evaluate unstructured qualitative feedback in educational settings. By integrating Natural Language Processing (NLP) tools like TextBlob and vaderSentiment, the study achieves an automated interpretation of open-ended survey responses from university faculty, categorizing academic sentiments with high correlation between different mining models.
TL;DR
As universities pivot toward digital-first environments, understanding the qualitative "pulse" of faculty is vital. This paper introduces the TSA (Text Analysis-Sentiment Analysis) methodology, a framework that leverages NLP to automate the evaluation of open-ended survey questions. By combining lemmatization, automatic translation, and dual-model sentiment scoring, the authors transform "noise" into structured, actionable insights regarding the adoption of virtual classrooms.
The "Unstructured Data" Bottleneck in Education
In the era of Big Data, educational institutions are swimming in information. However, a significant portion of this data—specifically open-ended suggestions and recommendations—is often ignored because it is "unstructured."
Traditional Educational Data Mining (EDM) focuses on quantitative metrics like grades or click rates. The authors argue that by ignoring qualitative opinions, institutions lose the "Why" behind the "What." The challenge lies in the subjectivity of human language and the scarcity of specialized NLP corpora for educational contexts, particularly in languages other than English.
Methodology: The TSA Framework
The paper proposes a streamlined pipeline to bridge the gap between raw text and institutional decision-making.
Phase 1: Text Analysis (The Foundation)
The goal here is to distill raw feedback into its core components through two critical techniques:
- Lemmatization: Reducing words to their linguistic roots (e.g., "computing" and "computer" are treated as a single concept).
- Tokenization: Segmenting phrases into individual lexical units while filtering out "noise" words (under four characters).
Phase 2: Sentiment Analysis (The Insight)
Since many sentiment libraries are optimized for English, the authors implement an Automatic Translation step. Once translated, the text is passed through two distinct engines:
- TextBlob: Focuses on linguistic patterns.
- vaderSentiment: Specifically tuned for social media and intensity-aware sentiment.
Figure 1: The TSA Methodology involving Text Analysis and Sentiment Analysis.
Case Study: Faculty Perception of Virtual Classrooms
The authors applied TSA to a diagnosis of teachers at the National Polytechnic School, Ecuador.
Key Findings
The mining process revealed that the primary "words of interest" for faculty were centered around training, instructional design, and platform stability.
By comparing two sentiment analysis methods, the authors found a high correlation, suggesting that automated tools are now reliable enough for academic administration. By averaging the scores, they grouped the sentiment into seven polarity classes based on the Galton-Mac Law.
Table 1: Minimal variance between TextBlob (x1) and Vader (x2) ensures reliable sentiment mapping.
Results Breakdown
- 90.57% of sentiments were classified as Neutral or Positive.
- 9.43% expressed Negative feelings (mostly related to the need for better training).
Figure 2: Distribution of polarity values across the two methodology pilots.
Critical Analysis & Conclusion
Takeaway
The TSA methodology provides a "lens" that allows university leaders to view faculty feedback not as 177 individual complaints, but as a statistical landscape of needs. The high correlation between sentiment engines proves that the "subjectivity" of English translation didn't significantly break the sentiment direction.
Limitations
- Translation Reliance: The model relies on Google Translate. While efficient for Big Data, it may miss specific cultural nuances or academic slang inherent in the original Spanish text.
- Small Sample Complexity: While the methodology is built for Big Data, the case study was relatively small (54 responses), which required the authors to manually retain low-frequency words.
Future Outlook
The next logical step for this research is the automation of the entire pipeline and the application of this method to social media monitoring (e.g., tracking a university's reputation on Twitter or Facebook in real-time). With the rise of LLMs, the "TSA" approach could soon move from simple polarity (-1 to 1) to nuanced emotional mapping (frustration, excitement, confusion).
