LeSeAC: Bridging the Gap Between Generic Sentiment Analysis and Mexican Academic Colloquialisms

Design of a Semantic Lexicon Affective Applied to an Analysis Model Emotions for Teaching Evaluation

2015-01-01
Guadalupe Gutiérrez, Lourdes Margain, Alejandro Padilla, Juana Canul-Reich, Julio Ponce
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces LeSeAC (Semantic Affective Colloquial Lexical), a specialized linguistic resource designed for sentiment analysis in Mexican Higher Education. It leverages a multi-layer lexicon approach, integrating grammatical and graphic components with WordNet to identify four key emotions in student evaluations.

Executive Summary

TL;DR: This paper presents the design of LeSeAC, an Affective Semantic Lexical model specifically tuned for the Mexican educational context. By integrating regional colloquialisms and emotional mapping into a structured 7-phase analysis pipeline, the researchers provide a framework for transforming raw student comments into actionable data regarding professor performance.

Strategic Positioning: This work occupies a niche intersection between Computational Linguistics and Educational Psychology. It moves beyond simple "positive/negative" polarity to identify specific emotions like Admiration and Anger by addressing the "lexical gap" in localized Spanish dialects.


Problem & Motivation: The Failure of One-Size-Fits-All NLP

While Sentiment Analysis (SA) is a mature field, most "off-the-shelf" tools are optimized for English or standard Castilian Spanish. The authors argue that these tools often fail when faced with:

  1. Context Sensitivity: Educational feedback uses different descriptors than product reviews.
  2. Regional Expressions: Terms like "Fresa" or "Gacho" carry heavy emotional weight in Mexico but may be misinterpreted or ignored by standard lexicons.
  3. Ambiguity: Without a "Colloquialism Dictionary," the nuance of a student's frustration is often lost in translation.

The motivation here is to provide professors with constructive criticism rather than just a numerical score, requiring a model that truly "understands" the colloquial frustration or appreciation expressed in open-ended survey questions.


Methodology: The LeSeAC Architecture

The proposed model operates through a rigorous 7-phase pipeline, transitioning from raw data to semantic understanding.

The 7-Phase Pipeline

The core of the methodology revolves around Phase 2 (Grammatical Labeling) and Phase 4 (Semantic Analysis).

  1. Data Acquisition: Using 21-question evaluations from the Polytechnic University of Aguascalientes.
  2. Cleaning: Normalizing "veeeeeeeery" to "very" and handling misspellings.
  3. POS Tagging: Utilizing the FreeLing system to identify nouns, verbs, and adjectives.
  4. Semantic Mapping: Linking words to the custom LeSeAC XML database.

Model Architecture

The Role of Mexican Colloquialisms

Unlike standard affective lexicons, LeSeAC explicitly defines regional terms. For instance, the adjective "Sangrón" is mapped to the emotion "Dislike" as it denotes an unfriendly person. This mapping is structured in XML to allow for interoperability with WordNet synsets.


Experiments & Results: Identifying the Student "Mood"

The authors utilized a Proof of Concept validated by a panel of psychology experts to categorize the most frequent emotions found in student evaluations.

Key Findings

The study identified four dominant emotions:

  • Positive: Like and Admiration.
  • Negative: Dislike and Anger.

Table 2 highlights how the model interprets specific Spanish feedback:

  • "La Maestra tiene un gran nivel..." Admiration.
  • "El profesor es muy irresponsable..." Anger.

Emotion Examples Table

Technical Depth: Syntactic Analysis Levels

The researchers categorized their analysis into three depths:

  • Simple: Associating a single word with an emotion.
  • Double/Multiple: Looking at word pairings to capture modifiers (amplifiers or reducers).

Critical Analysis & Conclusion

Takeaway

The value of LeSeAC lies in its cultural grounding. By acknowledging that language is not just a set of grammar rules but a localized set of emotional markers, the authors have created a tool that is significantly more useful for Mexican institutions than generic SOTA models.

Limitations & Future Work

  • Scale: The current lexical matching is heavily dependent on a pre-defined dictionary.
  • Automation: The authors acknowledge that the lexicon extension needs automation. Their future plan involves using Neural Networks and SVMs to auto-expand the colloquial database as new slang emerges.
  • Graphic Lexicon: While mentioned, the analysis of "graphics" (emojis/mood icons) remains a secondary focus in the current report and represents a major area for expansion.

Final Verdict: This paper serves as a vital blueprint for building context-aware NLP systems in regions where dialectal variations are the primary mode of expression.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize the SentiSense lexicon framework for non-English sentiment analysis tasks.
  • Which methodologies are currently used to automate the expansion of colloquial lexicons using Support Vector Machines and Neural Networks in Spanish NLP?
  • Explore the application of multimodal sentiment analysis that combines text feedback with graphic icons (emojis) in the context of academic performance evaluations.
Contents
LeSeAC: Bridging the Gap Between Generic Sentiment Analysis and Mexican Academic Colloquialisms
1. Executive Summary
2. Problem & Motivation: The Failure of One-Size-Fits-All NLP
3. Methodology: The LeSeAC Architecture
3.1. The 7-Phase Pipeline
3.2. The Role of Mexican Colloquialisms
4. Experiments & Results: Identifying the Student "Mood"
4.1. Key Findings
4.2. Technical Depth: Syntactic Analysis Levels
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work