Decoding Uyghur Metaphor: Multi-Attention and Emotional Consistency in Low-Resource NLP
Uyghur Metaphor Detection via Considering Emotional Consistency
This paper introduces the first deep learning framework for Uyghur metaphor detection, treating the task as a sequence tagging problem. The authors propose a Multi-Attention BiLSTM model that integrates word embeddings (GloVe and ELMo), part-of-speech (POS), position, and specifically emotional consistency features to achieve SOTA performance on a newly constructed Uyghur metaphor corpus.
TL;DR
This research tackles the challenging task of metaphor detection in Uyghur, a low-resource minority language. By treating metaphor detection as a sequence tagging task, the authors introduce a Multi-Attention BiLSTM framework. The core innovation lies in leveraging POS tagging, relative positioning, and emotional consistency—the idea that metaphors often stand out due to emotional "clashes" with their surrounding context.
Background & Motivation: Why Uyghur?
Metaphor is not just a rhetorical device; it is a cognitive mapping from one concept to another (e.g., "Knowledge is treasure"). While English metaphor detection has flourished with datasets like VUA and MOH-X, Uyghur has remained largely unexplored.
The difficulty stems from two factors:
- Linguistic Complexity: Uyghur is an agglutinative language with rich morphology (cases, voices, tenses) and specific metaphorical rules (e.g., marked similes using "dek" or "oxshash").
- Cultural Context: Animals or objects carry different metaphorical weights; for instance, "cat" in Uyghur implies "greed" rather than "cuteness," and "pumpkin" can mean a "fool."
Methodology: The Core Architecture
The proposed model moves beyond simple word embeddings to create a multi-dimensional representation of every token.
1. Emotional Consistency Embedding
The authors hypothesize that metaphors often disrupt the emotional flow of a sentence. They constructed an emotional dictionary (Positive, Negative, Neutral) and concatenated these embeddings () with GloVe and ELMo vectors.
2. Multi-Attention Layer
To capture the nuance of Uyghur's grid grammar, the model uses two specific attention mechanisms:
- POS Attention: Since metaphors are frequently tied to specific parts of speech (verbs and nouns), this layer prioritizes relevant syntactic categories.
- Position Attention: This encodes the relative distance between words, helping the model understand long-distance semantic dependencies.
Figure 1: The proposed Multi-Attention DNN framework, illustrating the fusion of emotional, word, and syntactic features.
Experiments and Results
The researchers curated a new dataset of 5,605 Uyghur sentences. By comparing different iterations, they demonstrated the necessity of each added feature:
| Model Configuration | Precision | Recall | F-score |
|---|---|---|---|
| Baseline (BiLSTM) | 57.8 | 65.1 | 61.2 |
| + POS Attention | 60.5 | 66.2 | 63.2 |
| Full Multi-Attention | 62.8 | 67.2 | 64.9 |
Key Insights:
- Emotional Clues: As shown in the ablation studies, the F-score significant drops when emotional features are removed. Metaphors are detected more accurately when the word’s emotional polarity conflicts with the overall sentence sentiment.
- POS is Vital: Adding POS information resulted in a significant 2% jump in F-score, proving that structural labels act as a powerful anchor in morphologically complex languages.
Figure 2: Performance drop after removing Multi-feature inputs (MTinput) and Multi-attention (MTatt).
Critical Analysis & Future Outlook
Takeaway: This work proves that for minority languages with limited data, "hard-coded" linguistic features like POS and emotional polarities are not just helpful—they are essential to compensate for the sparsity of word embeddings.
Limitations: The current F-score (64.9%) indicates there is still room for improvement. The reliance on a manually constructed emotional dictionary might limit the model's ability to generalize to modern internet slang or evolving metaphorical usage not captured in the static lexicon.
Future Work: Integrating Transformer-based models like BERT (pre-trained specifically on Uyghur) and exploring Graph Convolutional Networks (GCN) to model the dependency trees of Uyghur sentences could be the next frontier for this task.
