Opinion Mining for Romanian: Strategies for Navigating Low-Resource Social Media Data
Opinion mining for social media and news items in Romanian
This paper explores opinion mining for the Romanian language using a manually annotated corpus from social media and news. It evaluates several machine learning strategies, primarily focusing on Support Vector Machines (SVM) and n-gram probability models, to achieve high-accuracy sentiment classification in a low-resource linguistic context.
TL;DR
Analyzing sentiment in Romanian is significantly more challenging than in English due to a lack of specialized tools and the "noisy" nature of online text (missing diacritics and informal slang). This paper benchmarks three approaches—Bag-of-Words, Affective Scoring, and N-gram Probabilities—finding that Trigram Probabilities (88.44% accuracy) outperform more complex linguistic Parsing methods in real-world scenarios.
Positioning: This work serves as a foundational experimental study for Romanian sentiment analysis, moving beyond simple English-to-Romanian translation toward specialized local classifiers.
The Challenge: Why English Solutions Don't "Translate" to Romanian
Most sentiment analysis tools are optimized for English, taking advantage of massive datasets and precision POS taggers. For Romanian, researchers face three "walls":
- Resource Scarcity: Lack of reliable sentiment lexicons (like SentiWordNet) specifically tuned for Romanian nuances.
- The "Diacritics" Gap: Internet users in Romania often omit diacritics (e.g., writing pătrunjel as patrunjel), which breaks standard NLP pipelines.
- Noisy Contexts: Social media (Twitter, blogs) mixes facts and opinions, often utilizing sarcasm or brand-specific jargon that confuses general-purpose models.
Methodology: Three Paths to Sentiment Extraction
The authors didn't just stick to one method; they compared three distinct architectural philosophies:
1. Enhanced Bag-of-Words (BoW)
Using the WEKA framework, they built a pipeline that includes a POS filter. Crucially, they utilized a diacritics restoration service before sending text to the RACAI web service for lemmatization. This ensures that different forms of the same word (e.g., various verb conjugations) are treated as a single feature.

2. Affective Scores & Dependency Parsing
This was the most "academic" approach. It involved translating English word scores (SentiWordNet) into Romanian and using a Functional Dependency Grammar (FDG) parser to link sentiments to specific "target entities" (e.g., a specific brand name).
3. N-gram Probabilities
Instead of relying on dictionaries, this model calculates the conditional probability of an n-gram (unigram, bigram, or trigram) appearing in a positive versus a negative document. The resulting score ranges from -1 (strongly negative) to +1 (strongly positive).
Experiments & Results: Simplicity Wins
The researchers tested these methods on a dataset provided by ZeList, covering seven different entities (brands/companies).
Key Findings:
- Trigram Probabilities achieved the highest accuracy (88.44%).
- BoW + POS filtering was a close second at 81.31%.
- The Dependency Parsing approach failed significantly, yielding only 52.18%.
The failure of dependency parsing is a crucial insight: current Romanian parsers struggle with the informal structure of social media comments. When the parser fails, the sentiment-to-entity link breaks.
Performance Comparison Table

Critical Insight & Future Outlook
This paper highlights a common pitfall in NLP: Theoretical complexity does not always equate to practical performance. While dependency parsing is elegant, its sensitivity to grammar makes it brittle for the "wild West" of social media.
Takeaways for Practitioners:
- If you are working with a low-resource language, start with N-gram probability models; they are surprisingly resilient to noise.
- Always include a diacritics restoration step for Romanian text to maintain feature consistency.
- Future Work: The authors suggest expanding affective scores beyond just adjectives and improving the entity-linkage algorithms to handle the informal Romanian vernacular better.
Conclusion: While Romanian sentiment analysis is difficult, statistical methods like n-grams provide a robust bridge while we wait for more sophisticated, high-accuracy Romanian linguistic tools (such as native Romanian LLMs) to mature.
