MAS Corpus: Bridging the Gap Between Spanish Linguistics and Marketing Intelligence

MAS: A Corpus of Tweets for Marketing in Spanish

2018-01-01
María Navas-Loro, Víctor Rodríguez-Doncel, Idafen Santana-Pérez, Alba Fernández-Izquierdo, Alberto Sánchez
Summary
Problem
Method
Results
Takeaways
Abstract

MAS (Marketing in Spanish) is a novel expert-annotated corpus of 3,763 Spanish tweets designed for multidimensional marketing analysis. It moves beyond simple polarity by labeling data across three axes: fine-grained emotions, the "4Ps" Marketing Mix, and the stages of the Purchase Funnel.

TL;DR

The MAS Corpus is a specialized dataset of Spanish tweets meticulously annotated to serve the marketing industry. Unlike standard sentiment datasets, it categorizes social media posts according to Emotions, Marketing Mix (4Ps), and the Purchase Funnel. By releasing this as Linked Data, the researchers provide a robust foundation for building AI that understands not just how consumers feel, but where they are in their buying journey.

Problem & Motivation: Beyond "Positive" vs. "Negative"

In the world of professional marketing, knowing a tweet is "positive" is rarely enough. A brand manager needs to know if that positivity is coming from a loyal customer sharing their Post-purchase experience or a potential lead in the Evaluation phase.

Existing Spanish NLP resources have historically suffered from:

  • Simplicity: Focusing on polarity rather than complex emotions (e.g., distinguishing "Trust" from "Joy").
  • Lack of Context: Failing to link sentiments to specific business facets like Price or Distribution (Place).
  • Data Volatility: Twitter's API policies often make old datasets unreachable; the authors solve this by providing a structured, machine-readable RDF format.

Methodology: The Consumer Lens

The MAS project utilizes a sophisticated three-dimensional tagging criteria derived from established marketing theories.

1. The Marketing Mix (The 4Ps)

The corpus maps tweets to the classic McCarthy framework:

  • Product: Features, taste, or quality.
  • Price: Valorization, discounts, or affordability.
  • Promotion: Ads, sponsorships, and campaigns.
  • Place: Points of sale, customer service, and availability.

2. The Purchase Funnel (Customer Journey)

The experts labeled tweets based on the author's stage in the acquisition process: Awareness → Evaluation → Purchase → Post-purchase. This allows for "timeline" analysis of consumer behavior.

3. Emotion Taxonomy

The study uses a deep hierarchy of emotions (Table 2), mapping secondary feelings like "Ecstasy" to "Love" or "Anxiety" to "Fear," providing a much richer affective profile than standard SA.

Corpus Dimensions and Tags Figure 1: Overview of the annotation dimensions used in the MAS Corpus.

Experimental Insights & Results

The researchers didn't just label data; they analyzed sector-specific nuances. Their findings highlight why domain-specific training is vital:

  • Sector Variance: In the Banking sector, "Fear" is a dominant emotion (15.00%), likely tied to financial security concerns, whereas it is virtually non-existent in the Food sector.
  • Linguistic Ambiguity: The authors noted that "I like Heineken" (Food/Beverage) almost always implies a Post-purchase experience, whereas "I like BMWs" (Automotive) is often just an Evaluation or generic admiration, requiring different tagging logic.

Sector-Based Statistics Table 6: Distribution of emotions across different sectors, showing specific emotional signatures for industries like Banking and Telecom.

Critical Analysis & Conclusion

Takeaway

The MAS Corpus is a significant step forward for Spanish MarTech. By integrating marketing theory directly into the NLP pipeline, it transforms social media monitoring from a "mood ring" into a strategic tool for identifying friction points in the sales funnel.

Limitations & Future Work

One inherent challenge is the imbalance across categories—some sectors (like Telecom) have very high "NC2" (Not Classified) rates (91.63%), indicating that many tweets remain noise or purely corporate interaction. Future iterations could benefit from Active Learning to specifically target and enrich these underrepresented categories. Furthermore, as the corpus grows, it will serve as a vital benchmark for training large language models (LLMs) to perform zero-shot marketing analysis in Spanish.


Keywords: Spanish NLP, Sentiment Analysis, Marketing Mix, Purchase Funnel, Linked Data, Twitter Corpus.

Find Similar Papers

Try Our Examples

  • Which recent Spanish NLP datasets or models have improved upon the MAS corpus for detecting Purchase Intent specifically in social media?
  • Find the original paper by McCarthy (1978) on the "4Ps" of Marketing Mix and explore how its traditional definitions have been adapted for modern Aspect-Based Sentiment Analysis (ABSA).
  • Are there any studies that apply the Purchase Funnel classification framework to multi-modal data, such as Instagram posts or TikTok videos, for brand perception analysis?
Contents
MAS Corpus: Bridging the Gap Between Spanish Linguistics and Marketing Intelligence
1. TL;DR
2. Problem & Motivation: Beyond "Positive" vs. "Negative"
3. Methodology: The Consumer Lens
3.1. 1. The Marketing Mix (The 4Ps)
3.2. 2. The Purchase Funnel (Customer Journey)
3.3. 3. Emotion Taxonomy
4. Experimental Insights & Results
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work