Avocado: Bridging the Engagement Gap in Adaptive Vocabulary Learning
An adaptive vocabulary learning application through modeling learner's linguistic proficiency and interests
This paper presents "Avocado," an adaptive vocabulary learning application that personalizes English news article recommendations based on an individual's linguistic proficiency and topical interests. The system achieves personalization by integrating linguistic modeling with social media (Facebook) data and utilizing a dual-category word difficulty metric (general vs. domain-specific).
TL;DR
Most vocabulary apps treat learners like blank slates, feeding them generic word lists. Avocado, a research project from KAIST, flips this script by building a digital twin of the learner. By scraping your Facebook interests and tracking which words you find difficult in real-time news articles, it curates a personalized curriculum that feels less like "studying" and more like "browsing."
The Personalized Learning Paradox
Why do we find it easier to read a complex technical manual in our field than a "simple" children's story in a foreign language? The answer lies in Domain-Specific Exposure.
Prior methods for measuring word difficulty relied heavily on global frequency (e.g., how often "apple" appears vs. "quantum"). However, these metrics ignore the personal context. The researchers behind Avocado identified that a "one-size-fits-all" approach leads to cognitive overload or boredom, both of which are "app-killers" in the EdTech space.
The Core Engine: Dual-Faceted Proficiency Modeling
Avocado’s innovation lies in how it defines "difficulty" and "interest."
1. The "TD-DF" Metric for Difficulty
Instead of just counting words, the authors use a composite score:
- Frequency-based: Inversely proportional to Term Frequency (TF) and Document Frequency (DF).
- Structural-based: Linear combinations of character length, syllables, and the presence of rare character combinations.
- Domain Specificity: Every word is assigned a "general difficulty" and "domain-specific difficulties," allowing the system to recognize that a programmer might know "memoization" even if they are at an intermediate English level.
2. Social Interest Mapping
By using the Facebook Graph API, Avocado fetches a user’s "liked" pages. It doesn't just look at the titles; it crawls recent posts from those pages, converts them into Word Embedding Vectors (Word2Vec), and calculates a mean "Interest Vector."
Fig 1: Avocado personalizes recommendations by matching Article Vectors with User Interest Vectors.
From Vectors to Vocabulary: The Recommendation Loop
The recommendation process is a sophisticated filter:
- Selection: The system scrapes over 50 news sources.
- Cosine Similarity: It identifies articles whose "Article Vector" aligns with the user's "Interest Vector."
- Difficulty Match: From the relevant articles, it picks those closest to the user’s current vocabulary level.
- Active Highlighting: Words exceeding the user's proficiency are automatically highlighted (see Fig 1).
- Dynamic Update: As the user clicks a word to see its definition, the system updates the student’s level, creating a living model of their progress.
Fig 2: LDA (Latent Dirichlet Allocation) is used to extract and display key topics for each recommended article.
Critical Analysis & Future Outlook
While Avocado represents a significant step toward Context-Aware Learning, it faces challenges common in the 2020s:
- Data Privacy: Relying on Facebook APIs is increasingly difficult due to privacy regulations (like GDPR) and API restrictions.
- The Polysemy Problem: As the authors note, the word "make" is frequent but difficult due to its dozens of meanings. Current frequency-based metrics struggle with semantic ambiguity—a problem now being solved by contextual embeddings like BERT or GPT.
- Proper Noun Noise: Names like "Twitter" are often flagged as "difficult" simply because they are rare in older training corpora.
Takeaway: The future of language learning isn't in better flashcards; it's in better content-to-user matching. Avocado proves that by modeling both the what (interests) and the how (proficiency), we can transform language acquisition from a chore into a secondary effect of information consumption.
