Decoding Innovation: How Machine Learning Identifies Creative Sparks in Corporate Social Networks
Investigating Linguistic Indicators of Generative Content in Enterprise Social Media
This study investigates linguistic indicators of "generative content"—group interactions that foster innovation—within Enterprise Social Media (ESM). By applying machine learning classifiers like Random Forest to corporate communication data, the authors identify specific keywords and patterns that distinguish creative problem-solving from routine social exchanges.
TL;DR
Is your company's internal social network actually driving innovation, or is it just digital water-cooler talk? This preliminary study uses machine learning to identify generative interactions—conversations that create new ideas—within Enterprise Social Media (ESM). Using a Random Forest classifier, researchers achieved 76% accuracy in distinguishing generative content, revealing the specific "vocabulary of innovation" used by distributed teams.
The "Visibility" Opportunity
The heart of modern organizational agility lies in collaboration. Enterprise Social Media (ESM) platforms are no longer just optional tools; with over $100 billion invested globally, they are the central nervous system of team interaction.
The authors argue that ESM provides a unique "visibility affordance." Unlike private emails, ESM threads are persistent and traceable. This creates a goldmine for HCI (Human-Computer Interaction) researchers to study generativity—the process of creating, originating, or producing new knowledge—without the bias inherent in self-reported surveys.
Methodology: Detecting the "Aha!" Moment
The researchers focused on three types of generative conceptual changes:
- Expansion: Extending an existing concept to a new situation.
- Reframing: Deconstructing and reconstructing concepts to challenge the status quo.
- Combination: Merging existing ideas into something entirely new.
Using a dataset from a multinational organization with 11,000+ employees, the team labeled message threads and processed them through a machine learning pipeline involving lemmatization and TF-IDF vectorization.

Experimental Battleground: Random Forest vs. The Rest
The study compared five different models: Random Forest, AdaBoost, Naïve Bayes, SVM, and Logistic Regression. Random Forest emerged as the clear winner, particularly in its ability to handle the nuances of generative text.
Performance Metrics
| Model | Accuracy | AUC | F1-Score |
|---|---|---|---|
| Random Forest | 0.76 | 0.80 | 0.83 |
| Adaptive Boosting | 0.71 | 0.70 | 0.81 |
| Naïve Bayes | 0.44 | 0.59 | 0.53 |

Linguistic Indicators: The Vocabulary of Value
What does innovation "sound" like? The model extracted the top 20 terms that distinguish generative content. Interestingly, words like "Value," "New," "Different," "Need," and "Project" were core weights in the model. This suggests that generative interactions are often anchored in outcome-oriented and comparative language.

Critical Insight: Beyond the Binary
While this study effectively separated generative from non-generative content (finding that 28% of ESM interactions fit the generative mirror), the next frontier is distinguishing between the types of generativity.
Reframing is disruptive and challenges the status quo, while Expansion is incremental. For a manager, being able to detect which teams are merely refining ideas versus those that are fundamentally re-evaluating business models is the holy grail of innovation analytics.
Conclusion & Future Outlook
This work serves as a vital proof-of-concept. It proves that we can "listen" to the pulse of corporate innovation through automated text classification.
Practical Takeaway: In the future, ESM platforms might include dashboards for managers to see "generativity scores" for different projects, allowing for real-time nudges to shift conversations from simple information sharing to high-value creative synthesis. However, the reliance on a 1% subsample suggests that further research with larger, more diverse datasets (and perhaps modern LLM-based embeddings) is needed to refine these linguistic predictors.
