Deciphering Public Sentiment: Machine Learning Insights into New Zealand's COVID-19 Reaction
Twitter Sentiment Analysis using Machine Learning Algorithms for COVID-19 Outbreak in New Zealand
This paper evaluates public sentiment in New Zealand during the early COVID-19 outbreak by applying machine learning classifiers—Naive Bayes, KNN, CNN, and SVM—to Twitter data. The study utilizes both Python and RapidMiner platforms, identifying SVM and Naive Bayes as the most effective models for opinion mining in this context.
TL;DR
How did New Zealanders truly feel during the dawn of the pandemic? This research utilizes a multi-algorithmic approach—comparing Naive Bayes, KNN, CNN, and SVM—to analyze Twitter data. Utilizing both Python and RapidMiner, the study finds that while the majority of sentiment was positive, the Support Vector Machine (SVM) model stood out as the superior tool for capturing the nuances of public opinion with an accuracy of 86.82%.
Problem & Motivation: The Need for Real-Time Localized Insight
During the early stages of a global health crisis, governmental agencies often struggle to gauge public panic, compliance, and emotion. Traditional surveys are too slow. Social media offers a goldmine of data, yet the challenge lies in filtering noise and accurately classifying "human" emotions from messy, short-form text like Tweets.
The authors identified a gap: while global sentiment was being tracked, New Zealand's specific reaction—characterized by its unique "elimination strategy"—needed a dedicated deep dive using diverse machine learning tools to validate which technical approach works best for this specific demographic.
Methodology: A Multi-Model, Multi-Platform Rig
The researchers didn't just stick to one script. They implemented a comparative framework across two major environments:
- Python: Used for its flexibility, specifically utilizing
TextBlobandTweepyfor data extraction and CNN implementation. - RapidMiner: A GUI-based data science platform used to validate the results of the traditional algorithms (SVM and Naive Bayes) using a visual workflow.
The CNN Architecture
The study employed a Convolutional Neural Network (CNN) for text classification. Unlike image-based CNNs, this uses 1D convolution where filters slide over word embeddings (n-grams) to capture local contextual features.
Fig 1: The CNN model architecture utilized for sentiment feature extraction.
Traditional Machine Learning
The team contrasts deep learning with SVM and Naive Bayes. In the SVM model, the data domain is divided using linear and non-linear equations to define high-dimensional boundaries that separate positive from negative sentiments.
Experimental Results: SVM Takes the Crown
The results highlight a classic trade-off in machine learning: Computational Intensity vs. Accuracy.
- SVM Performance: Achieved the highest accuracy at 86.82%. By generating 48,607 support vectors, it effectively identified "biasing" words that indicate sentiment.
- Naive Bayes Performance: Faster but less precise, with an accuracy of 74.60%. The researchers noted that Naive Bayes struggled because only a few words in the distribution had high enough weights to clearly distinguish classes.
Fig 2: Comparison of positive vs. negative classifications across different models.
New Zealand Specific Findings
Interestingly, despite the global panic, the New Zealand subset of data showed a strong "Positive" bias. For instance, the SVM model identified 65 positive results compared to only 9 negative results within the NZ geolocation filter. This suggests that the early lockdown measures were met with more public support than criticism in the region.
Critical Analysis & Conclusion
Takeaway
The research proves that Sentiment Analysis (Opinion Mining) is a viable tool for public health officials. SVM remains a robust baseline for short-text classification, often outperforming simpler probabilistic models like Naive Bayes when given sufficient training data.
Limitations & Future Work
The study is currently limited to a Binary Classification (Positive/Negative). The authors acknowledge that a 3-way sentiment model (including "Neutral") would likely yield even higher accuracy, as many COVID-19 tweets are purely informational (e.g., "Corona", "Virus", "Quarantine") without inherent emotional weight. Future research should look into integrating Transformer-based models like BERT to better handle the linguistic nuances of sarcasm and context which traditional SVMs might miss.
In conclusion, during the "early stage" of the pandemic, New Zealanders on Twitter were surprisingly optimistic—a sentiment captured most effectively through the rigorous mathematical boundaries of Support Vector Machines.
