Deciphering Public Sentiment: Machine Learning Insights into New Zealand's COVID-19 Reaction

Twitter Sentiment Analysis using Machine Learning Algorithms for COVID-19 Outbreak in New Zealand

2021-11-06
Oras F. Baker, Jay Liu, Mayur Gosai, Suyog Sitoula
Summary
Problem
Method
Results
Takeaways
Abstract

This paper evaluates public sentiment in New Zealand during the early COVID-19 outbreak by applying machine learning classifiers—Naive Bayes, KNN, CNN, and SVM—to Twitter data. The study utilizes both Python and RapidMiner platforms, identifying SVM and Naive Bayes as the most effective models for opinion mining in this context.

TL;DR

How did New Zealanders truly feel during the dawn of the pandemic? This research utilizes a multi-algorithmic approach—comparing Naive Bayes, KNN, CNN, and SVM—to analyze Twitter data. Utilizing both Python and RapidMiner, the study finds that while the majority of sentiment was positive, the Support Vector Machine (SVM) model stood out as the superior tool for capturing the nuances of public opinion with an accuracy of 86.82%.

Problem & Motivation: The Need for Real-Time Localized Insight

During the early stages of a global health crisis, governmental agencies often struggle to gauge public panic, compliance, and emotion. Traditional surveys are too slow. Social media offers a goldmine of data, yet the challenge lies in filtering noise and accurately classifying "human" emotions from messy, short-form text like Tweets.

The authors identified a gap: while global sentiment was being tracked, New Zealand's specific reaction—characterized by its unique "elimination strategy"—needed a dedicated deep dive using diverse machine learning tools to validate which technical approach works best for this specific demographic.

Methodology: A Multi-Model, Multi-Platform Rig

The researchers didn't just stick to one script. They implemented a comparative framework across two major environments:

  1. Python: Used for its flexibility, specifically utilizing TextBlob and Tweepy for data extraction and CNN implementation.
  2. RapidMiner: A GUI-based data science platform used to validate the results of the traditional algorithms (SVM and Naive Bayes) using a visual workflow.

The CNN Architecture

The study employed a Convolutional Neural Network (CNN) for text classification. Unlike image-based CNNs, this uses 1D convolution where filters slide over word embeddings (n-grams) to capture local contextual features.

CNN Architecture Fig 1: The CNN model architecture utilized for sentiment feature extraction.

Traditional Machine Learning

The team contrasts deep learning with SVM and Naive Bayes. In the SVM model, the data domain is divided using linear and non-linear equations to define high-dimensional boundaries that separate positive from negative sentiments.

Experimental Results: SVM Takes the Crown

The results highlight a classic trade-off in machine learning: Computational Intensity vs. Accuracy.

  • SVM Performance: Achieved the highest accuracy at 86.82%. By generating 48,607 support vectors, it effectively identified "biasing" words that indicate sentiment.
  • Naive Bayes Performance: Faster but less precise, with an accuracy of 74.60%. The researchers noted that Naive Bayes struggled because only a few words in the distribution had high enough weights to clearly distinguish classes.

Performance Comparison Graph Fig 2: Comparison of positive vs. negative classifications across different models.

New Zealand Specific Findings

Interestingly, despite the global panic, the New Zealand subset of data showed a strong "Positive" bias. For instance, the SVM model identified 65 positive results compared to only 9 negative results within the NZ geolocation filter. This suggests that the early lockdown measures were met with more public support than criticism in the region.

Critical Analysis & Conclusion

Takeaway

The research proves that Sentiment Analysis (Opinion Mining) is a viable tool for public health officials. SVM remains a robust baseline for short-text classification, often outperforming simpler probabilistic models like Naive Bayes when given sufficient training data.

Limitations & Future Work

The study is currently limited to a Binary Classification (Positive/Negative). The authors acknowledge that a 3-way sentiment model (including "Neutral") would likely yield even higher accuracy, as many COVID-19 tweets are purely informational (e.g., "Corona", "Virus", "Quarantine") without inherent emotional weight. Future research should look into integrating Transformer-based models like BERT to better handle the linguistic nuances of sarcasm and context which traditional SVMs might miss.

In conclusion, during the "early stage" of the pandemic, New Zealanders on Twitter were surprisingly optimistic—a sentiment captured most effectively through the rigorous mathematical boundaries of Support Vector Machines.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing the performance of BERT and RoBERTa against traditional SVM for localized COVID-19 sentiment analysis on Twitter.
  • Which original papers established the methodology for cross-platform validation between Python Scikit-learn and RapidMiner for NLP tasks?
  • Explore how these sentiment classification techniques have been integrated into real-time public health dashboards for pandemic resource allocation.
Contents
Deciphering Public Sentiment: Machine Learning Insights into New Zealand's COVID-19 Reaction
1. TL;DR
2. Problem & Motivation: The Need for Real-Time Localized Insight
3. Methodology: A Multi-Model, Multi-Platform Rig
3.1. The CNN Architecture
3.2. Traditional Machine Learning
4. Experimental Results: SVM Takes the Crown
4.1. New Zealand Specific Findings
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work