Predicting Academic Performance: Bridging Behavioral Data and Social Influence

Predicting Academic Performance via Semi-supervised Learning with Constructed Campus Social Network

2017-01-01
Huaxiu Yao, Min Nie, Han Su, Hu Xia, Defu Lian
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a framework for predicting academic performance by constructing a campus social network from large-scale behavior data (14M records). The researchers introduce Label Propagation on Multiple Networks (LPMN), a semi-supervised learning algorithm that jointly learns the importance weights of different campus locations to achieve SOTA performance in student grade level classification.

TL;DR

Understanding why some students excel while others struggle often overlooks the social dimension—the "birds of a feather" effect. This paper constructs a hidden campus social network using 14 million smartcard records and introduces LPMN (Label Propagation on Multiple Networks). By learning which campus locations (like the library or teaching buildings) are the best indicators of meaningful friendship, the model significantly improves academic performance prediction without requiring explicit friend lists.

Motivation: The Missing Social Link

Prior research in academic forecasting has focused heavily on individual metrics: test scores, attendance, or smartphone usage patterns. However, academic performance is rarely an isolated variable; it is heavily influenced by one's peer group. The catch? Universities rarely have an accurate, up-to-date map of who is actually friends with whom.

The authors' central insight is that spatiotemporal co-occurrence (being in the same place at the same time) is a proxy for social ties. But not all co-occurrences are equal. Grabbing a quick meal in a crowded canteen might be a coincidence, but consistently studying in the same library nook is likely a sign of a strong social bond.

Methodology: From "Random Encounters" to "Social Ties"

1. Constructing the Network via Shuffling Tests

To differentiate between a random encounter and a real friend, the researchers built a null model. They shuffled the timestamps of student activities 20 times to see how often two people would "accidentally" meet. By setting a threshold (mean + 2σ), they filtered out the noise. For instance, in the canteen, one needs over 110 co-occurrences to be considered a friend, whereas, in the library, only 17 are needed.

2. LPMN: Multi-Network Label Propagation

Instead of treating all locations the same, the authors proposed the LPMN algorithm. It treats each location (Canteen, Library, Classroom, etc.) as a separate layer of a multi-layer graph.

The core optimization target is:

  1. Label Consistency: Students should have academic labels close to the ground truth.
  2. Social Smoothness: Friends (those with high weights ) should have similar performance scores.
  3. Weight Optimization: The model automatically learns which location-based network is the most reliable for prediction.

Model Overview: Co-occurrence Mapping Figure 1: Illustration of student co-occurrence across different campus facilities.

Experiments & Results

The study analyzed 5,388 students and verified the Homophily Phenomenon: friends indeed have closer GPAs than non-friends.

Key Findings:

  • Teaching Buildings are King: Among single-location networks, the "Teaching Building" network was the strongest predictor. This suggests that shared learning environments are the primary drivers of academic similarity.
  • The Power of Fusion: The proposed LPMN (Precision: 0.401) outperformed LP-All (equal weights) and all single-location models.

Performance Comparison Table 3: Accuracy comparison showing LPMN's superiority over baseline methods.

Critical Insight & Conclusion

The beauty of this work lies in its ability to infer latent social structures from "trash" data—ubiquitous, logs of smartcard swipes that are usually discarded.

Takeaway: For educational administrators, this provides a non-intrusive way to identify "at-risk" students. If a student's entire social circle is struggling, the student is statistically likely to follow, even if their current grades are acceptable. Early intervention can then be targeted not just at individuals, but at social clusters.

Limitations: The model currently uses a 4-level discretely ranked GPA to protect privacy, which may lose some granularity. Future work could integrate the temporal evolution of these friendships to see how social circles shift across semesters and how that correlates with academic trajectories.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Graph Neural Networks (GNNs) or Graph Attention Networks (GATs) for student academic performance prediction based on behavioral co-occurrence.
  • Which study first introduced the concept of the "shuffling test" or "null model" for inferring social ties from geographic coincidences, and how has this been adapted for privacy-preserving data mining?
  • Explore research that applies multi-view or multi-layer semi-supervised label propagation to other domains such as urban mobility or workplace productivity tracking.
Contents
Predicting Academic Performance: Bridging Behavioral Data and Social Influence
1. TL;DR
2. Motivation: The Missing Social Link
3. Methodology: From "Random Encounters" to "Social Ties"
3.1. 1. Constructing the Network via Shuffling Tests
3.2. 2. LPMN: Multi-Network Label Propagation
4. Experiments & Results
4.1. Key Findings:
5. Critical Insight & Conclusion