WisPaper
WisPaper
Search
Assistant
Pricing
TrueCite

Are the privacy and fairness risks of learning analytics systems being underestimated?

Evidence shows privacy and fairness risks in learning analytics are indeed underestimated, with students' concerns often overlooked and data sovereignty inadequately addressed.

Direct answer

Yes, the evidence strongly suggests that privacy and fairness risks of learning analytics systems are being underestimated. Across the studies reviewed, students' privacy concerns are a significant predictor of their willingness to share data [2], yet most institutional frameworks treat student data as a resource without considering student data ownership or rights [1]. Furthermore, privacy risks like re-identification are quantifiable and often higher than assumed [6], and technological fixes alone are insufficient because privacy is also a social and ethical issue [3]. The scale is substantial: in one study of 132 students, perceived privacy risk strongly predicted concerns, which in turn led to non-disclosure behaviors [2], indicating that underestimating these risks can undermine the very data collection learning analytics depends on.

6sources cited

This article was generated with WisPaper-powered search and paper analysis.

How big is the gap between perceived and actual privacy risks?

The gap is large and measurable. A 2022 study of 132 students across three Swedish universities found that students' perceived privacy risk was a 'firm predictor' of their privacy concerns, and those concerns directly led to non-self-disclosure behaviors—meaning students withheld data when they felt at risk [2]. This is not a minor effect: the model explained high variance in both privacy concerns and trusting beliefs, showing that privacy risk perception is a central driver of student behavior. If institutions underestimate this, they risk building analytics on incomplete or biased data.

On the technical side, a 2022 study using Markov models on real education datasets showed that re-identification risks—the chance of linking anonymized data back to a specific student—are often higher than traditional methods suggest, because they account for correlations between attributes and event-level data (e.g., multiple records for the same student) [6]. The authors explicitly note that most existing risk quantification methods assume an adversary with limited prior knowledge, which underestimates real-world threats. Their model provides a 'worst-case' risk assessment, which is precisely what institutions should be using but often do not.

Why do institutions systematically underestimate these risks?

A major reason is that most institutional frameworks for learning analytics do not view student data as part of a broader data ecology with dynamic power relations and interdependencies. A 2023 analysis of several leading frameworks found that while they understand learning analytics as a data ecosystem, there is 'very little evidence of a broader data ecological understanding' [1]. Crucially, the vast majority treat student data as a 'valuable resource' without considering student data ownership or their rights to self-determination [1]. This blind spot means privacy and fairness are not built into the system from the start.

Another reason is the over-reliance on technological solutions. A 2022 conceptual review mapped the assumptions behind privacy-enhancing technologies (PETs) and found they often rest on flawed premises—such as that individuals have full control over their data as a commodity, or that consent alone is sufficient [3]. The authors argue that privacy is not just a technical problem but a social and ontological one, and that approaches like 'contextual integrity' or group privacy are needed. In practice, this means that even when institutions deploy PETs, they may still fail to address fairness risks if they ignore the social context of data collection.

The problem is global but uneven. A 2022 scoping review of privacy regulations across 32 African countries found that while many have national and regional frameworks, African higher education institutions often use learning platforms from outside the continent, creating a 'data frontier' to be exploited [4]. The authors note that this raises unique fairness risks—such as data sovereignty and cross-border data transfer—that are easily underestimated by institutions in both the Global North and South.

Can privacy and fairness be preserved without killing the analytics?

Yes, but it requires intentional design. A 2024 study tested privacy-preserving mechanisms on a large-scale educational dataset and found that it is possible to preserve both privacy and utility—the usefulness of data for analytics [5]. The researchers applied several anonymization techniques and then evaluated how well common learning analytics models performed on the protected data. Their results provide a 'benchmark of utility loss' and prove that privacy protection can be an integral part of LA design, not an afterthought. However, the study also warns that this compatibility is not automatic; it requires careful calibration.

The same 2022 study on re-identification risk offers a practical workflow for data custodians to evaluate worst-case risks before releasing data [6]. This empowers institutions to make informed decisions about mitigation (e.g., data perturbation) rather than assuming the risk is low. Combined with the student-centered approach from the SPICE model [2]—which shows that enhancing perceived privacy control and reducing perceived risk builds trust—these tools suggest that underestimation is a choice, not a necessity. The key is to move from a 'data as resource' mindset to one that respects student data sovereignty [1] and treats privacy as a social, not just technical, challenge [3].

About These Sources

This answer is built on 6 peer-reviewed studies — published from 2022 to 2024, 1 from 2024 or later, 5 in Q1 journals, collectively cited 197 times — selected as the most relevant from 6 studies that passed quality screening, drawn from 77 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Learning analytics as data ecology: a tentative proposal

Analyzing several learning analytics frameworks, this 2023 study found that most treat student data as a valuable resource without considering student data ownership or rights to self-determination, and show little evidence of understanding learning analytics as part of a broader data ecology with dynamic power relations.

2

Students' privacy concerns in learning analytics: Model development

In a survey of 132 students across three Swedish universities, this 2022 study validated the SPICE model, showing that perceived privacy risk is a strong predictor of privacy concerns, which in turn predict non-self-disclosure behaviors and trust in institutions.

3

The answer is (not only) technological: Considering student data privacy in learning analytics

This 2022 conceptual review argues that technological solutions to student data privacy (e.g., PETs) are insufficient because they rely on flawed assumptions (e.g., data as commodity, individual control), and that privacy must be understood as a social, ontological, and contextual issue.

4

Data privacy on the African continent: Opportunities, challenges and implications for learning analytics

A 2022 scoping review of privacy regulations across 32 African countries found that while many have national frameworks, African higher education institutions often use platforms from outside the continent, creating risks around data sovereignty and cross-border data transfer that are easily underestimated.

5

Preserving Both Privacy and Utility in Learning Analytics

This 2024 study applied privacy-preserving mechanisms to a large-scale educational dataset and evaluated utility loss in common learning analytics models, demonstrating that privacy and utility can be compatible if protection is designed into the analytics process from the start.

6

Privacy risk quantification in education data using Markov model

This 2022 study proposed a Markov model to quantify re-identification risks in education datasets, accounting for correlations between attributes and event-level data, and found that worst-case risks are often higher than traditional methods assume, providing a workflow for data custodians to evaluate and mitigate these risks.