What counts as evidence that AI policies actually work?
The most direct evidence comes from measuring student behavior before and after a policy is in place, or comparing students who are aware of a policy against those who are not. A 2026 study of 168 undergraduate IT students in the Philippines did exactly that: it found that students who were aware of their institution's AI use policy reported significantly lower AI dependency than those who were unaware (a statistically significant difference, p = .026) [1]. Even more telling, when the researchers controlled for other factors like study habits and grades, policy compliance was the strongest predictor of lower AI dependency (β = −0.290, p < .001) — meaning the policy itself, not just being a good student, drove the effect [1]. This is the kind of controlled, quantitative evidence that directly answers the question.
Another study looked at a different outcome: academic integrity. A 2026 survey of 218 sports sciences students in Pakistan found that ChatGPT usage was actually positively correlated with academic integrity (r = .541, p < .001), and ChatGPT use predicted 29.3% of the variance in integrity scores [3]. However, the authors note that this positive relationship depended on the existence of an institutional policy framework — the dimensions of trust and fairness showed only weak associations with ChatGPT use, suggesting they rely on clear institutional rules [3]. This implies that policies don't just prevent cheating; they can channel AI use into honest, productive learning when the rules are clear.
What kind of policy works best — bans, surveillance, or redesign?
The evidence strongly favors policies that focus on assessment redesign and transparency over surveillance and bans. A 2025 analysis of 40 online distance learning institutions created an 'AI-aware Assessment Policy Index' (AAPI) that scored institutions on six dimensions: assessment redesign, disclosure norms, AI literacy, process safeguards, equity safeguards, and surveillance intensity [2]. Institutions with higher AAPI scores — meaning they emphasized redesign and transparency rather than detection and proctoring — had significantly lower AI-detector flag rates (Spearman ρ = −0.42, p < .01) without any evidence of increased misconduct [2]. This is a direct, quantitative link between policy design and real-world outcomes.
A 2024 study of undergraduate business students found that students preferred institutional policies that allowed AI use but set clear boundaries, and that requiring AI use in a structured assignment actually enhanced their learning experience [4]. The study also found that the perception that AI was 'cheating' reduced its usage, which suggests that policies framing AI as a legitimate tool (with rules) are more effective than outright bans [4]. A 2025 analysis of guidelines from the top 50 U.S. universities found that 94% had faculty guidelines emphasizing the importance of course-specific AI policies, and sentiment analysis revealed highly positive attitudes toward AI across all institution types [5]. This consensus among leading universities — that flexible, course-level policies are better than blanket bans — is itself a form of real-world evidence.
A 2024 study of 30 leading global universities identified a common ethical framework: student work must reflect individual knowledge, but instructors should have flexibility in how they use AI in their courses [6]. The typical response blended preventive measures (like modifying assessments) with soft, dialogue-based sanctioning procedures, rather than punitive surveillance [6]. This 'top-down requirement, bottom-up flexibility' model is what the evidence supports.
Where does the evidence conflict or fall short?
The studies agree on the main point — policies work — but they disagree on the mechanism. The IT student study [1] found that policy awareness reduced AI dependency, while the sports sciences study [3] found that AI use actually increased academic integrity. These aren't necessarily contradictory: the IT study measured dependency (overreliance), while the sports study measured integrity (honest use). It's possible that good policies reduce blind overreliance while simultaneously encouraging honest, productive use. But the conflict highlights that 'working' can mean different things — reducing misuse vs. promoting good use — and a policy might do one without the other.
The evidence also has important limitations. The strongest quantitative study [1] was a single-institution survey in the Philippines (n=168), not a randomized experiment. The policy index study [2] explicitly calls itself a 'pilot' with a modest sample of 40 institutions, cautioning against definitive causal claims. The sports sciences study [3] was conducted in Pakistan, where no explicit AI policies existed at the time, so its findings about the positive role of ChatGPT may not generalize to policy-rich environments. And the top-50 U.S. universities study [5] analyzed guidelines, not student outcomes — it tells us what institutions recommend, not whether those recommendations change behavior. Across all six studies, there is no large-scale, multi-institution randomized trial that proves causation. The evidence is consistent and suggestive, but not yet definitive.
About These Sources
This answer is built on 6 peer-reviewed studies — published from 2024 to 2026, 6 from 2024 or later, 1 in Q1–Q2 journals, collectively cited 320 times — selected as the most relevant from 6 studies that passed quality screening, drawn from 61 papers retrieved from a database of over 500 million.
Sources used in this answer
DO INSTITUTIONAL AI POLICIES RELATE TO STUDENT AI DEPENDENCY? EVIDENCE FROM INFORMATION TECHNOLOGY STUDENTS
In a survey of 168 IT students, those aware of their institution's AI policy reported significantly lower AI dependency (t-test p = .026), and policy compliance was a strong predictor of reduced dependency in regression analysis (β = −0.290, p < .001), while study habits and grades were not significant predictors.
Generative AI and Academic Integrity in Online and Distance Learning: A Policy Index and Evidence for Assessment Redesign in the Global South
In a meta-synthesis of 50 policy texts and data from 40 online distance learning institutions, higher scores on the AI-aware Assessment Policy Index (emphasizing redesign and transparency) correlated with lower AI-detector flag rates (Spearman ρ = −0.42, p < .01) without increased misconduct.
The Role of ChatGPT in Enhancing the Academic Integrity of Sports Sciences and Physical Education Students
A survey of 218 sports sciences students in Pakistan found a moderate positive correlation between ChatGPT use and academic integrity (r = .541, p < .001), with ChatGPT use predicting 29.3% of integrity score variance, though trust and fairness dimensions depended on institutional policy frameworks.
An investigation of generative AI in the classroom and its implications for university policy
A mixed-method study of undergraduate business students found that students used GenAI more when they perceived it helped learning and when it was a social norm, and less when perceived as cheating; students preferred policies that allow AI use with clear boundaries.
Investigating the higher education institutions’ guidelines and policies regarding the use of generative AI in teaching, learning, research, and administration
An analysis of guidelines from the top 50 U.S. universities found that 94% had faculty guidelines emphasizing course-specific GenAI policies, with topic modeling revealing four core themes including integration in learning and assessment, and sentiment analysis showing highly positive attitudes toward GenAI.
AI and ethics: Investigating the first policy responses of higher education institutions to the challenge of generative AI
A content analysis of policy documents from 30 leading global universities identified a common ethical framework: student work must reflect individual knowledge, with a blend of preventive assessment modifications and soft, dialogue-based sanctioning procedures rather than punitive surveillance.
