Why traditional performance measures fall short when students use AI
When students can use AI to generate essays, solve problems, or create content, grading solely on the final product no longer measures student ability—it measures the AI's ability. A 2022 experiment with 20 undergraduate students found that student-AI collaboration significantly boosted creativity in content and expressivity in expression on a public advertisement drawing task [4]. This means if you grade only the final drawing, you're rewarding the AI's contribution, not the student's skill. The same logic applies to written work: if ChatGPT can write a passing essay, then an 'A' on that essay tells you nothing about the student's understanding. The studies here converge on the idea that assessment must shift from what students produce to how they produce it—their process of prompting, evaluating, and refining AI outputs.
What should be measured instead: process, ethics, and critical thinking
The research points to three new dimensions of performance that AI policies should prioritize. First, ethical judgment and disclosure: a 2026 study of 864 sport-science students developed a validated scale that includes 'Ethics & Disclosure' and 'Trust & Verification' as core dimensions of AI literacy [5]. This means assessments should check whether students properly cite AI use and verify AI-generated facts. Second, critical thinking and collaboration: a 2025 survey of 982 students and 76 faculty found that students and faculty agreed that AI use could affect learning outcomes differently across groups, with male STEM students showing more positive attitudes than female non-STEM students [2]. This suggests that assessments need to be designed to capture how each student critically engages with AI, not just whether they used it. Third, emotional and strategic use: a 2026 mixed-methods study found that students' affective attitudes toward AI were the strongest predictor of their perceived learning, and that AI offered emotional support and boosted confidence [3]. So measuring performance should include how students strategically use AI for academic support while maintaining their own learning agency.
Practical changes to assessment design that AI policies demand
The studies offer concrete recommendations for redesigning assessments. The 2026 study on AI-integrated assessment proposes a 'two-lane model' where assignments are designed either to be completed with AI (lane one) or without AI (lane two), and each lane has different grading criteria [3]. For lane one, grading focuses on the quality of the student-AI interaction and the student's critical evaluation of AI outputs. For lane two, traditional measures still apply but must be administered in controlled settings. Additionally, a 2023 survey of 265 medical students found that while only 20.4% used ChatGPT for written assessments in medical school, 63.4% planned to use it in residency for research and exam prep [1]. This gap between current use and future intent highlights the urgency of updating policies now. The authors of that study explicitly call for 'structured curricula and formal policies and guidelines' to prepare learners [1]. Across all five papers, the consistent message is that policies must be paired with new assessment methods that measure AI literacy, ethical reasoning, and the ability to collaborate with AI—not just the final answer.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2022 to 2026, 3 from 2024 or later, 4 in Q1 journals, collectively cited 178 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 63 papers retrieved from a database of over 500 million.
Sources used in this answer
Medical Student Experiences and Perceptions of ChatGPT and Artificial Intelligence: Cross-Sectional Study
A cross-sectional survey of 265 medical students found that only 20.4% used ChatGPT for written assessments, but 63.4% planned to use it in residency, highlighting a gap between current use and future need that demands updated policies and assessment methods.
Examining Faculty and Student Perceptions of Generative AI in University Courses
A survey of 982 students and 76 faculty at a US university found that students and faculty had similar attitudes toward generative AI, but significant differences existed between male STEM and female non-STEM students, suggesting assessments must account for demographic diversity in AI use.
Using Student Perceptions of AI to Inform Academic Policy and Assessment Practices
A mixed-methods study found that students' affective attitudes toward AI were the strongest predictor of perceived learning, and that students wanted institutional guidance; the study recommends a 'two-lane' assessment model that separates AI-assisted from non-AI tasks.
Are Two Heads Better Than One?: The Effect of Student-AI Collaboration on Students' Learning Task Performance
A repeated-measure experiment with 20 undergraduates found that student-AI collaboration significantly improved creativity in content and expressivity in expression on a drawing task, showing that AI can enhance performance in ways traditional grading may not capture.
Development and validation of an AI use scale for sport and exercise science students.
A validation study of 864 sport-science students developed a 14-item AI-use scale covering Awareness, Ethics & Disclosure, Trust & Verification, and Course Expectations, providing a tool to measure the dimensions that should be assessed under new AI policies.
