Why traditional performance metrics (like test scores) become unreliable with AI
When AI can generate sophisticated answers, measuring performance by counting correct answers on multiple-choice tests no longer tells you what a person knows. A 2023 study of ChatGPT on ophthalmology board exam questions found the AI answered only 46% correctly in January 2023 and 58% a month later [1]. That means a student could get a passing score by letting the AI answer, but the score would reflect the AI's knowledge, not the student's. The same study found ChatGPT selected the same answer as human trainees only 44% of the time, showing the AI's reasoning often diverges from human clinical judgment [1]. So if you keep using traditional recall-based tests, you can't tell whether the student or the AI deserves the credit.
This problem is not limited to medicine. A 2026 conceptual analysis of medical education globally argues that traditional assessment models—built on individual authorship and knowledge recall—are 'increasingly strained' by generative AI that can produce convincing clinical narratives and analyses [4]. The authors call for a fundamental redesign of assessments to maintain trust and fairness. In other words, the old yardstick no longer measures what you think it measures.
What should replace old metrics: process, collaboration, and ethical reasoning
The emerging consensus across these studies is that performance should be measured by how well a person works with AI, not by whether they can beat it at recall. A 2026 survey of 420 students, instructors, and administrators in higher education found that stakeholders conditionally accept AI-based assessment tools—they value efficiency and personalized feedback, but they demand transparency about how AI grades, fairness in algorithms, and human oversight [2]. The strongest predictor of acceptance was perceived usefulness (a 0.42 increase in acceptance per unit of usefulness), followed by perceived fairness (0.29), while privacy concerns reduced acceptance (−0.31) [2]. This means new performance metrics must include things like: Did the student critically evaluate the AI's output? Did they use the tool ethically? Can they explain their reasoning?
A 2025 study on AI-powered graphic design tools reinforces this shift. It found that when students co-create with AI—using generative models and adaptive feedback systems—their creativity, engagement, and conceptual clarity improve compared to traditional teaching [3]. The authors argue that assessment should measure 'fluency, originality, and visual coherence' in the human-AI collaboration, not just the final product. Similarly, the medical education analysis advocates for 'AI literacy matched with professionalism' as a core competency to be assessed [4]. Across these fields, the new metrics are about process, judgment, and collaboration—not just the right answer.
The catch: fairness and integrity are harder to measure—and easy to get wrong
Redesigning assessment doesn't automatically make it fair or trustworthy. A 2022 study of AI practitioners at three technology companies found that when they tried to evaluate AI systems for fairness by breaking down performance by demographic groups, they struggled to choose the right metrics, identify which groups mattered, and collect enough data [5]. Business pressures often prioritized customers over marginalized groups, and practitioners lacked engagement with domain experts [5]. This is a direct warning for anyone redesigning assessments: if you don't build fairness into the metrics from the start, you risk creating a system that looks objective but actually amplifies bias.
The higher education survey found that STEM students trusted automated grading more than humanities students did (a statistically significant difference across disciplines) [2]. That means a one-size-fits-all AI assessment could be perceived as unfair by some groups, even if the algorithm is technically accurate. The medical education analysis adds that assessment redesign must include 'cohesive governance' and teacher development to maintain professional standards [4]. In short, new performance metrics need to be transparent, explainable, and tailored to the context—or they will fail the fairness test.
About These Sources
This answer is built on 5 peer-reviewed studies — published from 2022 to 2026, 3 from 2024 or later, 2 in Q1 journals, collectively cited 421 times — selected as the most relevant from 5 studies that passed quality screening, drawn from 54 papers retrieved from a database of over 500 million.
Sources used in this answer
Performance of an Artificial Intelligence Chatbot in Ophthalmic Knowledge Assessment
ChatGPT answered only 46% of 125 ophthalmology board exam questions correctly in January 2023 (rising to 58% in February), and matched human trainee answers only 44% of the time, showing AI's limitations on recall-based tests.
Perceptions of AI-Based Assessment Tools in Higher Education
In a mixed-methods study of 420 higher education stakeholders, acceptance of AI assessment was predicted by perceived usefulness (β=0.42) and fairness (β=0.29), while privacy concerns reduced acceptance (β=−0.31); students were more accepting than instructors.
AI-POWERED GRAPHIC DESIGN TOOLS: A PARADIGM SHIFT IN ART CURRICULUM
AI-powered graphic design tools improved student creativity, engagement, and conceptual clarity compared to traditional teaching, suggesting assessment should measure human-AI collaboration and creative process.
Artificial Intelligence, Assessment Integrity, and Professionalism in Medical Education: Global Disruption and Lessons from the Gulf Cooperation Council Region
Generative AI disrupts traditional medical assessment models based on recall and individual authorship, requiring redesigned assessments that integrate AI literacy, professionalism, and ethical accountability.
Assessing the Fairness of AI Systems: AI Practitioners' Processes, Challenges, and Needs for Support
AI practitioners at three tech companies faced challenges in designing fair disaggregated evaluations, including choosing metrics, identifying relevant demographic groups, and overcoming business pressures that prioritize customers over marginalized groups.
