Research shows how AI is getting better at exams: How universities can respond
When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat.
This article explores how advancements in artificial intelligence are impacting academic assessments, particularly in law programs. When ChatGPT emerged in late 2022, universities faced uncertainty over the extent to which generative AI could navigate assessments and potentially enable academic dishonesty. Despite early claims of AI achieving human-level performance, initial tests revealed limitations in critical legal analysis.
A 2023 experiment conducted by researchers tested AI on an Australian criminal law exam and found that AI was far from replacing human expertise in complex legal tasks. The new study revisits this experiment, examining nine AI models from five providers on two key law subjects at the University of Wollongong. By generating answers to identical exam questions, the researchers found significant improvement in AI's performance compared to 2023.
In criminal law, AI outperformed 82.5% of students with an average score of 76.3%, while in tort law, AI averaged 66%, surpassing 61% of students. Notably, seven AI-generated papers ranked within the top 10% of student performances. The study highlights the shift from AI struggling with critical analysis in 2023 to more competent performance now.
However, the researchers caution that AI is not yet reliable for independent legal analysis. Variability between models and subjects remains, with some models excelling in one area while performing poorly in others. The study also demonstrates that even highly capable AI models can produce flawed outputs, such as poor citations or fabricated authorities.
The researchers emphasize that while AI performance has improved, it does not equate to reliability. Further research is needed to understand AI's consistency across different assessments. The article concludes by proposing three approaches to address these challenges. First, some assessments should remain traditional, monitored, in-person exams to ensure students develop independent knowledge and reasoning.
Second, some assessments should involve collaboration with AI, focusing on the process rather than just the final product. Finally, a relay model, mirroring professional practice, could be adopted, where students submit drafts to AI for refinement and then critique the output using their own critical analysis.
Written by urgent.news from Phys.org's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.