GenAI cannot accurately grade student essays in higher education, study finds
New research from Cardiff University and the University of Melbourne has investigated whether GenAI could mimic human marking when evaluating written student work. The results have been published in the journal Assessment & Evaluation in Higher Education.
A study conducted by researchers from Cardiff University and the University of Melbourne has revealed that current generative artificial intelligence (GenAI) systems are unable to accurately grade student essays in higher education. The research, published in the journal Assessment & Evaluation in Higher Education, tested two popular GenAI platforms, ChatGPT, on marking 50 undergraduate bioscience essays.
The findings indicate significant discrepancies between the marks assigned by the AI and those given by human evaluators. While the overall essay marks were similar, the differences between LLM-assigned and human-assigned marks were substantial when assessing individual marking criteria and on an essay-by-essay basis. The AI system generally assigned higher marks, with differences up to 16.1 points on average and 40 points per essay.
Moreover, LLMs tended to reduce the marks for high-scoring essays and inflate them for low-scoring work, resulting in a systematic compression of scores toward the middle. The study's lead author, Dr. William Kay of Cardiff University's School of Biosciences, emphasized that at present, GenAI is not suitable for reliably assigning grades to students' written work.
While there is interest in exploring the potential of LLMs for objective grading, the research suggests that it is not advisable to rely on them for this purpose. The study also raises ethical concerns about submitting student work to GenAI tools without express consent.
Written by urgent.news from Phys.org's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.