A recent controversy surrounding an artificial intelligence (AI) system's grading for an English essay test has sparked concerns over fairness and accuracy in automated evaluations. The issue came to light when a student took the College Scholastic Ability Test (CSAT) mock exam, specifically the English section, and received an unexpectedly high score for an essay response that was later found to contain several instances of specific phrases or "buzzwords" that are often favored in evaluations. The student expressed frustration upon learning that despite their teacher grading the same essay as a C, the AI system awarded it an A. This discrepancy has raised questions about the reliability of AI in assessing student performance, particularly when it appears to prioritize certain keywords over the logical coherence or overall quality of the response. Further investigation revealed that the AI system seemed to give more weight to the presence of certain "buzzwords" rather than the content or logical structure of the essay. This approach has been criticized for potentially rewarding students who are adept at identifying and incorporating favored phrases into their work, rather than those who demonstrate a deeper understanding or more nuanced expression of the subject matter. The Korea Institute of Curriculum and Evaluation (KICE), the organization responsible for developing and grading the CSAT, has acknowledged the issue and is taking steps to address concerns over the AI grading system. KICE officials have stated that they are working to refine the AI evaluation algorithm to ensure that it assesses responses based on a more balanced and comprehensive set of criteria, rather than relying heavily on specific keywords. The incident highlights the challenges and limitations of using AI in educational assessments, particularly in evaluating subjective or creative tasks such as essay writing. As educators and policymakers continue to integrate AI into various aspects of education, ensuring the fairness, accuracy, and reliability of AI systems in grading and evaluation will be crucial. The goal is to create a system that supports and fairly assesses student learning, without inadvertently encouraging superficial strategies or disadvantaging students who approach problems differently.
South Korea's National Education Commission (National Education Commission) is currently considering the introduction of essay and argumentative writing questions for the nationwide high school entrance exams. However, as the country debates the use of artificial intelligence (AI) for grading, a recent report by the Korean Education Evaluation Institute has highlighted the limitations of AI grading systems.
In the report titled "Research on the Application of Artificial Intelligence Models for Automatic Grading of Written Expression," it was found that AI grading models, when applied to five core subjects such as Korean, social studies, mathematics, science, and technology, display certain shortcomings.
One example provided in the report involved a high school essay graded by a teacher at a C grade, which the AI system classified as the top grade, A. The teacher noted that the essay contained inconsistent tone of address and failure to accurately understand the meaning of the cited sources, resulting in a lower grade. Yet, AI deemed the lengthy essay and the use of specific vocabulary to be high-scoring.
The report also mentioned similar cases in the science subject, where incorrect explanations of principles were recognized as correct due to the presence of specific keywords.
Moreover, the AI grading system showed sensitivity to factors like sentence length, connecting words, and intonation, while human teachers emphasized the accuracy of technical concepts and logical reasoning. The grading committee, led by Ms. Yeom Seok, Minister of Education for AI Strategy, emphasized that AI grading should be used as a supplementary tool to assist teachers in the evaluation process, rather than replacing them entirely.
The AI grading system comprises four stages, with the final stage involving human review of the AI's scores and their justifications. The committee also suggested measures to address security concerns, such as separating student answers and personal information, and ensuring that student answers are not used for external AI learning.
Written by urgent.news from Hankyoreh's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.