A recent controversy surrounding an artificial intelligence (AI) system's grading for an English essay test has sparked concerns over fairness and accuracy in automated evaluations. The issue came to light when a student took the College Scholastic Ability Test (CSAT) mock exam, specifically the English section, and received an unexpectedly high score for an essay response that was later found to contain several instances of specific phrases or "buzzwords" that are often favored in evaluations. The student expressed frustration upon learning that despite their teacher grading the same essay as a C, the AI system awarded it an A. This discrepancy has raised questions about the reliability of AI in assessing student performance, particularly when it appears to prioritize certain keywords over the logical coherence or overall quality of the response. Further investigation revealed that the AI system seemed to give more weight to the presence of certain "buzzwords" rather than the content or logical structure of the essay. This approach has been criticized for potentially rewarding students who are adept at identifying and incorporating favored phrases into their work, rather than those who demonstrate a deeper understanding or more nuanced expression of the subject matter. The Korea Institute of Curriculum and Evaluation (KICE), the organization responsible for developing and grading the CSAT, has acknowledged the issue and is taking steps to address concerns over the AI grading system. KICE officials have stated that they are working to refine the AI evaluation algorithm to ensure that it assesses responses based on a more balanced and comprehensive set of criteria, rather than relying heavily on specific keywords. The incident highlights the challenges and limitations of using AI in educational assessments, particularly in evaluating subjective or creative tasks such as essay writing. As educators and policymakers continue to integrate AI into various aspects of education, ensuring the fairness, accuracy, and reliability of AI systems in grading and evaluation will be crucial. The goal is to create a system that supports and fairly assesses student learning, without inadvertently encouraging superficial strategies or disadvantaging students who approach problems differently.
We haven't written up this one. Hankyoreh has the full story — the link below goes straight to it.