DistantNews
Support us
๐Ÿ‡ฐ๐Ÿ‡ท South Korea /Technology

AI grading systems struggle with logic, favor keywords

From Hankyoreh · () Korean

Translated from Korean, summarized and contextualized by DistantNews.

At a glance

News Named sources Context piece
  • AI automated grading systems show limitations in evaluating subjective essays, sometimes awarding high scores to illogical arguments.
  • These systems tend to favor longer texts and the frequent use of specific keywords over logical coherence and accurate understanding.
  • Experts suggest AI should serve as an auxiliary tool for human graders, not a replacement, to ensure fairness and accuracy in assessments.

The application of artificial intelligence for automated grading of subjective essays is revealing significant limitations, with some systems awarding top marks to responses containing logical fallacies but abundant specific keywords. This finding emerges as discussions intensify around introducing essay and subjective questions into South Korea's College Scholastic Ability Test (CSAT) starting in 2027.

A report by the Korea Institute for Curriculum Evaluation, released in December, detailed experiments with AI grading models across five subjects: Korean language, social studies, mathematics, science, and technology. In a Korean language essay task on lowering the age of juvenile offenders, an AI graded a response as an 'A' (top grade) despite a human grader assigning it a 'C'. The human grader cited issues such as grammatical errors, inconsistent use of honorifics, and a failure to grasp the core meaning of cited sources. The AI, however, prioritized the essay's length and its extensive use of vocabulary found in the prompt, leading to the inflated score.

The possibility of receiving a high score increases if the answer similarity is high due to the frequent use of vocabulary from the provided material or related vocabulary, regardless of the validity of the argument or clarity of expression.

โ€” Research TeamDescribing how AI grading systems can be misled by keyword usage rather than logical coherence.

Similar discrepancies were observed in other subjects. In science, an AI incorrectly awarded full marks to an answer that explained the principle of reducing impact when force is equal, even though the student had reversed the concept. The AI recognized specific keywords, overriding the factual inaccuracy. In mathematics, AI struggled with open-ended questions requiring interpretation, such as analyzing graphs, where multiple correct answers or phrasing could exist. It performed well on questions with single numerical answers but faltered on more nuanced tasks.

These findings suggest that current AI grading systems are overly sensitive to superficial elements like sentence length and grammatical structure, while potentially overlooking deeper comprehension and logical reasoning. Experts, including those involved in national AI strategy, advocate for AI to function as a supportive tool for human educators, assisting in preliminary grading and providing feedback, rather than assuming full control. The proposed roadmap emphasizes a phased implementation, starting with lower-stakes school assessments before considering its use in high-stakes exams like the CSAT, with robust safeguards against misuse and a clear emphasis on maintaining human oversight.

AI should be designed as an auxiliary tool to assist teachers' judgment, not to have full control over grading.

โ€” Lee Min-seokA professor at Kookmin University and head of the education and talent division of the National AI Strategy Committee, emphasizing AI's role as a support system.
DistantNews Editorial

Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.