DistantNews
Support us
๐Ÿ‡ฐ๐Ÿ‡ท South Korea /Technology

South Korean AI project faces 'benchmaxing' controversy ahead of second evaluation

From Hankyoreh · () Korean

Translated from Korean, summarized and contextualized by DistantNews.

At a glance

News Named sources Context piece
  • Domestic AI companies participating in South Korea's "Dokpamo" project face allegations of "benchmaxing" with foreign firms.
  • Concerns arise that these collaborations may have involved using test data for model training, potentially inflating benchmark scores.
  • The government emphasizes a comprehensive evaluation beyond just benchmarks, including real-world usability.

Allegations of "benchmaxing" are swirling around South Korea's "Dokpamo" (Domestic AI Foundation Model) project as its second evaluation approaches. Several domestic AI companies reportedly received proposals from foreign AI firms, such as the U.S.-based AfterQuery, offering to "improve benchmark scores" for their models. This has sparked concerns that some participants might have collaborated with these external companies, potentially compromising the integrity of the evaluation.

AfterQuery, a company specializing in AI training data and model optimization, is said to have actively approached domestic developers. The company is known for "benchmaxing," a technique focused on optimizing models to achieve high scores on specific benchmarks. Industry insiders suspect that some Dokpamo participants may have pursued collaborations with AfterQuery, despite denials from all four finalists. Suspicious circumstances are reportedly visible in the technical reports submitted ahead of the evaluation.

While purchasing training data from external vendors is not a violation of the Dokpamo evaluation rules, the core of the controversy lies in whether evaluation data was used for model training during the optimization process. This issue echoes recent concerns about data contamination in other AI models, where unusually high benchmark scores suggested excessive model tuning. The principle in AI development is to separate "training data" from "test data" to ensure accurate performance measurement.

If test data, intended solely for evaluation, is mixed into the training process, AI models might "memorize" solutions rather than learn problem-solving principles. This could lead to inflated benchmark scores that do not reflect true performance on novel problems or in real-world applications. In response, the Ministry of Science and ICT stated that the evaluation will be comprehensive, considering not only benchmark scores but also real-world usability assessed by experts and the public.

We will comprehensively evaluate based on various criteria, including not only benchmark evaluations but also expert and user assessments of actual usability.

โ€” Ministry of Science and ICT officialAddressing concerns about potential data manipulation in the Dokpamo project's evaluation.
DistantNews Editorial

Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.