Costly Public AI Data Projects Hampered by Redundancy and Poor Quality
Translated from Korean, summarized and contextualized by DistantNews.
At a glance
- A government audit revealed significant issues with AI data construction projects, which cost 1.6 trillion won.
- Many datasets were found to be redundant or unusable for AI training due to a lack of coordination between public institutions.
- The audit recommended establishing a system for sharing and coordinating data construction plans to prevent duplication and improve quality.
A government audit has exposed significant flaws in South Korea's ambitious artificial intelligence data construction projects, which have consumed approximately 1.6 trillion won. The audit found that a substantial portion of the data collected is either redundant or unsuitable for AI training, undermining the government's push for an AI-driven transformation.
Since 2017, the Ministry of Science and ICT (MSIT) has invested 1.63 trillion won to build 908 types of AI learning datasets. However, a review by the Board of Audit and Inspection revealed that 26 public institutions, including the Seoul Metropolitan Government and Korea Expressway Corporation, independently built datasets without adequate prior sharing or coordination. This resulted in numerous instances where newly created data was strikingly similar to existing datasets.
For example, a dataset of 9,000 photos of domestic waste, built by the Seogu district of Daejeon at a cost of 130 million won, was found to be 57.8% similar to data previously compiled by the MSIT for 1.6 billion won in 2020. Similarly, data on pills, oral cavities, and autonomous driving, built by MSIT for 23.3 billion won between 2021 and 2023, overlapped significantly with datasets produced by three other agencies for 840 million won. Even a collection of 1,000 dental images, costing 308 million won, was identical to MSIT's data.
In response, the Board of Audit and Inspection recommended that public institutions first assess the potential for reusing existing data before undertaking new construction projects. It also advised implementing a preliminary review system to share and coordinate data building plans among different agencies.
Quality control was also found to be lacking. A 'CCTV learning dataset' prepared by the Korea Expressway Corporation was unusable for autonomous driving development because it omitted essential vehicle location information. Furthermore, four out of 20 institutions with high data construction performance lacked any quality control standards.
Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.