DistantNews
Support us
KAIST develops technology to eliminate AI 'hallucinations'
๐Ÿ‡ฐ๐Ÿ‡ท South Korea /Technology

KAIST develops technology to eliminate AI 'hallucinations'

From Dong-A Ilbo · () Korean

Translated from Korean, summarized and contextualized by DistantNews.

At a glance

News Sources not specified New plan
  • KAIST researchers have developed core technologies to reduce "hallucinations" in multimodal AI, preventing it from making up information.
  • The two key technologies focus on improving AI's understanding of various sensors and preventing "sensory confusion" between different data types.
  • These advancements can significantly enhance AI reliability without requiring extensive retraining, making them applicable to fields like autonomous driving and disaster response.

Researchers at the Korea Advanced Institute of Science and Technology (KAIST) have developed crucial technologies aimed at curbing "hallucinations" in artificial intelligence, effectively stopping AI from fabricating information it hasn't perceived. This breakthrough addresses the issue of multimodal AI systems falsely claiming to see or hear things that are not present.

The research team, led by Professor Noh Yong-man of the Department of Electrical and Electronic Engineering, introduced two core technologies. The first, dubbed 'Sensor Understanding AI' (DNA) optimization, enables AI to accurately grasp the physical meaning behind data from various sensors like thermal cameras, depth sensors, and X-rays. This allows AI to better interpret information in challenging conditions, such as darkness or smoke, by understanding physical properties like heat and distance.

The second technology, 'Sensory Confusion Prevention' (MAD), tackles the problem of AI misinterpreting or confabulating information when multiple sensory inputs, like visual and auditory data, are mixed. For instance, an AI might incorrectly state it hears an engine sound from a video showing a car, even if no sound was recorded. MAD allows the AI to first determine which sensory input is most relevant to a given question, focusing on that specific data to reduce errors.

These innovations offer a significant advantage by improving AI reliability without the need for costly and time-consuming large-scale retraining. The DNA technology requires less data to train AI on diverse sensors, while MAD can be applied to existing AI systems like a software add-on. This efficiency makes the technologies readily applicable to various industries, including autonomous vehicles that need to accurately identify pedestrians and vehicles in low visibility, and rescue robots operating in smoke-filled environments.

Professor Noh emphasized the importance of AI accurately understanding sensor characteristics and avoiding confusion between different senses for real-world applications. He stated, "This research lays the foundation for multimodal AI that can be reliably used in real life and industrial settings by reducing AI's sensory bias and hallucinations without large-scale retraining."

The research findings have been presented at leading academic conferences and published in relevant journals. The MAD research was presented at the International Conference on Computer Vision and Pattern Recognition (CVPR) in June, while the DNA research was published in the IEEE Transactions on Image Processing.

Multimodal AI needs to accurately understand the characteristics of various sensors and avoid confusion between different senses to be used in real environments. This research lays the foundation for multimodal AI that can be reliably used in real life and industrial settings by reducing AI's sensory bias and hallucinations without large-scale retraining.

โ€” Noh Yong-manExplaining the significance of the developed technologies for practical AI applications.
DistantNews Editorial

Originally published by Dong-A Ilbo in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.