New tool detects hidden biases in medical AI training data
Translated from English, summarized and contextualized by DistantNews.
At a glance
- Researchers from Johns Hopkins University and the FDA developed G-AUDIT, a tool to detect hidden biases in medical AI training data.
- AI systems can underperform and exhibit bias due to non-representative training datasets, leading to potential health disparities.
- G-AUDIT focuses on identifying data set problems before AI model training, addressing issues like shortcut learning and irrelevant pattern association.
A new tool called G-AUDIT, developed by researchers at Johns Hopkins University and the U.S. Food and Drug Administration (FDA), aims to identify hidden biases within the datasets used to train medical artificial intelligence systems. These AI systems hold significant promise for improving healthcare, but have frequently shown underperformance and serious implicit biases, often stemming from the data they are trained on.
The models that drive precision medicine learn to infer clinical outcomes from the data they're trained on. In many cases, that works great, but it can also lead to interesting failures that aren't immediately apparent.
The core issue lies in training datasets that are not representative of the diverse populations for whom the AI is intended. Mathias Unberath, an AI-assisted medicine expert at Johns Hopkins, explained that while AI models learn clinical outcomes from data, this can lead to "interesting failures that aren't immediately apparent." Such data bias can result in AI "shortcut learning," where the system associates irrelevant patterns with clinical outcomes, leading to biased predictions and a lack of generalization.
An example cited illustrates this problem: an AI learned to associate clinician markings on skin with malignant lesions. When tested, it produced 40% more false positives when these markings were present. Another AI, trained on data from two clinics with different imaging qualities and the frequent presence of rulers in cancer clinic images, learned to associate rulers and camera quality with cancer risk. Unberath warned, "You have just created health disparity" if such an AI is deployed with different equipment, like a smartphone camera, that lacks these markers.
Envision taking this algorithm into the real world with a smartphone camera. There's no ruler. The camera quality is completely irrelevant. But the model learned to associate 'ruler' and 'camera.' You have just created health disparity.
While much research focuses on adapting AI models themselves, G-AUDIT targets the data sets directly. The tool, officially named Generalized Attribute Utility and Detectability-Induced Bias Testing, is designed to flag potential problems in data before AI models are trained. Researchers tested G-AUDIT on both image datasets and health record texts, aiming to preempt the critical errors that AI can make when applied outside its original training data.
We train models to be the most predictive, but we have zero control over what the model uses to make the prediction.
Originally published by Jerusalem Post in English. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.