



Researchers studying AI dermatology tools found that accuracy drops sharply for darker skin because the models often key off surrounding skin tone rather than the lesion itself. In one test, digitally darkening the skin around a benign mole while leaving the mole unchanged caused OpenAI's GPT-4 to classify it as malignant melanoma. The same bias showed up with atopic dermatitis, which the models could recognize as pink on light skin but frequently missed when it appeared gray or violet on darker skin.
The root cause is training data: public medical image libraries and dermatology textbooks have historically skewed toward lighter-skinned patients, so the models never learned what these conditions look like on darker skin. Patients of color are already diagnosed with skin cancers like melanoma at more advanced stages with worse survival rates, and a tool biased toward lighter skin risks widening that gap rather than closing it.
One proposed fix is generating synthetic training images of conditions on darker skin using generative AI, which researchers say can match real-image performance in tests. But they caution these synthetic images might not accurately reflect how conditions actually present on real patients, risking a tool that looks diverse on paper while remaining blind in practice.
The full dispatch is available from the source below.