Foundation models are widely used for compressing complex data into vector embeddings (vembs), offering reduced storage and computational efficiency, and significant improvements in diagnostic accuracy, automation, and efficiency in medical imaging. However, concerns remain that these vembs may encode demographic features, which could provide a pathway for bias in AI models in medical imaging. This study investigates whether demographic attributes – including sex, age, ethnicity, and insurance type – are embedded in vembs derived from chest radiographs in the MIMIC-CXR and CheXpert datasets. We generate vembs using three different state-of-the-art contrastive learning-based foundation models, namely, CXR Foundation, MedCLIP, and BiomedCLIP, assessing demographic predictability. Through rigorous statistical analysis and machine learning evaluations, we demonstrate substantial demographic encoding, indicating a plausible pathway through which bias may propagate. Our findings provide cautionary evidence supporting the need for further investigation and auditing of potential biases in vemb-based medical imaging predictions.

Hidden in Plain Sight: Vector Embeddings give away Demographic Information

Quarta A.;Marzullo A.;Calimeri F.
2026-01-01

Abstract

Foundation models are widely used for compressing complex data into vector embeddings (vembs), offering reduced storage and computational efficiency, and significant improvements in diagnostic accuracy, automation, and efficiency in medical imaging. However, concerns remain that these vembs may encode demographic features, which could provide a pathway for bias in AI models in medical imaging. This study investigates whether demographic attributes – including sex, age, ethnicity, and insurance type – are embedded in vembs derived from chest radiographs in the MIMIC-CXR and CheXpert datasets. We generate vembs using three different state-of-the-art contrastive learning-based foundation models, namely, CXR Foundation, MedCLIP, and BiomedCLIP, assessing demographic predictability. Through rigorous statistical analysis and machine learning evaluations, we demonstrate substantial demographic encoding, indicating a plausible pathway through which bias may propagate. Our findings provide cautionary evidence supporting the need for further investigation and auditing of potential biases in vemb-based medical imaging predictions.
2026
AI fairness
demographic information
machine learning
medical imaging
vector embeddings
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.11770/413461
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact