Single-cell atlases, which are increasingly used to map the human body and train artificial intelligence models, may not be fairly representative of the world’s population, according to a study led by researchers at the Icahn School of Medicine at Mount Sinai. The analysis found that while people of European descent were consistently overrepresented, people of Asian and Latinx descent were underrepresented, and a large portion of the sample had no record of any ancestry. The findings were published on July 20th. cell genomics (10.1016/j.xgen.2026.101300).
Single-cell technology allows scientists to profile the biology of one cell at a time, revealing rare cell types and disease-related changes that older, high-volume methods missed. A large-scale international effort is using these tools to build reference maps, vast datasets intended to serve as universal references for research and medicine. But the researchers say the demographic composition of these resources has not been systematically investigated, raising questions about whether the maps reflect all of humanity or just some of us.
The team reviewed more than 13,500 samples from three major single cell resources: Human Cell Atlas, Human Tumor Atlas Network, and PsychAD Consortium. The researchers handpicked the reported ancestry, race, ethnicity, and gender of the sampled individuals and compared each dataset to global population data, U.S. cancer incidence data, and disease-specific reference data, looking for patterns across tissues, cancer types, and brain disease categories.
These atlases are becoming reference maps in biology and medicine, and are increasingly used to train AI models that will shape future research and care. We wanted to ask a simple but important question. Do these maps fairly represent different groups of people, or are some groups left out? What we’ve discovered is that some of our most important datasets aren’t as representative as they should be. ”
Kuan-lin Huang, Ph.D., senior corresponding author, associate professor of genetics and genomic sciences and artificial intelligence and human health, Icahn School of Medicine
The gap was consistent across all three resources. In the Human Cell Atlas, nearly 70% of samples had no ancestry information recorded at all. Among the detected samples, individuals of European descent were approximately six times more represented compared to global expectations, while individuals of Asian, African, and Latinx ancestry were underrepresented. In this study, “Latino” refers to the ethnicity reported in the available data, rather than a single genetic ancestry, and its underrepresentation highlights the need for single-cell datasets that better reflect the communities they aim to serve.
Approximately 69% of the Human Tumor Atlas Network tumor samples were European, and the PsychAD brain dataset was nearly two-thirds European. Several cancer types also showed unexpected gender differences in disease incidence.
The researchers note that the overrepresentation of Europeans in the human cell atlas persists even under the most conservative assumptions about missing data and cannot be explained by incomplete records alone.
“The most surprising finding was how much ancestry information is simply missing from some of these expensive studies,” Dr. Huang says. “These gaps can be passed on to the AI model trained on the dataset, often without the user of the AI model realizing it. That’s why it’s so important to build diversity and complete demographic information into single-cell studies from the beginning.”
The concerns, the authors say, are one of fairness and reliability. If future research, AI tools, biomarkers, or treatments are built on datasets that are not representative of all people, the benefits of those advances may not reach everyone equally. The problem reflects a long-standing problem in human genomics, in which studies conducted primarily in people of European descent produce risk scores that do not translate well to other populations.
This study is one of the first attempts to systematically check who is included in major single-cell atlases. Rather than just asking about the size and detail of the dataset, the researchers asked whether the dataset was fair and generalizable across the population. The team also provides practical, discipline-specific checklists that research groups can use when planning future research. This checklist includes how to plan for recruitment, record demographic information, balance samples, and report whether your AI model performs equally well across ancestry and gender groups.
“The important point is that representation matters. If these are to become reference maps of human biology, they need to reflect the diversity of humanity,” says Dr. Huang. “These resources are already extremely valuable, and we wanted to help the field improve them so that we can serve more people equitably. This ultimately means more reliable science and more equitable precision medicine.”
The analysis was conducted by a team of student researchers including co-lead authors Katrina Yang of the University of Oxford, Kavisarini Saravanan of the University of North Carolina at Charlotte, and Aryan Saharan of Saint Louis University, in collaboration with Dr. Huang of the Icahn School of Medicine.
The authors acknowledge that this study relies on demographic information reported in public datasets rather than direct measurements of genetic ancestry, and that broad ancestry categories do not fully capture the complexity of identity. Because this study focused on three major public consortia, the findings may not reflect all single-cell datasets in the field. As next steps, the team plans to extend this type of audit to additional datasets, track whether representation improves over time, and assess whether AI models trained on single-cell data perform differently across populations.
sauce:
Mount Sinai Health System
Reference magazines:
Yang, C. others. (2026) Are different populations fairly represented in single-cell omic atlases? Cell genomics. DOI: 10.1016/j.xgen.2026.101300. https://www.cell.com/cell-genomics/fulltext/S2666-979X(26)00162-X

