The MOSAIC dataset
Large-scale spatial omics data fuels cancer research

There is an urgent need for the whole cancer community to have access to these data to reveal groups of patients with distinct tumour/immune biology.
MOSAIC will overcome this limitation by creating a large scale dataset of tumor spatial transcriptomes, allowing scientists to conduct research on data from cohorts larger than what is currently possible.
The MOSAIC dataset today
Data modalities and cancer therapy areas
MOSAIC captures the complexity of disease biology across six data modalities and selected cohorts spanning ten cancer indications. Here is a breakdown of the current number of patient samples across each modality and therapy area in the dataset.
How to access the MOSAIC dataset
MOSAIC Window, a subset of the MOSAIC dataset, is available via the European Genome-Phenome Archive (EGA).
Explore the MOSAIC dataset via Owkin’s AI Scientist, K Pro.
K Pro is the first AI agent connecting research and care. It continuously learns from real-world data, user feedback, and clinical validation, evolving toward fully automated R&D.