Impact of tissue staining and scanner variation on the performance of pathology foundation models a study of sarcomas and their mimics

Binghao Chai1, Jianan Chen1, Paul Cool2,3, Fatine Oumlil1,4, Anna Tollitt4, David F. Steiner5, Tapabrata Chakraborti1,6, Adrienne M. Flanagan1,4

1Research Department of Pathology, UCL Cancer Institute, University College London, London, UK.
2The Robert Jones and Agnes Hunt Orthopaedic Hospital, Gobowen, UK.
3Keele University, Keele, UK.
4Royal National Orthopaedic Hospital, Stanmore, UK.
5Google Health, Mountain View, CA, USA.
6The Alan Turing Institute, London, UK.

Paper Code Dataset

Interactive Figures

Diagnosis-Level Performance

t-SNE Embeddings

About the Study

Histopathological analysis is considered the gold standard for the diagnosis and prognostication of cancer. Recent advances in AI, driven by large-scale digitisation and pan-cancer foundation models, are opening new opportunities for clinical integration. However, it remains unclear how robust these foundation models are to real-world sources of variability, particularly in H&E staining and scanning protocols. In this study, we use soft tissue tumours, a rare and morphologically diverse tumour type, as a challenging test case to systematically investigate the colour-related robustness and generalisability of seven AI models. Controlled staining and scanning experiments were utilised to assess model performance across diverse real-world data sources. Foundation models, particularly UNI-v2, Virchow and TITAN, demonstrated encouraging robustness to staining and scanning variation, particularly when a small number of stain-varied slides were included in the training loop, highlighting their potential as adaptable and data-efficient tools for real-world digital pathology workflows.