Multitask Vision Transformer Predicts Molecular and Prognostic Features Across 32 Solid Tumors
Researchers developed an AI-based Vision Transformer model to analyze routine hematoxylin and eosin (H&E)–stained whole-slide images and then simultaneously predict cancer subtype, TP53 mutation status, and survival outcomes across 32 solid tumors. Study findings were published in The American Journal of Pathology.
“This approach could help identify patients who may benefit from confirmatory molecular testing, support triage in settings with limited genomic testing, and provide additional decision support to clinicians. Importantly, this method should be viewed as complementary to molecular testing, not a replacement. Its potential impact is strongest as a screening, prioritization, or decision-support tool within broader diagnostic pathways," noted co-lead investigator Abadh K. Chaurasia, PhD, of the Menzies Institute for Medical Research, University of Tasmania, Australia and Pandani Solutions Pty Ltd.
The study authors also highlighted the potential of computational pathology and deep learning to be able to extract more clinically relevant morphologic, prognostic, and molecular features from whole-slide images.
Study and Model Methods
Researchers developed a Vision Transformer–based multi-instance learning model to identify the presence of a TP53 biomarker, detect 32 solid tumor types, and predict survival based on whole-slide images.
They gathered 11,000 primary tumor data from The Cancer Genome Atlas’ Pan-Cancer Atlas and somatic mutation, RNA sequencing, and clinical outcome data as training data for the model. The whole-slide images underwent preprocessing to extract patches, normalize staining, and implement quality control. Then with a Vision Transformer encoder, features were encoded from the slides for standardized input into the AI framework.
For the predictive tasks, seven task heads with a single, shared backbone were developed to account for cancer type, TP53 mutation status, TP53 RNA expression level, overall survival event, progression-free survival event, and time-to-event variables for both progression-free and overall survival. This created a weakly supervised architecture that minimized computational and training costs while focusing on the practicality of the model’s clinical utility.
“Standard molecular profiling for TP53 mutations is often costly and inaccessible in underprivileged or remote clinical settings,” explained co-lead investigator Alex W. Hewitt, PhD, also of the Menzies Institute for Medical Research and School of Medicine, University of Tasmania, Australia. “We wanted to develop a more practical tool for pathologists. Currently, most deep learning-based models are used for single-model concepts; one model for one task. We developed a single model that can generate seven outputs simultaneously from the whole histopathology image, including TP53 mutation status, TP53 RNA expression, tumor type, and survival-related outcomes at the slide level.”
Model training was first done on tumor-only patches at multiple magnifications and then the model was fine-tuned on whole-slide images with a content-aware strategy. Eighty percent of the Pan-Cancer Atlas cohort was used for model training and the remaining 20% for validation in both stages.
Performance evaluations were conducted on an independent validation group of 1,729 slides.
Key Findings
The model achieved an area under the receiver operating characteristic curve of 0.766 for TP53 mutation status detection on the independent validation set across 32 cancer types.
“In this study, molecular labels such as TP53 mutation status were available at the patch level, but whole slide images containing TP53-associated morphological information were not manually labeled. Weak supervision enabled the model to learn from slide-level labels and identify relevant patterns across image patches without requiring exhaustive pixel- or region-level annotations,” Dr. Hewitt noted.
The model was also able to infer TP53 RNA expression levels, with reasonable alignment between predicted and observed expression levels, and tumor taxonomy based on the histopathology images.
In terms of tumor classification, the fine-tuned model achieved an overall accuracy of 0.659 on internal validation, which was considered strongly generalizable for the different tumor types. With the exception of ovarian cancer (n = 1), most tumor types achieved area under the receiver operating characteristic curves ranging from 0.887 to 0.999. Ovarian cancer, meanwhile, had an area under the receiver operating characteristic curve of 0.782; accuracy, sensitivity, and specificity were also the lowest for ovarian cancer.
The largest cohort of slides represented breast invasive carcinoma (n = 114), which demonstrated an area under the receiver operating characteristic curve of 0.98, with an accuracy, sensitivity, and specificity of 0.939.
For survival prediction, the model was able to clearly separate high- and low-risk groups with highly significant differences, indicating a consistent association between predicted risk and disease progression.
“Accurate molecular profiling from routine histopathology slides, already widely used in cancer care, could transform clinical oncology,” Dr. Hewitt concluded. “This new AI-based model integrates diagnostic, molecular, and prognostic tasks, and could help clinicians obtain more information from existing pathology workflows, ultimately supporting more accessible precision cancer care and early intervention.”
DISCLOSURES: This study was supported by an Australian National Health and Medical Research Council Leadership Award. Drs. Chaurasia, Bennett, and Hewitt are cofounders of Pandani Solutions Pty. Ltd., which develops artificial intelligence–based tools in ophthalmology and computational pathology. For full disclosures of the other study authors, visit ajp.amjpathol.org.
ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.