From Task-Specific Tools to Generalist Models: Charting Oncology's Next AI Shift
A proposed “Leave No Data Behind” paradigm suggests that foundation models and large language models may make better use of oncology’s fragmented clinical, imaging, pathology, and molecular data by bringing all such data together to support clinically meaningful applications. The paradigm was laid out in a narrative review published in Cell Reports Medicine.
"Oncology generates enormous amounts of heterogeneous data, from clinical records and radiological imaging to digital pathology and molecular profiles, but much of this information remains fragmented or underused," corresponding author Arsela Prelaj, MD, PhD, a medical oncologist and Head of the AI-ON-Lab at the Fondazione IRCCS Istituto Nazionale Dei Tumori di Milano in Italy, told ASCO AI in Oncology. "Our review introduces the concept of 'Leave No Data Behind' to describe how foundation models and large language models could move oncology AI away from narrow, task-specific systems toward models that learn from much broader and more diverse sources of information."
A potential advantage of the "Leave No Data Behind" paradigm is the ability to integrate information across data types rather than analyze each source separately. However, the authors emphasized that just processing more data is not enough. "What matters is whether these models can turn that information into outputs that are genuinely useful for diagnosis, biomarker discovery, treatment prediction, research, and clinical workflows," Dr. Prelaj said.
Review
The authors conducted a broad, nonsystematic literature review of foundation models and large language models across four major oncology data domains: clinical language, digital pathology, radiology, and molecular biology. They examined whether current evidence supports translating the value of these data into improved generalizability and clinical applicability.
The authors contrasted the generalist approaches of foundation models and large language models with traditional machine learning and deep learning models developed for specific tasks using task-specific architectures and labeled data.
The meaning of "Leave No Data Behind" varies by modality. For clinical language, the paradigm encompasses sources including medical literature, guidelines, electronic health records, diagnostic reports, research databases, and clinical trial documents. In digital pathology and radiology, it extends to heterogeneous imaging data, while molecular applications include genomic, transcriptomic, proteomic, and single-cell information.
Key Findings
Across the applications reviewed, evidence was more mature for diagnostic and workflow-oriented tasks than for more complex predictive applications, including biomarker prediction, treatment-response estimation, prognostic modeling, and therapy selection.
"One of the clearest findings was that bigger is not necessarily better," Dr. Prelaj said. "Larger models, larger datasets, and greater computational resources did not consistently translate into better clinical performance. The composition, diversity, representativeness, and relevance of the training data appeared to matter enormously."
Clinical language provides one example of how previously underused data may be made more accessible. In a multicenter study included in the review, a fine-tuned large language model extracted 31 clinical variables from unstructured records of 10,327 patients with lung cancer. The model had a 7% error rate vs 14.2% for manual abstraction, with F1 performance scores of 0.97 vs 0.86 for key variables.
Digital pathology provided evidence for another premise underlying “Leave No Data Behind”: the importance of data diversity. In a benchmark of pathology foundation models across 31 morphology, biomarker, and outcome-prediction tasks, multimodal models consistently outperformed image-only models, particularly in low-resource settings. The CONCH model had the highest average area under the receiver operating curve at 0.71 vs 0.66 for Virchow, and data diversity contributed more strongly to performance than data set size alone.
Performance was less consistent for more complex predictive tasks. In radiology, cancer-detection accuracy ranged from 80% to 94% depending on cancer type, while the area under the receiver operating curve for prognosis ranged from 0.52 to 0.70, which the authors characterized as “promising, but still far from maturity.”
The paradigm also extends to research applications. Molecular foundation models have been evaluated for tasks including variant interpretation, biomarker discovery, and therapeutic hypothesis generation. The C2S-Scale Gemma-2 27B model identified silmitasertib plus low-dose interferon as a potential strategy to increase antigen presentation in tumor cells, with experimental validation reported in the original study. The review cautioned, however, that experimental validation of model-generated hypotheses remains uncommon and that findings may reflect dataset biases or known biological relationships rather than new causal mechanisms.
Multimodal foundation models offer another direct expression of the paradigm by combining sources such as imaging, text, and molecular data. The review found the strongest evidence for segmentation, abnormality detection, disease classification, and report generation, with more preliminary evidence for prognosis and prediction. Clinical translation remained limited by heterogeneous acquisition protocols, scarce multimodal annotations, image-text misalignment, computational demands, uncertain generalizability, and insufficient real-world validation.
Limitations and Future Research
The authors identified several limitations across the evidence reviewed, including fragmented benchmarks, limited external validation, insufficient clinically meaningful endpoints, and evaluation of models on data from the same repositories used for pretraining, which can limit assessment of generalization. Many models trained on private institutional data sets were also difficult to reproduce or validate externally. Larger models and data sets did not consistently produce better downstream performance; results depended on data composition, representativeness, and alignment with the intended clinical use.
The authors also cited hallucinations, overconfident outputs, bias, privacy and security concerns, model drift, and the need for human oversight as barriers to clinical implementation.
Before wider clinical use, the authors called for oncology-specific benchmarks, open multi-institutional datasets and stronger clinical validation, including multicenter prospective studies, silent trials, and randomized trials. They added that future research should determine when human-AI collaboration outperforms either clinicians or AI alone and address uncertainty, explainability, accountability, and postdeployment monitoring.
"The practical message for oncology is that foundation models and large language models are likely to become powerful clinical teammates, but their deployment needs to be evidence-based, transparent, continuously monitored, and governed with appropriate human oversight," Dr. Prelaj said.
DISCLOSURES: For full disclosures of the study authors, visit cell.com.
ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.