News Research Prognostic & Predictive Models Breast Cancer

Two Studies Evaluate AI-Based TIL Scoring for Breast Cancer Prognostication

September 11, 2026 Julia Cipriano 7 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Two analyses published in The Lancet Oncology, both with Sherene Loi, MBBS (Hons), PhD, FRACP, FAHMS, GAICD, of Peter MacCallum Cancer Centre, Melbourne, Australia, as senior and corresponding author, found that AI-based assessment of tumor-infiltrating lymphocytes (TILs) provided prognostic information across early-stage triple-negative and HER2-positive breast cancers. The findings highlighted potential uses for computational TIL scoring alongside conventional pathologist assessment and in settings where such expertise may be limited.

The CATALINA study independently evaluated computationally assessed TIL scores against pathologist-scored stromal TILs using prospectively collected whole-slide images and long-term clinical outcome data pooled from several prospective, randomized trials. In a secondary analysis of the phase III APHINITY trial, investigators compared manual, digital, and AI-based stromal TIL quantification, together with AI-derived spatial metrics, and examined their associations with prognosis and treatment benefit.

CATALINA

In CATALINA, an independent, external validation study, investigators evaluated an analytical cohort of 220 digitized hematoxylin-and-eosin (H&E) whole-slide images from patients with early-stage triple-negative or HER2-positive breast cancer who had previously been scored by trained pathologists to assess correlations between computationally assessed TIL scores and the mean pathologist-scored stromal TIL score. They also evaluated digitized H&E whole-slide images and long-term outcome data from a separate clinical cohort of patients with early-stage triple-negative breast cancer pooled from seven prospective, randomized adjuvant trials to assess prognostic performance; individual data were collated from 1,759 patients, including 1,356 with complete clinicopathologic, pathologist-scored stromal TIL, and computationally assessed TIL data.

Two fully automated AI pipelines that were previously independently developed, trained, and validated—AI-TIL and MuTILs—were independently deployed as locked models without retraining on the CATALINA slides, with algorithm developers masked to clinical data. Together, the pipelines generated five prespecified computationally assessed TIL scores per slide, each quantifying lymphocytes relative to stromal or tumoral compartments.

The deep-learning AI-TIL model identified all cells, segmented viable tumor tissue and stromal areas, classified cells as cancer cells, lymphocytes, stromal cells, or other cells, and used unsupervised clustering to categorize lymphocytes based on their proximity to tumor regions. It computed three scores: percentage_lymphocyte, representing lymphocytes as a proportion of all detected cells; AI_TIL, representing tumor-adjacent lymphocytes relative to stromal cells; and the intratumoral immune infiltration ratio, representing intratumoral lymphocytes relative to cancer cells.

The MuTILs algorithm, an interpretable multiresolution panoptic segmentation neural network, used convolutional networks at two resolutions to segment tissue regions and nuclei. It generated two scores: calMeanAnyStroma, representing lymphocyte nuclei density per stromal area, and calMeanAllStroma, representing lymphocyte nuclei as a proportion of total fibroblast and lymphocyte nuclei within stromal tissue.

According to the investigators, the correlation between computationally assessed TIL scores and the mean pathologist-scored stromal TIL score was modest (Spearman’s correlation coefficient = 0.375–0.473).

Both stromal and computational TIL scores appeared to be independently associated with 5-year invasive disease–free survival, distant disease–free survival, and overall survival rates after adjustment for clinicopathologic factors, with hazard ratios (HRs) of 0.73 (95% confidence interval [CI] = 0.66–0.82; q < .0001), 0.70 (95% CI = 0.61–0.79; q < .0001), and 0.72 (95% CI = 0.63–0.82; q < .0001), respectively, for stromal TIL scores, and 0.80 (95% CI = 0.73–0.89; q < .0001), 0.77 (95% CI = 0.69–0.86; q < .0001), and 0.79 (95% CI = 0.70–0.88; q = .0002), respectively, for percentage_lymphocyte scores. In models adjusted for clinicopathologic variables and stromal TIL score, the prognostic association of computational TIL score was no longer statistically significant, the investigators wrote.

Additionally, stromal and computational TIL scores were found to improve 5-year prognostic discrimination compared with clinicopathologic variables alone; however, the investigators noted that the computational TIL score did not further improve discrimination significantly when combined with clinicopathologic variables and stromal TIL score.

Nevertheless, they concluded that “these findings support the application of computationally assessed TILs as a reproducible prognostic biomarker, particularly in settings where routine or widespread pathologist assessment is unavailable.”

APHINITY

The secondary analysis of APHINITY included H&E slides from 4,262 participants with paired TIL data available in the randomized, double-blind phase III trial, in which 4,805 patients with operable HER2-positive early-stage breast cancer received standard adjuvant chemotherapy plus trastuzumab with either pertuzumab or placebo. Stromal TILs were quantified manually by an expert pathologist, with a fully automated non-AI digital approach, and with a pretrained AI model.

The deep-learning Case45 pipeline was applied without retraining or optimization on APHINITY data to classify and spatially map cancer cells, lymphocytes, stromal cells, and other cells. The AI-derived lymphocyte percentage score (AI-percentage lymphocyte) quantified AI-detected lymphocyte nuclei as a proportion of all AI-detected nucleated cells within the tumor-associated tissue, whereas two AI-derived spatial features, AI-TIL and immune hotspot, captured the ratio of mononuclear inflammatory to stromal cells in peritumoral stroma and the fraction of immune-only cell aggregates relative to all detected spatial clusters, respectively.

To assess interobserver reproducibility, five pathologists independently scored 262 randomly selected tumor samples. Manual scoring demonstrated high interobserver reproducibility (intraclass correlation coefficient = 0.84 [95% CI = 0.79–0.88]), the investigators wrote, although concordance between manual and automated methods was modest across all evaluable cases.

AI-percentage lymphocyte scoring reclassified 11.6% of node-positive tumors in the pertuzumab group from immune low (≤ 75th percentile) by manual assessment to immune high (> 75th percentile). Patients with these discordantly classified tumors showed greater separation in 5-year invasive disease–free survival curves between pertuzumab and placebo than those with tumors classified as immune low by both approaches.

Higher levels of TILs appeared to be associated with improved invasive disease–free survival across all stromal TIL measurement approaches and spatial measurements, with HRs ranging from 0.41 to 0.93. An association between pertuzumab and improved invasive disease–free survival was observed in immune-high groups defined by manual assessment, digital scoring, and AI-percentage lymphocyte scoring (HRs = 0.36–0.48), but not in those defined by the spatial measures.

The largest 6-year absolute improvements in invasive disease–free survival, distant recurrence–free interval, and overall survival with pertuzumab were observed among patients with node-positive disease whose tumors had the highest level of immune infiltration by manual stromal TIL scoring (≥ 70%), with a mean absolute improvement of 12.1 percentage points. In nested prognostic and predictive modeling, combining immune hotspot scores with any stromal TIL measurement yielded the most consistent additional information (all P < .01).

The investigators concluded, “Overall, our findings support a complementary model in which manual stromal TIL scoring remains a robust standard, whereas computational methods provide scalable and biologically informative extensions.” They noted that models combining immune quantity and spatial organization performed best and called for independent validation of the studied approaches and further evaluation of their clinical utility for stratifying contemporary HER2-directed therapies.

DISCLOSURES: CATALINA was funded by the Breast Cancer Research Foundation. The secondary analysis of APHINITY reported no funding source; the parent trial was funded by Roche. For full disclosures of the authors of the CATALINA study and APHINITY secondary analysis, as well as expanded funding and data and code availability, visit thelancet.com. 

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content