News Features Lung Cancer Prognostic & Predictive Models Diagnostics & Imaging

How AI Is Shaping Decision-Making in Lung Cancer Across Disciplines

October 07, 2026 Lisa Astor 15 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

AI has had a significant impact on clinical decision-making across medicine, from computational pathology to auto-segmentation to risk prediction models. However, no AI tool is currently approved to determine how a patient with lung cancer should be treated.  

In a scientific plenary session at the International Association for the Study of Lung Cancer (IASLC) 2026 World Conference on Lung Cancer (WCLC), three presenters considered the role of AI in clinical decision-making for lung cancer, each from the perspective of their own specialty: medical oncology, radiation oncology, and surgery.  

Across these three perspectives, the physicians outlined considerations for when AI can be trusted in clinical scenarios, where it can currently be used, and where the evidence does not yet support an expanded role. Automation can be audited and corrected, but a delivered radiation dose cannot be reversed, nor can the consequences of a decision not to operate. Responsibility for clinical decisions, all three agreed, remains with the clinician. 

Medical Oncology Perspective 

Mihaela Aldea, MD, PhD
Mihaela Aldea, MD, PhD

Mihaela Aldea, MD, PhD, Assistant Professor of Medical Oncology at Gustave Roussy in France, explained that in medical oncology, AI has become useful for streamlining processes, including reviewing and proposing management options for patients in routine practice. When oncologists must choose between treatments without clear data establishing a preferred option, AI may help identify predictive biomarkers to guide treatment selection. And for unresolved challenges in oncology, such as undruggable targets, AI may aid in drug discovery. 

There are currently no AI tools approved for treatment decision-making in lung cancer, although many are in development. Dr. Aldea noted that one tool that may be among the closest to approval is the TROP2 RxDx Device companion diagnostic, which measures the TROP2 normalized membrane ratio using quantitative continuous scoring. In the phase III TROPION-Lung01 trial, patients with nonsquamous non–small cell lung cancer (NSCLC) without actionable genomic alterations whose tumors were classified as positive for the biomarker (≥ 75% of tumor cells expressing a TROP2 normalized membrane ratio ≤ 0.56) demonstrated improved outcomes with datopotamab deruxtecan vs docetaxel. The computational pathology device was granted a Breakthrough Device designation by the U.S. Food and Drug Administration (FDA) in April 2025.  

Dr. Aldea noted that the companion diagnostic is a true predictive, rather than prognostic, biomarker; however, prospective validation of the biomarker has yet to be published.  

A number of factors should be considered when assessing an AI-based biomarker for use in practice, including population, ground truth, performance, generalizability, and fairness. When determining whether an AI tool can be used for an individual patient, Dr. Aldea stressed that “we need to make sure that the patient is similar to the population in which the AI test was developed.” 

She also noted that the comparator used in the analysis and verification of a biomarker should represent the ground truth, whether based on expert pathologist scoring, molecular testing, international guidelines, or known biomarkers. How the biomarker's performance is assessed is also important. Predictive biomarkers should ideally be validated in a randomized trial with two or more study arms that demonstrates a treatment–biomarker interaction. However, she noted that clinically significant thresholds may vary according to patient characteristics or preferences.  

To illustrate the importance of generalizability and applicability to a health system’s patient population, she pointed to the EAGLE and DeepGEM models for identifying EGFR mutations from hematoxylin and eosin (H&E) slides. The EAGLE model was trained on samples from U.S. patients and achieved an area under the curve (AUC) of 0.87, whereas the DeepGEM model was trained on samples from Chinese patients and achieved an AUC of 0.86. When both models were tested for ancestry-associated performance variability, researchers found that DeepGEM’s performance declined in U.S. and European patients, with an AUC of 0.68. Both models performed worse on samples from patients of Asian ancestry.  

“For these AI models we have this unmet need to be able to quantify the uncertainty at the patient level and for the AI to be able to produce some confidence score where we could evaluate how sure is the AI for its prediction for the patient in front of me,” Dr. Aldea said.  

She also stressed that changing patient management plans based on an AI prediction requires prospective validation of the model’s clinical utility to fully understand the potential consequences for the patient if the prediction is wrong.  

“Who will be responsible for erroneous predictions? For sure it will not be ChatGPT or Claude,” Dr. Aldea commented. “So, to bring these AI tools from proof of concept to decision making, we, as oncologists, need to start asking questions that are relevant for the patients today and for tomorrow.”  

To further advance responsible AI tools for clinical decision-making, Dr. Aldea encouraged developers to demonstrate generalizability across health-care settings, particularly in low- and middle-income countries, including by making their models open source for further testing. She also called on the industry to expand access to large, globally representative datasets for testing, and she emphasized the need to educate hospitals about AI tools so they are better prepared to integrate them into clinical workflows.  

“It is all of our responsibility today that AI should narrow, not widen, the gaps between patients across the world,” Dr. Aldea concluded.  

Radiation Oncology Perspective

Fabio Ynoe De Moraes, MD, PhD, MBA
Fabio Ynoe De Moraes, MD, PhD, MBA

Fabio Ynoe De Moraes, MD, PhD, MBA, Associate Professor of Oncology and Executive Director of the Global Oncology Program at Queen’s University in Kingston, Ontario, Canada, raised the question of what evidence is needed for radiation oncologists and other health-care professionals to act on an AI recommendation or flag. If a model raised a red flag, is that enough reason for a physician to change their treatment plan?  

The threshold for evidence rises with the level of clinical consequence, he suggested. If a model only automates or audits a task, as with auto-contouring, it first needs to demonstrate technical validity and a measurable workflow benefit. With auditing, errors can be identified and corrected. If a model predicts or stratifies risk, such as for pneumonitis or cardiac events, it needs to undergo external validation and calibration and demonstrate a net benefit. At the highest level, models that guide treatment selection—including dose, fractionation, or modality—need to demonstrate comparative clinical utility, improved patient outcomes, and accountability.  

“The more autonomy that we give to the system, the greater the consequences, and the higher the burden of proof,” Dr. Moraes explained. He illustrated how the need for more evidence increases as AI models move from perception to quantification and prediction and, ultimately, to decision-making, especially when an action cannot be reversed once it is taken. “But who is going to decide on the plan?” he asked. “That’s up to us and we must remain accountable.” 

"Who is going to decide on the plan? That's up to us and we must remain accountable."
— Fabio Ynoe De Moraes, MD, PhD, MBA

In addition to autonomy, capability also increases with different levels of AI, including generative and agentic AI. There is currently no clear evidence supporting the use of these forms of AI for decision-making in lung cancer. Generative AI models have been evaluated only in pilot studies, whereas the use of agentic AI in lung cancer decision-making remains conceptual, with no outcomes data. These higher-level forms of AI will also require greater governance.  

Dr. Moraes used AI-generated contouring as an example of how AI can affect clinical decision-making in lung cancer. He explained that AI-assisted contouring can have an impact on clinical outcomes while also saving physicians time. In a prospective multicenter trial of deep learning–based auto-segmentation for organs at risk in thoracic radiotherapy, the median contouring time with manual physician review was 55 minutes vs 10 minutes with AI-assisted review. AI assistance also reduced variability associated with differences in expertise.

Dr. Moraes also stressed the need for local acceptance and emphasized that these models should be evaluated in both difficult and routine cases and assessed for potential consequences related to radiation dose.  

Returning to the question of acting on an AI recommendation or flag, he pointed to a study that used AI to identify the risk of toxicity in patients with locally advanced NSCLC from the RTOG 0617 study. The investigators used Shapley Additive Explanation (SHAP) and AUCs to assess model predictions and feature weighting. Dr. Moraes explained that AUC indicates how well the model predicts, whereas SHAP indicates what drove the prediction; neither, however, establishes causation. 

The generalizability and applicability of the model are also important considerations. “Will this prediction be used in my department? Does the model match my patients and the techniques that I have available?” Dr. Moraes noted. “Does the predicted probability match the observed risk? And what action follows, and who benefits?” 

He emphasized that AI can also aid research by identifying signals that are not typically measured. For example, analyses of data from RTOG 0617 found that AI was able to classify patients by survival and predict toxicity from esophageal radiation doses. 

He also noted that general-purpose large language models (LLMs) are catching up to clinical AI, pointing to a study in which a general-purpose LLM outperformed specialized AI tools on medical benchmarks, including assessments of clinical knowledge and clinical cases.  

The question is no longer whether to use a medical or general-purpose model, but rather which system performs best on the actual clinical task, Dr. Moraes explained.  

He also suggested that clinical prompting can constrain the evidence pathway. Rather than simply asking AI for an answer, clinicians should define what evidence can be used, what must be checked first, and what the model is not permitted to decide. He provided an example prompt for a lung cancer decision brief that included a task, facts, sources, first, output, and limits.  

When considering whether an AI flag or recommendation provides sufficient evidence to change a treatment plan, Dr. Moraes explained that such a flag should prompt a review that includes verifying the inputs, anatomy, and clinical risk factors; validating local calibration, uncertainty, and subgroup performance; and comparing feasible alternatives against established criteria.  

He concluded that an AI model cannot establish that an alternative treatment is better; rather, it can inform clinical reasoning, but its use in treatment decisions requires demonstrated clinical utility. 

He stressed that AI assists the physician in making the decision, but the radiation oncologist ultimately owns the context, trade-offs, and responsibility.  

Surgical Perspective 

Kelvin Lau, BA, MA, MBBS, PhD, FRCS
Kelvin Lau, BA, MA, MBBS, PhD, FRCS

Kelvin Lau, BA, MA, MBBS, PhD, FRCS, Consultant Thoracic and Robotic Surgeon at St Bartholomew's Hospital in London, divided surgical decision-making into two categories: preoperative and intraoperative.  

“We’ve been predicting risk for many, many years, but thoracic surgeons are particularly bad at predicting risk,” Dr. Lau commented. He noted that the Thoracoscore clinical risk model has been widely used for many years but “is a very blunt instrument and it largely only predicts mortality, which is a very rare outcome.”  

Studies have demonstrated that the score does not correlate with observed mortality across different populations. AI has therefore been adopted in an effort to improve risk prediction models for thoracic surgery.  

One such model, Predicthor, outperformed the Thoracoscore risk model in predicting 30-day mortality and complications after thoracic surgery for lung cancer. Another model, which used unsupervised deep learning to predict postoperative pulmonary complications following surgery for NSCLC, achieved an area under the receiver operating characteristic curve of 0.84, “which is quite impressive,” Dr. Lau said. However, he noted that the model’s generalizability is questionable.  

Neither model, however, incorporated data from electronic patient records. He explained that in the United Kingdom, the Barts Health Data Platform provides researchers with access to primary care patient health data, potentially expanding the data available for AI training. “If you can pump that into the learning, it will find unexpected trends, and so this would be much more powerful to use rather than just running it on your existing database,” he commented.  

After data from the CheckMate 816 trial showed that patients who achieved a pathologic complete response were “practically cured of cancer,” with median overall survival not reached regardless of treatment arm, questions remained about whether surgery was still necessary.  

An AI model from China, NeoPred, suggested that AI may be able to predict pathologic complete response based on pretreatment and presurgical CT scans. When expert radiologists used AI assistance, the AUC increased to 0.829 and accuracy increased to 0.820. However, Dr. Lau commented that even an AUC of 0.83 might not be sufficient for a surgeon to choose not to operate.  

Turning to decision-making during surgery, including robotically assisted surgery, Dr. Lau explained that while the robot is essentially a remote-controlled arm, the surgeon is still responsible for the decision-making and for operating the robotic system.  

He explained that digital AI, such as commonly used LLMs, is not effective for performing surgery: “In order for AI to be able to help to do a surgery, you need it to interact with the outside world. It needs to break the jail bars off of the digital world into the real world, rather than taking a text in a prompt. It needs to see where you are. It needs to know how my tools are interacting with you. And it needs to output an action to a robot. So an LLM is not going to do the job. You need a VLA [vision-language-action model].” 

However, Dr. Lau said that one of the biggest barriers to training AI to perform surgical tasks is the lack of an “internet of operations” from which it can learn. Modern robotic systems are better equipped to record surgical actions and capture environmental variables that can be used for training. “These robots are learning and they’re adapting to the environment,” he said.  

However, he cautioned that these robots cannot be trained on rare events, and researchers are not yet sure how to train them to handle such situations without causing harm. Although several digital solutions are being considered—including simulation, oversampling, transfer learning, anomaly detection, and human supervision—the solution may ultimately involve the robot stopping if it predicts that something will go wrong. In such situations, surgeons need to communicate with the rest of the operating team to correct the problem, but coordinating the team’s response requires a different set of capabilities from the robot.  

“So we’re still quite far away, but it is coming,” Dr. Lau said of autonomous robotic surgery.  

He concluded that AI can improve patient selection for surgery and the prediction of postoperative complications and treatment response. However, autonomous robotic surgery requires further development of physical AI, including the incorporation of tool kinematic and force data and methods for enabling robots to respond safely to rare and serious events.  

DISCLOSURES: Dr. Moraes reported receiving prior teaching honoraria from AstraZeneca. No other disclosures were reported.  

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content