News Research Decision-Making Support Gynecologic Cancers

Head-to-Head Benchmarking of Three LLMs in Endometrial Cancer Decision-Making

October 08, 2026 Julia Cipriano 5 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Gemini demonstrated greater concordance with expert recommendations and fewer major safety issues than ChatGPT and Claude in postoperative endometrial cancer decision-making, although none of the large language models (LLMs) performed well enough for autonomous clinical use, according to findings reported by co-first authors Emanuele Perrone, MD, and Giuseppe Parisi, MD, both of Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Rome, and colleagues in JCO Clinical Cancer Informatics.

Describing the models’ alignment with practice guidelines as “inconsistent and potentially unsafe,” Frank Po-Yen Lin, MBChB, PhD, FRACP, FAIDH, the journal's Deputy Editor, wrote, “Specialist oversight is mandatory before any LLM output informs multidisciplinary oncological decision-making.”

Study Details

The case-based in silico benchmarking study included 35 fabricated postoperative endometrial cancer vignettes developed by two expert gynecologic oncologists to represent a broad range of management scenarios across the 2025 European Society of Gynaecological Oncology (ESGO)–European Society for Radiotherapy and Oncology (ESTRO)–European Society of Pathology (ESP) risk groups. The two experts also determined the guideline-based reference standard for each vignette, including the preferred postoperative management recommendation.

Between March 21 and 23, 2026, each vignette was submitted to OpenAI's GPT-5.3 Instant, Google’s Gemini 3.1 Flash-Lite, and Anthropic’s Claude Sonnet 4.6 in a new, independent chat session using the same standardized unassisted prompt. The models were asked to assign the FIGO 2023 stage, identify the ESGO-ESTRO-ESP 2025 risk group, and recommend postoperative management.

Concordance with the prespecified reference standard was the primary endpoint, with responses scored on a three-level ordinal scale as discordant (0; management recommendation did not align with the reference standard), partially concordant (1; correct management despite incomplete agreement in staging assessment), or fully concordant (2; correct identification of the reference postoperative management, together with the corresponding stage and risk group). Model responses were independently assessed by two experts blinded to model identity, with scoring disagreements adjudicated by consensus with a third expert.

Secondary endpoints included major safety issues—defined as recommendations plausibly leading to clinically meaningful overtreatment, undertreatment, or otherwise inappropriate management—and recognition of missing critical information. In a post hoc subgroup analysis of 12 vignettes, in which none of the models initially achieved full concordance, the models were retested after being provided with the full ESGO-ESTRO-ESP 2025 guidelines. 

Key Findings

Significant differences in concordance were observed between the models (P < .001). The median concordance score was highest with Gemini at 2, vs 1 with ChatGPT and 0 with Claude. Fully concordant recommendations were generated in 65.7%, 2.9%, and 8.6% of cases with these respective LLMs.

Major safety issues were reported in 22.9% of Gemini responses, 48.6% of ChatGPT responses, and 54.3% of Claude responses, with a significant difference across models (P = .007). For the subset of six vignettes with intentionally omitted decisive information, Gemini identified the need for additional data in 83.3% of cases; ChatGPT and Claude each did so in 50.0%.

Guideline-informed prompting appeared to significantly increase concordance for Gemini and Claude in the post hoc subgroup analysis, with median scores increasing from 0 to 2 for both models (P < .001 and P = .008, respectively); no significant change was observed with ChatGPT (P = .312). The number of cases with major safety issues decreased from 8 to 0 with Gemini (P = .008) and from 10 to 4 with Claude (P = .031) but remained at 5 with ChatGPT.

Opportunities and Takeaways

“Although [Gemini 3.1 Flash-Lite] achieved higher concordance and fewer major safety issues than [GPT-5.3 Instant] and Claude Sonnet 4.6, the observed frequency of discordant and clinically unsafe recommendations precludes support for any autonomous clinical use,” the investigators concluded. “These findings suggest that current LLMs may have a role as supervised adjunctive tools, but not as substitutes for expert multidisciplinary decision-making.”

The findings should be interpreted in the context of several limitations. Fabricated clinical vignettes cannot fully reflect the complexity of real-world multidisciplinary decision-making, including the systematic consideration of patient-level factors such as performance status, comorbidities, and treatment preferences. The relatively small number of vignettes also limited the precision of subgroup analyses. In addition, the study evaluated publicly accessible base models rather than more advanced, optimized versions, and the use of independent chat sessions prevented cumulative adaptation, contextual memory, and iterative refinement. Finally, because LLMs evolve rapidly, model performance may change over time, even with similar prompting.

The investigators called for future research using larger case sets, independent expert panels, and real-world multidisciplinary data that incorporate patient-level factors. They also recommended evaluating LLMs for more granular treatment-planning decisions, while emphasizing the need for clear governance standards, transparent disclosure policies, and continued expert oversight to mitigate the potential spread of misinformation.

DISCLOSURES: Dr. Perrone has served as a consultant or advisor for Olympus Medical Systems. Dr. Parisi reported no conflicts of interest. For full disclosures of the other study authors, visit ascopubs.org.

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content