News Research Decision-Making Support Hematologic Malignancies

LLM Agent Shows Concordance With Tumor Board Decisions in Hematologic Malignancies

August 31, 2026 Julia Cipriano 4 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

A locally deployable, modular large language model (LLM) agent, called HemaGuide, provided treatment recommendations with a traceable rationale across hematologic malignancies that maintained concordance with tumor board decisions across institutions and under real-time conditions, according to findings published in Nature Medicine.

Late-line hematologic malignancies bring together many of oncology’s most difficult challenges, requiring decisions that account for long, sequence-dependent treatment histories, therapeutic windows narrowed by cumulative toxicity, and rapidly evolving standards of care. At the same time, access to subspecialty tumor board deliberation may be uneven.

“We show that a retrieval-first, workflow-aligned approach by grounding LLM outputs in curated guideline flowcharts, a contemporary case memory, and structured molecular interpretation can convert unstructured clinical documentation into auditable treatment recommendations…,” wrote study author Mirco J. Friedrich, MD, PhD, of the German Cancer Research Center, Heidelberg, and colleagues.

Study and Model Details

HemaGuide is an openly available clinical decision support architecture with a four-step workflow: case extraction from routine clinical documents, enrichment and routing, skill mode selection, and evidence aggregation into an auditable treatment recommendation.

In practice, the workflow begins by converting routine unstructured tumor board briefings or case summaries into semistructured case representations (nine clinical sections; optional molecular layers). It then performs contextual enrichment and autonomously routes cases to one of three specialized decision modes:

  • Guideline mode uses disease-specific clinical guideline and standard operating procedure flowcharts;

  • Advanced mode retrieves clinically similar longitudinal cases from a clinical memory of more than 2,000 real-world tumor board discussions spanning leukemias, lymphomas, and plasma cell dyscrasias, along with targeted literature;

  • Molecular mode performs variant interpretation and targeted literature retrieval.

The resulting context is aggregated to generate a conference-style treatment recommendation with a verifiable rationale.

HemaGuide’s model-agnostic architecture allows different LLMs to serve as its backbone, with the investigators evaluating the architecture across six foundation models (gpt-5-mini, gpt-5-nano, gpt-oss-120b, gpt-oss-20b, qwen3-next-80b, and qwen3-32b). This flexibility enables deployment on local infrastructure and allows model choice to be optimized for cost, latency, and regulatory constraints without altering clinical logic.

Performance and Validation

The investigators benchmarked the six LLMs both as plain models and as backbones within HemaGuide, using 45 high-complexity hematologic malignancy cases. Blinded expert hematologists evaluated outputs across six dimensions, including concordance with tumor board decisions, which was found to improve with HemaGuide. A systematic ablation analysis across 11 levels, ranging from a plain LLM to the full agent with autonomous routing, showed that concordance gains were routing type–dependent, with no single architectural component sufficient across all case types.

For 70 clinically relevant missense variants, HemaGuide’s automated classifications showed high concordance with expert standards, with no oncogenic variant downgraded to benign. The full workflow had a median latency of 39 seconds on commodity hardware under real-time conditions, compared with hours typically required for manual molecular board workflows, according to the investigators.

A small simulated practice study found that HemaGuide-assisted resident physicians achieved near–senior physician level concordance and, in some comparisons, performed better than senior physicians assessing cases within their primary disease expertise.

In external validation at a second academic center, HemaGuide achieved 81.8% concordance across 555 independent cases spanning 47 distinct hematologic entities. Concordance was 82.8% in a prospective 1-month silent trial that included 64 consecutive, unselected cases. Across these cohorts and the benchmarking cohort, hallucinations were identified in 2 of 664 evaluated cases (0.3%; both in the external validation cohort).

The investigators concluded, “Future studies are required for multicenter validation, formal evaluation under uncertainty, electronic health record integration, and demonstration of improved clinical outcomes; however, these results support a practical path toward locally deployable decision support that complements expert tumor board deliberation on commodity hardware, at a cost and latency compatible with routine preconference preparation.”

DISCLOSURES: For full disclosures of the study authors, funding information, and data and code availability, visit nature.com.

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content