News Features Decision-Making Support Pancreatic Cancer Operational Efficiency

An Academic Lab's Autonomous EHR Agent, and the Case for Deploying It in Oncology

August 18, 2026 Meg Barbor 11 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

For years, many medical AI tools have been built to answer one question at a time: read this scan, classify this image, summarize this note, respond to this prompt. MIRA was designed to test a different idea—whether an AI system could work through a patient case more like a clinician does.

Jakob Nikolas Kather, MD
Jakob Nikolas Kather, MD

In a recent Nature article, Jakob Nikolas Kather, MD, and colleagues described MIRA, short for Medical Intelligence for Reasoning and Action, as an autonomous AI agent operating in a sandboxed electronic health record environment. In simulations using real patient cases from the MIMIC-IV database, MIRA could take patient histories; order and interpret laboratory, imaging, and microbiology tests; generate differential diagnoses; and formulate treatment plans, including medications, procedures, and admission decisions.

In the controlled setting, MIRA performed at or above physician level on several measures, including diagnostic accuracy and treatment quality. But the authors did not frame the system as one to replace clinicians. They emphasized that prospective, real-world studies are needed to establish generalizability, safety, and governance.

ASCO AI in Oncology spoke with Dr. Kather, who holds the W3 Professorship for Clinical Artificial Intelligence at TU Dresden and serves as a senior physician in medical oncology at University Hospital Dresden in Germany, about studying MIRA's performance, where autonomous agents might fit in oncology, and why evidence-based medicine still has to guide clinical AI.

What led you to develop MIRA?

I run an AI research lab, and for many years we have been developing AI tools for oncology. We have built tools for image analysis, diagnosis, prognosis, and prediction of treatment response using radiology data, pathology data, and other kinds of data.

What has changed recently is that the technical capabilities in AI have advanced tremendously. Of course, there are chatbots that everybody is using, but there is also this newer paradigm called AI agents, where large language models can be used to perform more complex procedures.

AI agents can go beyond what we have done before. They are being used in many other domains, and we wanted to systematically apply these AI agents to the clinical workup of patients. That became MIRA, and the system worked very well.

Tell us more about the study, and what results surprised you most in the sandbox experiment.

We built this agent, which we called MIRA, and gave it real patient cases in a sandboxed setting. Typically, a human doctor would look at the case, start the workup, order laboratory tests and imaging tests, make a diagnosis, and make a recommendation in the end.

We built a system that allowed a large language model in this agent platform to basically do the same thing, and it worked very well. It made very reasonable choices. We benchmarked it extensively against human doctors, and we found that it performed essentially on par.

We also reviewed the agent’s actions afterward with human experts. It made some mistakes, of course, but many of the mistakes were debatable—things that were equally valid compared with what we had designated as the ground truth. It almost never made severe mistakes or severely harmful recommendations. From the start, it performed very well and appeared safe, which was very encouraging.

Why did you include pancreatic cancer among the eight conditions tested?

We tested MIRA on very common cases for the first presentation of patients, for example in the outpatient clinic or emergency department. Many of these cases involve symptoms such as abdominal pain or shortness of breath, and statistically, very few will ultimately turn out to be cancer.

We tested pancreatic cancer cases, and the model performed decently. But what we have to do, and what we are doing now, is take AI agents like MIRA and specifically test them in oncology.

Last year, before the MIRA system, we had our agentic tumor board, which was technically a little simpler. There, we already found that AI agents are very good at making recommendations for patients with cancer. The next step is to bring this all together.

What would need to happen for an agent like MIRA to be applied to and useful in oncology?

In an oncology setting, agents have tremendous potential. There are many things that we do in oncology that could and should be agentified, including administrative and clinical procedures.

What we really have to do is look at the real world: What are clinicians doing in oncology, and what do they need help with? Some tasks could be handled completely by autonomous agents. For example, one thing we oncologists have to do here in Germany is review chemotherapy orders for the coming days and then go through them one by one and clear them. I think that is a task that is ripe for automation by agents, of course with some level of human oversight.

In other cases, agents could serve as a second opinion or provide an initial evaluation of a case. What needs to happen is that we identify these situations by really looking at the real world of oncology. The main barriers are practical ones, especially the interoperability of computer systems, which are currently not ready to deploy these agents.

I would not say oncology is necessarily harder. I think it could actually be the first real clinical application of these agents, because many parts of oncology are quite standardized. We know how to do them, and we know how to do them well. Oncology is actually a very good test ground for AI agents in health care.

How do you view MIRA in relation to Google’s AMIE system?

Of course, we know the team at Google, and we have interacted with them in many ways, but we were not aware that this work was coming out at the same time.

I think these papers are somewhat similar, but also complementary. For us, it was also very encouraging that we are a small academic lab with almost zero funding, and that it is still possible from academia to have novel ideas, validate them, and drive the field forward. That is very rewarding.

Of course, they take slightly different approaches, but the common theme is that there are three steps in the evolution of AI: The first step was using AI to read images and give us a result. That is a very simple, narrow task. The second is chatbots. You can talk to a chatbot and ask any question, but you ask one question and get one answer. The chatbot will not autonomously do anything without you explicitly prompting it. The third step is multistep approaches. That is essentially what we do in medicine. We do not ask just one question of our patient. We have a multistep process in which we ask many questions, get many responses, and consult external information sources. 

It is very encouraging that we now have AI tools that can engage in the same approach. I think these AI systems—AMIE, MIRA, or whatever comes next—could really help us in our real-world work.

What are the next steps before systems like MIRA can be used in clinical care, and what limits would you place on autonomous systems like these?

We are working hard to bring these autonomous systems into clinical application. We want to deploy these systems in clinical routine and run prospective trials where we evaluate their performance and safety. I think that is the most important thing we have to do right now.

The main barriers are not the technological limitations, legal limitations, or regulatory issues. The main problems are the digital infrastructure we are facing in hospitals. It is fragmented, and in many cases it is not accessible to self-developed software or external software.

We have to invest in healthcare infrastructure and make it modular, interoperable, and modern, so that we can ultimately apply these agents. And of course, we have to be fast. We cannot wait 1 or 2 years, because even right now, oncologists and patients are using commercial tools on their phones. They are putting cases into ChatGPT and asking it for recommendations.

In many cases, that is probably reasonable, but it is still not ideal to use consumer software for medical decision-making for many reasons. We hope to provide real, validated, official software solutions. Then we can also make them, to some degree, autonomous to help us cope with the workload in oncology.

Where should autonomous agents fit within the broader spectrum of AI assistance in clinical care?

The main questions are how and where we deploy this. Are these systems that will run in the background and prepare things for humans to sign off on? Will we give them more autonomy to do things on their own? Or will they just give us suggestions that run in parallel to human-centered workflows?

All of this still needs to be figured out, and for that, we need the oncology community. Given that this technology exists now, what are the most valuable applications?

Do you find that physicians are generally accepting of the idea of autonomous agents in a clinical setting? 

We are in dialogue with clinical colleagues and with patients, and everybody is very proactive and optimistic.

My oncology colleagues come to me and ask, “Why do I still have to do these things manually? Why can we not just use an agent?” When I am planning my holidays, I can have an agent plan the whole trip through Italy for me, so why do I still have to do this here?

I give talks all the time at oncology conferences, and what we hear is that people want this.

People want this technology to help them. We have a mismatch between the workload and the number of oncology professionals, and that mismatch is going to increase in the next few years. I think everyone knows that we need to optimize how oncology works.

One part of that is working more efficiently across different centers and improving collaboration between the ambulatory and inpatient sectors. But another part is technical innovation. Everyone who has ever used a chatbot knows these systems are incredibly intelligent and can be very helpful if we embed them correctly in our workflows.

How should evidence-based medicine apply to autonomous AI systems?

Evidence-based medicine is super important. If we look back over the last couple of decades, we have cured some cancer types and increased survival for the vast majority of patients with cancer. There has been so much progress in medicine, and all of that has only been possible because of evidence-based medicine, because new approaches have to go through trials and then be adopted.

I think we should keep that. We should uphold this in the age of AI. It is difficult because AI is moving much faster than trials. If you run a prospective trial for an AI system, by the time the first patient is in the trial, the AI system is outdated. So, we have to be much more agile.

But we cannot just blindly trust the technology. We have to run these trials. They might have to be small, pragmatic, and maybe noninterventional, but we need evidence, and we need to uphold evidence-based medicine in this age of AI. I think that is very, very important.

DISCLOSURES: Dr. Kather reported consulting services for Bioptimus, Panakeia, AstraZeneca, and MultiplexDx; stock ownership in StratifAI, Synagen, and Ignition Lab; an institutional research grant from GSK; and honoraria from AstraZeneca, Bayer, Daiichi Sankyo, Eisai, Janssen, Merck, MSD, BMS, Roche, Pfizer, and Fresenius.

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content