Why More, and Better, Benchmarks Are Needed for Clinical AI
As AI is increasingly being tested in clinical scenarios, the question over how best to evaluate the performance of an AI model has led to significant debate and disagreement. AI has already demonstrated potential benefit across a variety of use cases in the health-care field, yet a model that shows strong performance in testing has not always translated to clinical benefit when used in practice.
As Ethan Goh, MD, Executive Director of the ARISE (AI Research and Science Evaluation) Network and Program Director of the Healthcare AI Leadership Program at Stanford University, explains it, evaluating clinical AI is complex because models have different capabilities and are used across different clinical use cases. A model that performs well on a medical knowledge exam may still struggle with another task, such as interpreting an image, while knowledge-exam-based evaluations may miss subtle safety issues, including errors of omission. Different capabilities, he argues, require different benchmarks and different approaches to evaluation.
ARISE is a research group based primarily out of Stanford and Harvard labs focused on evaluating and validating clinical AI tools, including their safety. The group grew out of work comparing physicians using traditional reference materials, physicians using AI, and AI alone, and has since expanded to focus on questions around human-AI performance and clinical AI safety. Today, the group runs multicenter studies and builds benchmarks that evaluate medical AI tools in relation to their true functionality and clinical utility.
In an interview with ASCO AI in Oncology, Dr. Goh discusses why clinical AI requires multiple benchmarks; what errors of omission show about current approaches to safety testing; why specialized clinical tools may retain advantages for specific use cases; and how benchmarks could help researchers, physicians, regulators, and patients understand where AI tools should, and should not, be used.
Can you tell me about ARISE—its mission and how it started?
ARISE is a research group based primarily out of Stanford and Harvard labs. We started as physician-scientists and technologists wanting to study, evaluate, and validate clinical AI tools. We started back in 2022. I think ChatGPT was just beginning to become very popular, and I was like, all right, this could actually be really, really helpful in a lot of our clinical work.
We won a grant from the Gordon and Betty Moore Foundation to do a series of randomized controlled trials testing physicians using traditional reference text vs physicians using AI. As a side thing, we also tested AI alone. I think what was most surprising was that the AI alone did better than even the physicians using AI. The physicians using AI also did better than physicians using non-AI tools.
The three of us—myself; Adam Rodman, MD, MPH, FACP, at Harvard Medical School; and Jonathan Chen, MD, PhD, of Stanford University—just had so much fun working together. Those studies have obviously spun off into a lot of interesting follow-up questions and grant funding that we've started working on. One is, for example, instead of just measuring the model alone, how do we measure physicians using AI? Because we want to improve that. Unfortunately, studies today have mostly found that AI alone outperforms the physician using AI.
Other directions we are focused on include clinical AI safety. How do we know these tools are safe when used by both physicians and patients?
One way we're different from traditional academic groups is that we try to learn from some of the more frontier AI labs. They are very good at not just doing the research, but science communication, whether that's through design, interactive products, explainers, or social media. That's something that I think we spend a bit more time and resources on.
Why is it so hard to evaluate clinical AI? Is it simply that we haven't landed on the best way to do it yet?
One, it's because the technology is so new, so, evaluate it on what? There are so many different evolving capabilities. That's why it's science, right? It's the scientific process and improving ways to validate.
Two would be because there are so many different types of clinical AI tools. We talk about clinical AI as if it's a monolith—just one thing. But no. Do you mean a medical reference clinical AI tool like OpenEvidence? Or do you mean a clinical AI tool put in front of patients, like ChatGPT Health?
Even within that, I think each of the products can rightfully claim that they are used differently. A physician could use both OpenEvidence and ChatGPT, but they could use them in different ways. A lot of that today is based on the product claims that it does well; a lot is branding and marketing. But, there are so many different types of tools and use cases. We haven't even begun to explore all of them, because it's likely that all these interactions and types of users can make a difference in the collective performance.
ARISE currently has several different benchmarks, including the recent Medical AI Superintelligence Test (MAST). Why are so many benchmarks needed? Could we eventually reach a point where only one is needed for all of clinical AI?
We need many, many benchmarks. Different benchmarks test different capabilities. Imagine if you went to ChatGPT with an X-ray and your medical question was, “Do I have a fracture, and what should I do next?”
It's actually two things. One, you're testing whether the AI can recognize and interpret an image. And, you're testing whether the AI can answer a clinical reasoning and knowledge question. So there are actually two different capabilities being tested there. Often benchmarks only test one component. Is it good at text reasoning, medical knowledge exams, understanding different drug formulations?
But what we're increasingly finding is that just because a model is good at text on the MAST benchmark—they can still be terrible at image interpretation. So I tell my friends: If you want to ask a text-based medical question, it's likely going to be correct, although it's important to be aware of some of the harms because our NOHARM study, led by David Wu, MD, PhD, and Fateme Nateghi, PhD, found that they still occur more than 20% of the time.
But if you're going to upload a medical image now, it's still very bad. Often it can't even tell that it's a left-hand X-ray from a right-hand X-ray, as one example. So, there's still some catching up to do.
Remember how these AI models are trained. They're ingesting a lot of data sources, and predominantly that has been more medical text than images. Over time, though, they start getting better at images as well because model developers realized that's where they're weak, and wanted them to get better at reading and interpreting images.
That said, I do think companies are already rolling out products to patients and physicians that are not tested as robustly as one would hope. Therefore, I think it's very, very important to test and break down all these advertised product capabilities. Someone could say, “Look, this model got 100% right on the USMLE medical knowledge exam that students and residents have to take.” But that is not the same thing as being good at interpreting an X-ray image.
Where could benchmarks fit into the regulatory process, and how should regulators consider them?
I think they are one signal. They shouldn't be the only signal, though, because benchmarks have many limitations. Think of it like an exam. Just because a medical student scores 95% on an exam doesn't mean the medical student is immediately ready to see patients.
Now, does that make them useless? No, it doesn't. What it does mean is that if benchmarks were eventually deployed to guide a regulatory position, they would need to be upstream of deployment.
Let's say you want to roll out ChatGPT Health, and it has an imaging capability as well as text reasoning. If we had a benchmark that said, in text reasoning it achieves 95%, and maybe for images it does only 55%, and this is where it gets it wrong, that might cause some change in thinking about how it's regulated. At least it's a good signal to model developers, to physicians, and to patients using these tools to understand what's good vs what's weak. Benchmarks give us a common, reproducible signal to test on.
Looking at the NOHARM paper, one of the most striking findings was that most errors were errors of omission. Why was that finding so important, and why are these errors so dangerous?
It's dangerous because it's so subtle. Imagine you're following a chicken soup recipe. It looks 80% or 90% correct, but the model misses out on one really important thing. That's way harder to identify than if you're making chicken soup and it tells you something silly, like, “Add cheese,” instead of the chicken. That's definitely wrong.
The reason this evaluation is important is because, to date, the validation and evaluation of whether clinical AI tools are safe has been predominantly based on knowledge exams. The thinking is, because I scored well, I know a lot of medical information or drug information, so I'm safe. But that is not the same, and that's what I think we set out to prove and test.
If you are relying more on a knowledge exam, you are going to be able to pick up more errors of comission. You pick up something that is said that the AI should not have said. You've missed out on all these omissions—things that should have been said but were missed.
In your study, clinical retrieval-augmented generation (RAG) tools performed better than general frontier large language models, although other studies have shown the opposite. What are the implications, and how do you think these findings could affect clinical AI development?
Out-of-the-box general large language models are good at a ton of things, everything from medicine to law to finance. But the more downstream and application-specific you go—is it a physician using it to interpret X-rays? Is it a physician using it to ask for a medical citation?—I generally like the idea of what the popular clinical AI companies (OpenEvidence, Doximity, etc.) are doing, which is to invest more resources into expert annotations, better prompting, harnessing, data feedback loops, and clinical validation.
To the point about why different studies say different things, I think that ties back to, again, benchmarks test different things, just as exams test different things. Can you say a driver's license exam in the U.S. is the same as one in Australia? Cars are a little bit different. They drive on different sides of the road.
There's a whole ton of nuance there, but broadly and directionally, my expectation is that the more clinical validation that goes into these specialized tools—the more testing, the more pulling on accurate sources, the more investment that goes into improving these AI models—the better they should be on that specific use case. I think that should be fairly intuitive.
Don't frontier model developers have more resources to put into their models?
They have more resources, sure, but they're also not as focused on the end user because they're trying to build for lawyers, physicians, medical students, and pathologists. So when there are so many cooks in the kitchen, what happens to the recipe? You may not be able to have that prioritization.
For the longest time I was asking my friends at OpenAI, although they have since corrected this, “Why do you still hallucinate citations? Why is OpenEvidence better at citations?” They were like, “Oh, we know it's a big gap.” But among all the other things they're trying to do, it's just prioritization.
I think increasingly the performance gap between frontier model companies and clinical AI companies will be narrowed. So, it will be dependent on these specialized companies to design a better user experience, designed more around the actual end user.
Take ASCO members: oncologists. Oncologists have very, very nuanced information needs. Let's say you want to interpret a chart, you want diagnostic advice, you want citations. You want it from a specific source. That is very different from what OpenAI or Anthropic would be training its model to do today, which may be more for general physicians than a specific specialty.
There's a school of thought that says because one model eventually has more computing, better teams, better talent, eventually it'll be the one model to do everything. I'm not solely convinced by that idea, because insofar as the ones that build on top of these tools keep using the latest, best models and learn where the weaknesses are and patch them, they can move just a bit faster. Although it's on them to do a lot to continue staying ahead. But it's an exciting, open question at the moment.
Large language models can match or exceed physician output on some clinical reasoning tasks, but they can also break down with uncertainty, changing context, or missing information. What does that tell us about what these models haven't captured yet?
Well, there are two parts to it. One, the big models are so capable that researchers are just trying to catch up. We were surprised that they did so well on a medical knowledge exam 3 years ago. Even the developers didn't know themselves.
A lot of this is because of what's called emergent capabilities, which is this direction that the more data, the more computing power you dump at training a model, the more it's able to abstract new capabilities. Who knew it could become good at composing poems? The same applies to health care.
The other part that both model developers and researchers are highly interested in, is what are the failure modes? We have obviously identified some gaps—for example, harm testing and gaps of omission. That's not really a gap that is closed yet today, and that could be one cause of harm.
There also needs to be so many more benchmarks in different specialties than exists today, which is a direction we are enabling at ARISE via open-sourced tooling (ARISEKit). For example, there needs to be an oncology benchmark. Because we have no idea today how well AI could manage a complex oncology patient. There could be another benchmark for women's health. Maybe in creating those tests, we find out that, actually, it's great at oncology, it's great at women's health. But without having actually done the test yet, it's quite hard to say what else are the specific gaps that models need to close.
Another point is that tying it back to how it's actually going to be used in the real world at the end of the day is the most valuable. How would an oncologist want to use an AI tool? Is it, “Help me brainstorm a diagnosis,” or “Help me review and summarize all my patient's chart notes”? The more we can articulate the exact end-user use case, the more we can test for that, and the more we can articulate the gap and go back to the developers and say, “This is where we need to make them stronger.”
What level of performance would AI need to demonstrate before you would feel comfortable with autonomous AI?
Our group put out a perspective piece recently on measuring progress towards medical superintelligence in Nature Medicine. This could form part of a conversation on how to know if AI is eventually good enough to autonomously manage patients.
Industry has already been using the term ‘medical superintelligence.’ Often what they point to is that studies—and this is true—have shown that the AI does better than a generalist physician doing their test. I think, first, we need to define what medical AI superintelligence is. You can't just be better than a sole generalist physician doing a hard oncology case. You have to be better than teams of the best specialists working together.
Once we have the definition, we can start moving—tracking progress and measuring toward that, because the more we move toward that, the more tasks the AI can take on. And even if, for example, we could test a product, and it is indeed medically superintelligent, it doesn't mean we ever need it to actually treat patients autonomously.
Think of the potential of deploying it on much lower-stakes tasks, such as medication refills (for more information on this subject, see "What an AI Prescription Pilot in Utah Could Mean for Oncology"). If an AI is "superintelligent," even in the rare edge cases it should know when to reason and flag that a physician should come and review the case.
Ultimately, it has to be tied to patient outcomes. I don't care if your AI is medically superintelligent. If you're a cancer patient, did you get the best treatment? Did you improve? That's most important. But outcomes are really hard to track. Randomized controlled trials are hard to do.
I think benchmarks help us measure progress along that axis. The more we can do that, the more we can tell regulators, the more we can tell physicians, the more we can tell patients: Use this AI for that. Nope, don't use that AI for X-ray interpretation. Or yes, use that AI for pregnancy medication side effect questions.
Who should set these benchmarks and ultimately determine where autonomous AI should be used in patient care?
Who's the best judge? I think everyone has a role to play. Developers should make benchmarks, and some do share them. I think that's great because they know best how the tools are being used, and they have the data. Researchers should do it as well because they have the scientific expertise and credibility as independent evaluators.
Physicians, institutions, regulators, policymakers—they all have different considerations and points of view on how these tools should be tested. So, I think the TL;DR ("too long; didn't read") is everyone should, if not create benchmarks themselves, at least start being more vocal and sharing in that scientific process of how we should be measuring and testing these things.
DISCLOSURES: Dr. Goh receives funding from the Gordon and Betty Moore Foundation, Macy Foundation, Stanford Artificial Intelligence in Medicine and Imaging—Human-Centered Artificial Intelligence Partnership Grant, Stanford Bio-X Interdisciplinary Initiatives Seed Grants Program (IIP). Dr. Goh reports consulting fees or other compensation from Google, Hello Heart, Grow Care Inc, and Faculty Connection.
ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.