News Research Decision-Making Support

Experimental Study Finds Physicians Susceptible to Causal Illusions From AI-Generated Patient Classifications

August 10, 2026 Meg Barbor 8 min read
Share Share via Email Share on Facebook Share on LinkedIn Share on Twitter

Physicians in two experiments often followed incorrect AI-based patient classifications, even after receiving outcome feedback that contradicted those classifications. 

In the study, published in PLOS Digital Health, researchers tested how professional physicians responded when an AI system classified fictitious patients as either highly or lowly sensitive to a treatment. The AI classification was wrong by design—in both experiments, the label did not predict which patients would benefit from the treatment. Patients labeled highly sensitive were no more likely to benefit than patients labeled lowly sensitive.

The physicians’ task was not to diagnose a real disease or select among real therapies. Instead, the authors used a fictitious rare disease and a fictitious treatment so that participants’ prior clinical knowledge would not influence the results.

Aranzazu Viñas Gomez, PhD, an adjunct professor in the Department of Economics and Management, University of the Basque Country, Spain, and the corresponding author of the study, said the findings suggest that physicians, like other people, may be susceptible to illusions of causality: the belief that one event causes another when it does not.

“Despite their experience, physicians also seem to have a tendency to develop causal illusions in their professional practice,” Dr. Viñas Gomez said in an interview with ASCO AI in Oncology. “This can have negative consequences for some patients.”

She added that the findings do not mean clinicians should avoid AI-based tools, but they do emphasize the importance of training and workflow design when these tools are used in clinical decision-making.

How the Experiments Worked

Both experiments recruited physicians through the Prolific survey platform. Participants were required to speak fluent English and to work as doctors in the health-care sector. The authors excluded participants who answered control questions incorrectly.

In Experiment 1, 120 physicians were recruited, but 15 were excluded after failing control questions, leaving 105 participants. The final cohort included 58 women and 47 men; mean age was 38.6 years, mean professional experience was 13.1 years, and the most common specialties were general medicine and pediatrics.

In Experiment 2, the authors recruited 121 physicians; after 3 were excluded for failing a control question, 118 were included in the analysis. The final sample had a mean age of 36.6 years and mean professional experience of 10.6 years; the most common specialties were general medicine and internal medicine.

At the beginning of each experiment, participants were told to imagine they were doctors treating a rare disease called Lyndsay syndrome. The treatment was still under development, and its effectiveness had not yet been proven. Participants were also told that an AI agent had classified patients as either highly sensitive or lowly sensitive to the treatment. A highly sensitive patient, according to the classification, would be expected to improve with higher probability after receiving the treatment than a lowly sensitive patient.

The authors gave participants no additional information about the AI classification system or the fictitious patients, such as comorbidities, disease trajectory, or patient history. According to the authors, this was meant to ensure that the only information available to participants was the AI label, the physician’s treatment decision, and the patient outcome shown after that decision.

In each experiment, participants saw 60 fictitious patients in random order: 30 classified by the AI as highly sensitive and 30 classified as lowly sensitive. For each patient, the physician decided whether to administer the treatment. Immediately afterward, the screen showed whether the patient healed.

The authors measured how often physicians administered the treatment to each patient group and how effective they judged the treatment to be for each group after all 60 patients had been seen. Effectiveness judgments were made on a scale from zero (completely ineffective) to 50 (moderately effective) to 100 (completely effective).

Experiment 1: Treatment Was Effective for Both Groups

In the first experiment, the treatment itself worked, but the AI classification did not. Patients labeled highly sensitive were no more likely to benefit from the treatment than patients labeled lowly sensitive.

The actual healing rates were the same in both AI-labeled groups. Among patients who received the treatment, 70% healed. Among patients who did not receive the treatment, 20% healed. That meant the treatment improved the chance of healing by 50 percentage points, regardless of the AI label.

If physicians had learned from the outcome feedback, they should have judged the treatment as equally effective in the two groups. But they did not. They gave the treatment more often to patients labeled highly sensitive than to those labeled lowly sensitive, with mean treatment-administration probabilities of 0.893 and 0.557, respectively.

They also rated the treatment as more effective in patients labeled highly sensitive. The mean effectiveness judgment was 69.5 for patients classified as highly sensitive and 51.0 for patients classified as lowly sensitive, even though the correct judgment was 50 for both groups.

The perceived reliability score was above the midpoint, meaning physicians generally viewed the AI classification as somewhat reliable. But perceived reliability and general attitudes toward AI did not significantly influence treatment decisions, so the authors did not include those variables in the final analysis.

Experiment 2: Treatment Was Ineffective for Both Groups

The second experiment used the same basic task but changed the true effectiveness of the treatment.

In Experiment 2, the treatment did not work for either group. Seventy percent of patients recovered whether or not they received the treatment, so treatment did not increase the probability of healing. The correct effectiveness judgment was therefore zero.

The authors described this setup as a laboratory model of pseudomedicine because the treatment was presented as potentially useful but had no relationship with recovery. They used a high rate of spontaneous recovery because previous research has shown that causal illusions are more likely when recovery is common.

The AI label was still wrong by design, but physicians again treated the two AI-labeled groups differently. They administered the ineffective treatment more often when the AI had labeled patients highly sensitive than when it had labeled them lowly sensitive, with mean treatment probabilities of 0.783 and 0.368, respectively.

Physicians also rated the treatment as effective, even though it had no effect on recovery. On average, they rated the treatment’s effectiveness as 67.6 in patients classified as highly sensitive and 47.2 in patients classified as lowly sensitive. Both values were significantly higher than the correct value of zero.

The authors wrote that physicians trusted the AI classifications and did not use the available outcome information to override the erroneous labels. In Experiment 2, they added, physicians also failed to recognize that the treatment was completely ineffective.

Training, Workflow, and Study Limitations

Dr. Viñas Gomez said training and workflow design may be needed in addition to human oversight when clinicians are using AI-based classification or decision-support tools.

“Training and workflow design are basic,” she said. “Training is essential to make professionals aware of the importance of critical thinking, the pros and cons of AI, and how AI can affect human decision-making.” Workflow design may also matter, she added, because it can prompt clinicians to stop and think before making a decision.

The authors emphasized several limitations. The physicians’ professional status was self-reported, although the researchers used Prolific’s professional filter and additional screening questions. The disease and treatment were fictitious, the task was simplified, the treatment decision was binary, and participants did not have access to the kind of clinical information physicians usually use.

The AI system was also not described in detail. Participants were not told whether it had been validated or was still experimental. The authors noted that trust and decisions might differ depending on how an AI system is presented.

Dr. Viñas Gomez said the experimental design was useful because it allowed the authors to control for outside variables.

“Experiments have the benefit that they allow us to control for extraneous variables, thus, in this case, helping to make the causal claim that it is the misclassification of patients that influences both the decision to use the treatment and the effectiveness judgments,” she said.

But the design also limits how far the results can be applied. 

“We conducted our experiments with a sample of real doctors. But our research is still experimental,” she said. “Precisely because they are experiments, their ecological validity is lower than that of other types of research, and results might be task-dependent.”

According to Dr. Viñas Gomez, more research is needed to test whether the findings hold with real AI-based patient classification systems in clinical settings.

DISCLOSURES: Support for this research was provided by a grant funded by MICIU/AEI and by ERDF A way of making Europe, as well as through a grant funded by the Basque Government. A.V. was supported by Fellowship FPU20/01009 funded by MICIU. The authors declared no competing interests.

ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.

KOL Commentary
Watch

Related Content