FDA-Cleared AI Faces an Evidence Gap
Clinical evaluation has not kept pace with the regulatory authorization of AI-enabled medical devices, according to a review published in PLOS Digital Health.
Researchers identified 1,357 AI and machine learning–enabled devices cleared or approved by the U.S. Food and Drug Administration (FDA) through December 5, 2025. Only 34 devices (2.5%) were linked to registered prospective clinical trials, which included 12 (0.9%) with results posted on ClinicalTrials.gov and 12 (0.9%) with an associated peer-reviewed publication. Only three devices (0.2%) were evaluated using primary outcomes directly related to patient health, such as mortality, major morbidity, functional status, or quality of life.
The findings do not indicate that the other devices are unsafe or ineffective, or that there was deficient regulatory oversight. However, the review highlights the limited publicly available evidence clinicians and health systems can use to assess their clinical effects.
Review Methods
Rawan Abulibdeh, PhD, a postdoctoral research fellow with University Health Network in Toronto, Canada, and colleagues reviewed medical devices listed in the FDA database and the American College of Radiology (ACR) Data Science Institute’s AI Central catalog. They searched FDA summaries, ClinicalTrials.gov, and PubMed for associated trials, posted results, and peer-reviewed publications.
The researchers distinguished between patient-centered outcomes and surrogate measures. Mortality, major morbidity, hospital admission or readmission, functional status, quality of life, and symptom burden were considered patient-centered when specified as primary endpoints. Sensitivity, specificity, area under the curve, and other measures of diagnostic or technical performance were classified as surrogate endpoints.
Among the 34 devices with registered trials, 20 were intended for diagnosis, and seven for screening, three provided treatment or therapeutic guidance, while four performed other clinical tasks. Twenty-eight trials (82.4%) used diagnostic accuracy or another surrogate measure as the primary endpoint, three assessed process measures, such as workflow efficiency, and three evaluated patient-centered outcomes.
This distinction is relevant to oncology, where an AI tool may improve the detection or classification of a radiologic or pathologic finding without necessarily improving treatment decisions, reducing complications, or extending survival.
Radiology Dominates Authorized Devices
Radiology accounted for 1,059 of the AI devices (78%), but only three radiology devices were linked to registered trials. Twelve of 126 cardiovascular devices and six of 62 neurology devices had registered trials. None of the 22 anesthesiology devices had a related trial.
The available studies also consisted of small sample sizes. Nine of the 34 trials enrolled fewer than 100 participants, and 16 enrolled between 100 and 500. Additionally, 23 were conducted only in the United States, 10 were international, and the location of one was unclear. Industry led 32 of the trials; one was academically sponsored, and one involved an academic-industry partnership.
Evidence across patient populations was also limited. Only nine trials reported any subgroup analysis, five examined performance by sex, four by age, and three by race or ethnicity. None reported subgroup findings by language.
Pediatric patients were excluded from almost all trials. Several studies also excluded pregnant women, non-English speakers, and people with mental impairment. These criteria may limit the ability to determine how devices perform among patients encountered in routine clinical practice.
For cancer care, representative evaluation may need to account for differences in tumor type, disease stage, comorbidities, clinical setting, and patient demographics. Performance demonstrated at an academic center may not transfer directly to community practices or other health systems.
Regulatory Clearance and Clinical Benefit
Most AI medical devices reach the U.S. market through the FDA’s 510(k) pathway, which generally requires manufacturers to demonstrate substantial equivalence to an existing device. Independent prospective trials showing improved patient outcomes are not routinely required.
The authors identified several barriers to further evaluation. Prospective trials are costly and must integrate changing software into clinical workflows and electronic health records. AI systems may also be updated while a study is underway, complicating assessment of the version ultimately used in practice.
The review noted that limited public evidence does not mean regulatory monitoring is absent. The FDA uses manufacturer reporting requirements and postmarket surveillance systems. Nevertheless, without accessible clinical studies, clinicians and health systems may have difficulty independently evaluating whether a device improves care.
Proposed Evaluation Framework
The authors proposed a three-stage approach to evidence development. Before clearance, devices would undergo retrospective validation using diverse, representative data sets, with demographic performance reported.
Around the time of clearance, developers would conduct prospective studies in clinical workflows, with prespecified safety and usability outcomes. After clearance, larger multicenter trials would assess patient-centered outcomes and examine performance across groups defined by factors such as age, sex, race, language, comorbidities, and geography.
Additional recommendations included prospective trial registration, public protocols, and wider use of SPIRIT-AI and CONSORT-AI reporting standards. The authors also called for postmarket monitoring of algorithmic drift and bias, clear software version tracking, and defined requirements for revalidation after updates.
The analysis was limited to publicly available FDA records, ClinicalTrials.gov, and PubMed. It may have missed internal, unpublished, or unregistered studies, as well as trials not linked to accessible FDA summaries. The researchers did not compare AI devices with conventional devices or classify the entire cohort by risk class, regulatory pathway, or degree of clinical autonomy.
DISCLOSURES: Funding sources included the National Institutes of Health, National Science Foundation, Korea Health Industry Development Institute and several institutional and competitive grant programs. The authors declared no competing interests.
ASCO AI in Oncology is published by Conexiant under a license arrangement with the American Society of Clinical Oncology, Inc. (ASCO®). The ideas and opinions expressed in ASCO AI in Oncology do not necessarily reflect those of Conexiant or ASCO. For more information, see Policies.