Business

NOHARM medical AI study finds omission errors across leading tools

A benchmark of OpenEvidence, Doximity, OpenAI and Anthropic tools found most harmful AI mistakes came from missing information.

Daniel Okafor

By Daniel Okafor · Business Editor

3 min read

NOHARM medical AI study finds omission errors across leading tools
Photo: Fortune

The NOHARM medical AI study found that leading clinical AI systems still make potentially harmful mistakes, with omissions accounting for most errors across every tool tested. The finding matters because doctors and health systems are adopting AI assistants while regulators and courts are still deciding how much oversight is enough.

The independent benchmark tested Doximity Ask, OpenEvidence, OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5, according to the study. Researchers from Stanford, Harvard and the ARISE network built the benchmark and ran 1,100 real clinical cases through each system, then used about 13,000 physician annotations to score the outputs for possible patient harm.

What did the NOHARM medical AI study find?

Doximity Ask ranked first in the benchmark, according to the NOHARM study. OpenEvidence disputed the accuracy of Doximity’s score, with CEO Daniel Nadler telling Fortune by email that he did not believe the methodology would pass peer review and noting that the study itself had not been peer-reviewed.

The broader result was less about the order of the ranking than the type of failure the systems shared. Across all AI systems tested, 76.6% of harmful errors were omissions, according to NOHARM, meaning the AI left out relevant information rather than making a false statement.

An omission error in clinical AI is a missing piece of information that could affect care. That distinction is important because an answer can sound correct while still failing to include something a physician needed to consider.

Eric Topol, a cardiologist, Scripps Research scientist and co-chair of Doximity’s PeerCheck program, told Fortune that omission errors need to be brought as close to zero as possible. He said current medical AI systems can create an “illusion of readiness,” even as the technology improves.

NOHARM also found that doctors using AI delivered better care than doctors without AI, according to the study. That puts the debate in a harder place: the tools may help clinicians, but the remaining errors can be difficult to spot if the answer appears complete.

How OpenEvidence and Doximity are using medical AI

OpenEvidence, founded in 2021, offers doctors a free, ad-supported AI search engine that draws from peer-reviewed medical journals and labels the strength of evidence, according to the company information cited by Fortune. Investor demand has risen quickly: a Sequoia-led round valued the company at $1 billion in February, followed by reported valuations of $3.5 billion in July, $6 billion in October and $12 billion in January.

CNBC reported that OpenEvidence raised about $700 million over roughly a year, including a round co-led by Thrive Capital and DST Global. Fortune described the company as a prominent example of the medical AI boom.

Doximity, best known as a professional network for physicians, has built its AI business around enterprise contracts. Its Ask assistant helps doctors summarize patient notes, check drug interactions and draft documentation, and Doximity says it is included in paid contracts with more than 150 health systems.

Doximity says Ask uses a human-review layer called PeerCheck, in which physicians compare AI outputs with the cited original sources. The company reported $145.4 million in quarterly revenue this spring, up 5% from a year earlier, according to StockAnalysis data cited by Fortune.

Why regulators are watching clinical AI

The study arrives as oversight of medical AI remains unsettled. KevinMD reported that the FDA loosened its position in January on AI-powered clinical decision-support tools, allowing more room for use when doctors can independently check the system’s reasoning.

State lawmakers have moved in a more restrictive direction. The Transparency Coalition reported that states passed more than a dozen healthcare AI laws in 2026, many of them requiring human signoff before an AI-assisted decision reaches a patient.

Liability remains unresolved. KevinMD reported that courts are still sorting out whether doctors, hospitals or AI vendors are responsible when an AI recommendation is wrong, a question likely to shape how hospitals buy and deploy these systems.

This story draws on original reporting from Fortune.