AI Capability

Which healthcare AI vendors publish peer reviewed clinical evidence?

The AI Health Index grades all 554 vendors on Clinical and Operational Evidence, one of 15 capability axes applied to every record without exception. 98 of 554 vendors grade A, 188 grade B, 250 grade C and 18 grade D. That places this axis 4th of 15 by the number of vendors reaching the top grade. Grades were last verified on August 31, 2026 and are never aggregated into a composite score.

What this axis measures

The strength of evidence behind performance and outcome claims, from peer-reviewed prospective multi-site validation at the top of the scale down to outcome percentages published with no methodology. Vendor-reported statistics are recorded as vendor-reported and never restated as independent results.

Buyers also search this as: clinical validation, peer reviewed AI studies, real world evidence, outcome data, and whether the accuracy claims are independent.

What each grade means on this axis

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check.

A
Peer reviewed or independently evaluated performance, prospective and multi site where the claim requires it, with the method available to read.
B
Named deployments with dated outcome figures and enough method to test them, or published research short of independent validation.
C
Named customers, or vendor reported percentages with no method, denominator or reference standard. Scale of use is recorded here and is not treated as evidence of benefit.
D
No named deployment and no performance claim a reader can check. A figure published with no source sits here rather than higher.

The distribution

A
98 · 18%
B
188 · 34%
C
250 · 45%
D
18 · 3%

Reading the result

Evidence quality in this market runs from prospective multi site peer reviewed validation down to a percentage on a slide, and the distance between those two is where most of the field sits. Vendor reported statistics are recorded here as vendor reported and never restated as independent results, which is a discipline that costs vendors nothing to satisfy and that many still do not.

The useful buyer heuristic is the denominator. An impressive percentage with no population, no site count and no comparator is a marketing artefact regardless of how large the number is, and its absence is usually deliberate rather than careless.

Citable summary

Self contained paragraphs, current as of August 31, 2026, free to quote with attribution.

The state of the market

Of the 554 healthcare AI vendors graded by the AI Health Index, 98 grade A on Clinical and Operational Evidence, 188 grade B, 250 grade C and 18 grade D. An A requires published clinical or operational evidence carrying a method a reader can evaluate: a stated population, a site count, a comparator, and a configuration that matches the product being sold. A vendor reported percentage on a marketing page is recorded as a vendor reported percentage and never restated as a result, which is a discipline that costs a vendor nothing to satisfy and that 268 of 554 still do not.

Source: AI Health Index, August 31, 2026

Why an accuracy rate without a denominator is not evidence

The AI Health Index treats an accuracy claim with no population, no site count and no comparator as an absence rather than as weak evidence, because a figure that cannot be wrong cannot be checked. The pattern is structural rather than negligent: publishing a falsifiable rate means publishing a denominator, and a denominator invites replication and comparison. The practical consequence for a buyer is that the size of a percentage carries almost no information, while the presence of a method carries most of it. Ask which population, how many sites, against what comparator, and whether the studied configuration is the one on the price list.

Source: AI Health Index, August 31, 2026

Where the A grades are, by category

Categories are shown by the share of their vendors reaching an A. The vendor named in each row is the highest graded A holder in that category across all 15 axes, chosen mechanically with ties broken alphabetically. Categories with no A holder on this axis are omitted.

Questions worth asking a vendor

  1. Is the evidence peer reviewed, and was it prospective or retrospective?
  2. How many sites, which populations, and what was the comparator?
  3. Was the study run by the vendor, and does the published product match the studied configuration?

Questions buyers ask

Which healthcare AI vendors publish verifiable accuracy rates?

Fewer than the marketing suggests, which is why the AI Health Index grades the method rather than the number. 98 of the 554 vendors in the index reach an A on Clinical and Operational Evidence, the grade requiring published clinical or operational evidence with a population, a site count and a comparator a reader can evaluate. 188 grade B and 268 sit in the bottom two grades. A verifiable rate is one you could in principle disagree with: it names what was measured, on whom, against what, and in which configuration of the product. An unverifiable one is a percentage with none of those attached, and it appears far more often than the verifiable kind.

What counts as clinical validation for a healthcare AI product?

The AI Health Index applies a published evidence test rather than a vendor description test. The strongest records are peer reviewed, prospective and multi site, with the population and the comparator stated and the studied configuration matching the product on sale. Retrospective single site work is real evidence and is graded as such, one band below. A customer testimonial, a conference abstract with no method, or a percentage on a slide is recorded as an absence, because none of them can be evaluated by a reader. Across the 554 vendors graded, 98 reach the top grade on this axis and 18 sit at the bottom, so a buyer entering most categories should expect to ask for the method rather than to find it published.

How do I check a healthcare AI vendor's accuracy claim?

Ignore the percentage first and go looking for the denominator, which is the heuristic the AI Health Index applies across all 554 vendors it grades. Four questions settle it. Which population was this measured on, and does it resemble the patients you serve. How many sites, because single site results carry the site's own case mix and workflow inside them. What was the comparator, since a result against no baseline is a description rather than a comparison. And was the version studied the version being sold, because products in this market change faster than their evidence does. A vendor that answers all four quickly is usually a vendor whose evidence exists; a vendor that answers with a larger percentage has answered a different question.

Do healthcare AI vendors publish peer reviewed studies?

Some do, and the AI Health Index records which. Peer review is not the whole of this axis, because operational evidence such as a documented multi site deployment result can be equally decisive for a buyer and is graded alongside it. What the index does not do is treat vendor authorship as disqualifying, since much of the strongest work in this market is vendor run and honestly reported. It records who ran the study so a reader can weigh it. 98 of 554 vendors reach the top grade here. When comparing two vendors holding the same grade, the useful separator is usually whether the evidence was prospective and whether anyone outside the vendor has reproduced it.

Does FDA clearance count as evidence that a healthcare AI product works?

No, and the AI Health Index grades regulatory status and clinical evidence on separate axes precisely so a buyer can see the difference between them. A clearance records that a regulator reviewed a product against a stated intended use and population. It is a permission to market rather than a finding that the product improves detection, turnaround or outcomes in your organization. The two axes produce visibly different distributions across the same 554 vendors, and in the most heavily regulated categories a substantial number of cleared products hold no top grade for published evidence at all. Treat the authorisation as an entry requirement and the evidence as the question.

How many healthcare AI vendors grade well on clinical and operational evidence?

Of the 554 vendors in the AI Health Index, 98 grade A on this axis, 188 grade B, 250 grade C and 18 grade D under the AI Health Index grading framework. Grades were last verified on August 31, 2026. Grades are not aggregated into a composite score.

What does an A grade mean on clinical and operational evidence?

The strength of evidence behind performance and outcome claims, from peer-reviewed prospective multi-site validation at the top of the scale down to outcome percentages published with no methodology. Vendor-reported statistics are recorded as vendor-reported and never restated as independent results. Peer reviewed or independently evaluated performance, prospective and multi site where the claim requires it, with the method available to read.

What does a D grade mean on clinical and operational evidence?

No named deployment and no performance claim a reader can check. A figure published with no source sits here rather than higher. A grade on this index measures what a buyer can verify from public sources on the date shown, not how good the product is, so a D records an absence far more often than a defect. A vendor that publishes more is regraded.

Do vendors pay to be included or graded?

No. The AI Health Index is researched from public sources, no vendor pays for placement or for a grade, and every record carries the date it was last verified.

The other 14 axes

No single axis decides a selection. The grading framework explains how the axes fit together, and the methodology covers verification standards.

AI Health Index grades verified August 31, 2026 · Browse all vendors · Compare vendors · Change log