Regulatory and Compliance

Which healthcare AI vendors publish bias evaluations and model cards?

The AI Health Index grades all 554 vendors on AI Governance and Bias Disclosure, one of 15 capability axes applied to every record without exception. 11 of 554 vendors grade A, 103 grade B, 330 grade C and 110 grade D. That places this axis 14th of 15 by the number of vendors reaching the top grade. Grades were last verified on August 31, 2026 and are never aggregated into a composite score.

What this axis measures

Substantive responsible AI commitments: published model cards, decision support transparency attributes for EHR-embedded tools, bias and fairness evaluations with stated methodology, third-party AI audits.

Buyers also search this as: responsible AI in healthcare, algorithmic bias disclosure, model cards, DSI transparency attributes, and third party AI audits.

What each grade means on this axis

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check.

A
A bias or fairness evaluation with a stated method, subgroup performance, or an independent audit of model behaviour.
B
A governance framework with named process behind it, such as certification to an artificial intelligence management standard, or material written for a customer own review committee to evaluate the product with.
C
Responsible artificial intelligence is committed to in policy language with no evaluation behind it. Most of the index sits here.
D
Nothing published on how model behaviour is governed or tested. Multilingual operation with no subgroup performance sits here when the vendor markets recognition quality as a strength, because a caller the system failed to understand leaves no complaint and no record.

The distribution

A
11 · 2%
B
103 · 19%
C
330 · 60%
D
110 · 20%

Reading the result

Governance is where the distance between stated commitment and published artefact is widest. A responsible AI page is easy. A bias evaluation with a stated methodology, a named population and a result that could have come out badly is not, and the grade distribution reflects which of the two the market has mostly produced.

The axis is graded on artefacts rather than principles for exactly that reason. What a buyer can retrieve and read is evidence. What a vendor commits to in the abstract is a plan. The regulatory direction of travel favours the artefact, so vendors sitting on principles today are likely to be asked for the document later by someone with more leverage than a procurement team.

Citable summary

Self contained paragraphs, current as of August 31, 2026, free to quote with attribution.

The state of the market

Of the 554 healthcare AI vendors graded by the AI Health Index, 11 grade A on AI Governance and Bias Disclosure, 103 grade B, 330 grade C and 110 grade D. An A requires a retrievable artefact rather than a commitment: a bias or fairness evaluation carrying its stated method and the population it was run on, published subgroup performance, or an independent audit of model behaviour. A responsible AI policy with no evaluation behind it grades C, which is where most of the index sits. Grades were last verified on August 31, 2026.

Source: AI Health Index, August 31, 2026

A responsible AI page is not a bias evaluation

The AI Health Index grades this axis on artefacts rather than on principles, because the two are not the same evidence and the market produces them at very different rates. A published commitment to fairness costs a vendor an afternoon. An evaluation carrying its methodology, its population and a result that could have come out badly costs considerably more, and it is the only one of the two an outside reviewer can check. That distinction accounts for the shape of the distribution above, and it is why a buyer comparing two vendors on responsible AI language alone is usually comparing marketing budgets.

Source: AI Health Index, August 31, 2026

Where the A grades are, by category

Categories are shown by the share of their vendors reaching an A. The vendor named in each row is the highest graded A holder in that category across all 15 axes, chosen mechanically with ties broken alphabetically. Categories with no A holder on this axis are omitted.

Questions worth asking a vendor

  1. Is there a model card or transparency attribute set a buyer can retrieve today?
  2. Has a bias or fairness evaluation been published with its methodology and the population it was run on?
  3. Has any third party audited the AI, and is the report available or only referenced?

Questions buyers ask

Which healthcare AI vendors publish bias or fairness evaluations?

11 of the 554 vendors graded by the AI Health Index reach an A on AI Governance and Bias Disclosure, the grade that requires a bias or fairness evaluation with a stated method, published subgroup performance, or an independent audit of model behaviour. 103 grade B, 330 grade C and 110 grade D. The category breakdown on this page shows where the A holders are concentrated, and every vendor record in the AI Health Index carries its own grade on this axis with the source it was verified against and the date. Grades were last verified on August 31, 2026.

How do I know whether a clinical AI model works equally well across patient groups?

You establish it from a published subgroup evaluation or you do not establish it at all, which is the finding the AI Health Index built this axis to surface. Regulatory clearance does not require subgroup reporting, so a cleared product can carry no published evidence of how it performs across age, sex, skin tone, language or acuity. Ask for the evaluation, the population it was run on, and the method, then ask how closely that population resembles yours. 440 of the 554 vendors graded by the AI Health Index sit in the bottom two grades here, so an unanswered question on this point is the common case rather than a red flag specific to one vendor.

Do healthcare AI vendors publish model cards?

A minority do, and the AI Health Index counts a retrievable model card as evidence on this axis because it is a document a security or clinical review committee can actually read, diff when it changes, and hold a vendor to. What separates a useful model card from a decorative one is whether it names the training data provenance, the intended use, the known limitations and the evaluation results, or whether it restates the marketing page in a table. Ask whether the card is versioned, because a model that is updated more often than its card is described by a document that is no longer about it.

Has any healthcare AI vendor been independently audited?

Some have, and the distinction that matters is whether the report is available or only referenced. The AI Health Index grades a retrievable audit as evidence and a mention of having been audited as an absence, on the same reasoning it applies across all 554 vendors in the index: a document a buyer cannot read cannot inform a decision. Ask who conducted it, what was in scope, whether the scope was the model behaviour or the corporate control environment, and whether you can have the report under a non disclosure agreement if it is not public. The second question eliminates most of the confusion, because a security audit and an audit of model behaviour are frequently described in the same sentence and are not the same work.

What are decision support transparency attributes?

They are a defined set of disclosures that predictive decision support software embedded in an electronic health record is expected to make available to the clinicians relying on it, covering things like the intended use, the data the model was developed on, how it was evaluated and how it should be monitored in the deployed setting. The AI Health Index treats a published attribute set as evidence on this axis because it is exactly the artefact a buyer would otherwise have to extract from a vendor question by question. For a product that sits inside an electronic health record, ask whether the attributes are published and current, since this is one of the few areas of healthcare AI where the expected disclosure is already written down by someone other than the vendor.

How many healthcare AI vendors grade well on ai governance and bias disclosure?

Of the 554 vendors in the AI Health Index, 11 grade A on this axis, 103 grade B, 330 grade C and 110 grade D under the AI Health Index grading framework. Grades were last verified on August 31, 2026. Grades are not aggregated into a composite score.

What does an A grade mean on ai governance and bias disclosure?

Substantive responsible AI commitments: published model cards, decision support transparency attributes for EHR-embedded tools, bias and fairness evaluations with stated methodology, third-party AI audits. A bias or fairness evaluation with a stated method, subgroup performance, or an independent audit of model behaviour.

What does a D grade mean on ai governance and bias disclosure?

Nothing published on how model behaviour is governed or tested. Multilingual operation with no subgroup performance sits here when the vendor markets recognition quality as a strength, because a caller the system failed to understand leaves no complaint and no record. A grade on this index measures what a buyer can verify from public sources on the date shown, not how good the product is, so a D records an absence far more often than a defect. A vendor that publishes more is regraded.

Do vendors pay to be included or graded?

No. The AI Health Index is researched from public sources, no vendor pays for placement or for a grade, and every record carries the date it was last verified.

The other 14 axes

No single axis decides a selection. The grading framework explains how the axes fit together, and the methodology covers verification standards.

AI Health Index grades verified August 31, 2026 · Browse all vendors · Compare vendors · Change log