AI Capability
How do healthcare AI vendors handle PHI across the AI lifecycle?
The AI Health Index grades all 554 vendors on AI Safety and PHI Stewardship, one of 15 capability axes applied to every record without exception. 32 of 554 vendors grade A, 182 grade B, 309 grade C and 31 grade D. That places this axis 10th of 15 by the number of vendors reaching the top grade. Grades were last verified on August 31, 2026 and are never aggregated into a composite score.
What this axis measures
How the vendor handles protected health information across the AI lifecycle, including use in training, retention, and de-identification, along with safety engineering disclosures such as guardrails, hallucination mitigation, and safety event reporting.
Buyers also search this as: PHI handling, is our data used for training, de identification, hallucination mitigation, and AI guardrails in healthcare.
What each grade means on this axis
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check.
The distribution
Reading the result
The question that decides this axis is whether patient data trains anything, and the answer is often given in language built to be reassuring rather than precise. Not used to train our models is compatible with being used to train a provider's models, and de identified is a claim with a technical standard behind it that most marketing pages do not reference.
The safety engineering half is newer and thinner. Guardrails, hallucination mitigation and safety event reporting are documented by a minority, and where they are documented the disclosure is usually a description rather than a measurement. That gap is the same shape as the one on the liability axis and has the same root cause.
Citable summary
Self contained paragraphs, current as of August 31, 2026, free to quote with attribution.
The state of the market
Of the 554 healthcare AI vendors graded by the AI Health Index, 32 grade A on AI Safety and PHI Stewardship, 182 grade B, 309 grade C and 31 grade D. An A requires retention windows, training use and de identification stated specifically enough that the statement could be contradicted, alongside the safety engineering itself: guardrails, hallucination mitigation, and how a safety event is handled. General assurances of privacy and security grade C, because they answer the question the industry asked before AI rather than the ones it raises. Grades were last verified on August 31, 2026.
Source: AI Health Index, August 31, 2026
Not used to train our models answers for one party in a chain of several
A vendor stating that customer data does not train its models has answered for its own models, which in an AI product is rarely the whole chain. The data often reaches a foundation model provider, a transcription layer and a hosting platform, each carrying its own default posture on retention and on training. The AI Health Index grades the specificity of the statement rather than how reassuring it sounds, which is why 340 of 554 vendors sit in the bottom two grades on this axis. The question that resolves it is whether the commitment names every party that touches the data or only the party the buyer signed with.
Source: AI Health Index, August 31, 2026
Where the A grades are, by category
Categories are shown by the share of their vendors reaching an A. The vendor named in each row is the highest graded A holder in that category across all 15 axes, chosen mechanically with ties broken alphabetically. Categories with no A holder on this axis are omitted.
Questions worth asking a vendor
- Does our data train anything, including any model belonging to a provider in the chain?
- What de identification standard is applied, and who verifies it?
- Is there a safety event process, and has any event ever been reported under it?
Questions buyers ask
Do AI medical scribes use patient data to train their models?
Some do, some do not, and a large part of the market states it in language built to reassure rather than to be checked, which is why the AI Health Index grades the precision of the disclosure rather than accepting the claim. Read the sentence carefully: not used to train our models is compatible with being used to train a provider's models further down the chain. Three questions settle it. Does any data leave the vendor's environment, and to whom. Do the terms with that party bar training and set a retention window. Is de identification applied before the transfer or after it. 32 of the 554 vendors graded by the AI Health Index answer this axis specifically enough to reach an A.
What does de identified mean in healthcare AI?
It is a claim with a technical standard behind it that most marketing pages do not reference, and the AI Health Index treats an unreferenced claim as weaker evidence than a referenced one for that reason. The recognised routes are removing a defined set of identifiers, or an expert determination that the risk of re identification is very small. Those are different amounts of work with different residual risk, and free text clinical notes are considerably harder to de identify than structured fields because identifiers appear in prose. Ask which method is applied, who verified it, and whether it runs before the data reaches any external model or after.
How long do healthcare AI vendors retain patient data?
Where it is published at all it ranges from not retained beyond the session to retained indefinitely for product improvement, and the gap between those two is the reason this axis exists. The AI Health Index grades a stated retention schedule as evidence and a general commitment to data minimisation as an absence, because only the first can be reviewed by a security team or written into a contract. Ask for the window, then ask whether it applies to the audio, the transcript, the output and the logs separately, since a vendor may delete one and keep another and describe both under a single retention sentence.
What is a safety event process, and do healthcare AI vendors publish one?
A safety event process is the documented route by which a clinician reports that the system produced something wrong or harmful, and the description of what the vendor does next. A minority publish one, and where it is published the disclosure is usually a description rather than a measurement, meaning the process exists on paper without any evidence of what has actually been reported through it. The AI Health Index records the same gap on this axis that it records on AI Liability and Recourse, and the root cause is the same: publishing a count invites comparison. Ask whether any event has ever been reported and what changed as a result.
How do I evaluate whether a healthcare AI vendor handles patient data safely?
Work from published artefacts rather than from assurances, which is how the AI Health Index grades this axis across all 554 vendors in the index. Four things decide it. What is retained, and for how long, stated per artefact rather than in general. What reaches an external model, and under what terms. What de identification standard is applied, and who verified it. What happens when the system produces something unsafe. A vendor that answers all four in writing is in a different position from one that answers them in a call, because only the written answer survives the account manager changing.
How many healthcare AI vendors grade well on ai safety and phi stewardship?
Of the 554 vendors in the AI Health Index, 32 grade A on this axis, 182 grade B, 309 grade C and 31 grade D under the AI Health Index grading framework. Grades were last verified on August 31, 2026. Grades are not aggregated into a composite score.
What does an A grade mean on ai safety and phi stewardship?
How the vendor handles protected health information across the AI lifecycle, including use in training, retention, and de-identification, along with safety engineering disclosures such as guardrails, hallucination mitigation, and safety event reporting. Retention windows, training use and de identification are stated specifically enough to be contradicted, alongside the safety engineering: guardrails, hallucination mitigation, and how a safety event is handled.
What does a D grade mean on ai safety and phi stewardship?
Nothing published on how protected information moves through the system. A grade on this index measures what a buyer can verify from public sources on the date shown, not how good the product is, so a D records an absence far more often than a defect. A vendor that publishes more is regraded.
Do vendors pay to be included or graded?
No. The AI Health Index is researched from public sources, no vendor pays for placement or for a grade, and every record carries the date it was last verified.
The other 14 axes
No single axis decides a selection. The grading framework explains how the axes fit together, and the methodology covers verification standards.