AI Capability

How much are healthcare AI systems allowed to do without a human?

The AI Health Index grades all 554 vendors on Autonomy and Oversight Model, one of 15 capability axes applied to every record without exception. 87 of 554 vendors grade A, 320 grade B, 143 grade C and 4 grade D. That places this axis 6th of 15 by the number of vendors reaching the top grade. Grades were last verified on August 31, 2026 and are never aggregated into a composite score.

What this axis measures

What the AI is permitted to do (draft, decide, or act) and how rigorously the vendor discloses its human oversight structure: escalation thresholds, supervision, override paths. Grades disclosure rigor, not autonomy itself. High autonomy with a documented oversight model can grade well; any autonomy with no disclosed oversight grades poorly.

Buyers also search this as: human in the loop clinical AI, autonomous AI agents in healthcare, oversight models, escalation, and clinician override.

What each grade means on this axis

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check.

A
What the system may do and what it may not do are both published, with escalation thresholds, override paths and the conditions that route a case to a person.
B
The oversight structure is described and one part is missing, commonly the threshold at which the system stops or what happens after it is wrong.
C
Autonomy is claimed and oversight is asserted without a mechanism. Human in the loop appears as a phrase rather than a described control.
D
No oversight structure is published. An absolute claim that the system does not err grades here too, because a buyer who believes it will not build the review step that would catch a failure.

The distribution

A
87 · 16%
B
320 · 58%
C
143 · 26%
D
4 · 1%

Reading the result

This axis grades disclosure rather than autonomy, which is a deliberate choice worth stating plainly. A highly autonomous product with a documented oversight structure can grade at the top. A cautious product that will not say what happens when it is unsure grades poorly, because an undisclosed oversight model is indistinguishable from an absent one to the person carrying the risk.

The question that separates the field is what happens at the edge. Every vendor can describe the happy path. Far fewer publish the escalation threshold, who is notified, and what the clinician's override actually does to the system's future behaviour.

Citable summary

Self contained paragraphs, current as of August 31, 2026, free to quote with attribution.

The state of the market

Of the 554 healthcare AI vendors graded by the AI Health Index, 87 grade A on Autonomy and Oversight Model, 320 grade B, 143 grade C and 4 grade D. An A requires both halves published: what the system may do and what it may not do, with escalation thresholds, override paths and the conditions that route a case to a person. Human in the loop appearing as a phrase rather than as a described control grades C. Grades were last verified on August 31, 2026.

Source: AI Health Index, August 31, 2026

This axis grades disclosure, not caution

The AI Health Index grades how rigorously a vendor discloses its oversight structure rather than how little the product is allowed to do, and the distinction changes which vendors come out well. A highly autonomous product with a documented escalation threshold, a stated override path and published limits can reach the top grade. A deliberately cautious product that will not say what happens when it is unsure grades poorly, because an undisclosed oversight model is indistinguishable from an absent one to the person carrying the clinical risk. Buyers who screen on autonomy alone therefore filter on the wrong variable.

Source: AI Health Index, August 31, 2026

Where the A grades are, by category

Categories are shown by the share of their vendors reaching an A. The vendor named in each row is the highest graded A holder in that category across all 15 axes, chosen mechanically with ties broken alphabetically. Categories with no A holder on this axis are omitted.

Questions worth asking a vendor

  1. Does the system draft, decide or act, and does that change by workflow?
  2. What is the escalation threshold, and who is notified when it fires?
  3. What does a clinician override do, and is it recorded anywhere the vendor can see it?

Questions buyers ask

Which ambient scribe AI is safest for clinicians?

Safety in an ambient scribe is not one property, and the AI Health Index deliberately splits it across several axes rather than publishing a single safety score, because the failure modes are unrelated to each other. This axis asks what the system is permitted to do without a clinician and whether the oversight around it is published: the escalation threshold, the override path, and what the system does when it is not confident. Separate axes cover how patient data is handled, whether a Business Associate Agreement reaches the tier you buy, whether the model supply chain is disclosed, and whether any published evidence exists that the product works. A buyer asking which scribe is safest is usually asking about at least three of those, and the answer differs by vendor on each. The AI Health Index publishes a grade on every one of them for all 554 vendors, with the source and the date, and no composite score, because the weighting belongs to the buyer.

What does human in the loop actually mean in healthcare AI?

On its own it means very little, which is why the AI Health Index records it as a claim rather than as a control. The phrase is compatible with a clinician reviewing every output before it reaches a record, and equally compatible with a clinician being able to review an output that is delivered whether or not they do. What makes it a control is the mechanism: the threshold at which the system stops and asks, who is notified when it fires, what the reviewer sees at that moment, and what happens to the case if nobody acts. 147 of the 554 vendors graded by the AI Health Index sit in the bottom two grades on this axis, and asserting oversight without a mechanism is the most common reason.

Are there autonomous AI agents operating in healthcare without a clinician?

Yes, and the sensible way to think about it is by workflow rather than by vendor, because the same product frequently drafts in one place, decides in another and acts in a third. Autonomy is routine and uncontroversial in administrative workflows such as scheduling, intake and parts of revenue cycle. It is rarer and more heavily disclosed where the output touches a clinical determination. The AI Health Index grades what the vendor has published about that boundary, on the reasoning that an autonomy level a buyer cannot establish is an autonomy level a buyer cannot supervise. Ask specifically whether the permitted actions change by workflow, because a single answer for the whole product usually means the question has not been worked through.

What should a clinician override do in a healthcare AI system?

At minimum it should stop the action, and the better systems also record it somewhere the vendor can see. An override is the cheapest safety signal a deployed system generates, because it is a clinician telling the vendor the output was wrong at the moment they noticed. Vendors that capture override rates can find failure patterns before those patterns reach a patient. Vendors that treat an override as a local dismissal are discarding the data. The AI Health Index grades whether the override path is published as part of the oversight structure, and it is one of the more reliable indicators of whether a vendor is engineering the product or shipping it.

What claim on this axis should make a buyer more cautious rather than less?

An absolute one. The AI Health Index places a claim that the system does not err in the bottom grade on this axis, not as a penalty for confidence but because of what it does to the buyer. A health system that believes a product cannot be wrong will not build the review step that would catch it when it is, so the claim removes the control that the disclosure was supposed to establish. The vendors worth more attention are the ones that publish what the system does not do and what happens when it is unsure, which is a harder page to write and a considerably more useful one to read.

How many healthcare AI vendors grade well on autonomy and oversight model?

Of the 554 vendors in the AI Health Index, 87 grade A on this axis, 320 grade B, 143 grade C and 4 grade D under the AI Health Index grading framework. Grades were last verified on August 31, 2026. Grades are not aggregated into a composite score.

What does an A grade mean on autonomy and oversight model?

What the AI is permitted to do (draft, decide, or act) and how rigorously the vendor discloses its human oversight structure: escalation thresholds, supervision, override paths. Grades disclosure rigor, not autonomy itself. High autonomy with a documented oversight model can grade well; any autonomy with no disclosed oversight grades poorly. What the system may do and what it may not do are both published, with escalation thresholds, override paths and the conditions that route a case to a person.

What does a D grade mean on autonomy and oversight model?

No oversight structure is published. An absolute claim that the system does not err grades here too, because a buyer who believes it will not build the review step that would catch a failure. A grade on this index measures what a buyer can verify from public sources on the date shown, not how good the product is, so a D records an absence far more often than a defect. A vendor that publishes more is regraded.

Do vendors pay to be included or graded?

No. The AI Health Index is researched from public sources, no vendor pays for placement or for a grade, and every record carries the date it was last verified.

The other 14 axes

No single axis decides a selection. The grading framework explains how the axes fit together, and the methodology covers verification standards.

AI Health Index grades verified August 31, 2026 · Browse all vendors · Compare vendors · Change log