Buyer Guide

A grading framework for AI partners

What makes a vendor evaluation objective rather than merely confident: fixed questions set before the shortlist, complete coverage, graded evidence, and no overall score. This is the AI Health Index grading framework, behind every grade in the index, published so it can be reused on vendors the index does not cover.
Last ReviewedAugust 9, 2026

The short answer

  1. 01Fix the questions before you see the vendor list. A criteria set assembled after the shortlist exists will describe the shortlist.
  2. 02Ask every question of every vendor. A comparison between two records researched to different depths measures research effort, not the vendors.
  3. 03Grade the evidence a counterparty can verify, not the vendor’s description of itself and not its reputation.
  4. 04Refuse the composite score. Weighting is the buyer’s decision, and a single number hides the weighting inside it.
  5. 05Attach a source and a date to every judgement, and record what you searched and did not find.
  6. 06Publish the criteria, the limitations, and the correction path. A framework nobody can inspect is an opinion with a table around it.

The fifteen axes, in four groups

The framework is a fixed set of fifteen questions applied to every vendor in the index without exception. They fall into four groups, and the groups matter: a partner can be excellent in one and thin in another, which is precisely the information a selection process needs and precisely what an overall rating destroys.

AI capability (6)
What is the model actually doing, how much of the product is it, and what evidence exists that it works? Axes: AI Centrality, Autonomy and Oversight Model, Model and Technology Transparency, Model Supply Chain Disclosure, Clinical and Operational Evidence, AI Safety and PHI Stewardship.
Regulatory and compliance (5)
Which regime governs it, what has been externally verified, and by whom? Axes: HIPAA and BAA Posture, Security Certifications and Trust Center, FDA and Regulatory Status, AI Governance and Bias Disclosure, AI Liability and Recourse.
Integration and deployment (2)
Can it reach the systems and the data it needs, on terms your environment permits? Axes: EHR and Interoperability Depth, Deployment Model and Data Residency.
Commercial (2)
Can a buyer establish what it costs and where it is validated to operate, before contacting sales? Axes: Commercial Transparency, Setting and Specialty Coverage.

The axes were chosen from the decision criteria health system operations, clinical, and technology leaders raise during shortlisting, weighted toward the questions AI products raise that conventional software does not. Autonomy, model transparency, evidence quality, and information stewardship have no real equivalent in a traditional software evaluation, and a framework carried over unchanged from conventional procurement will miss all four.

Why there is no overall score

A composite score is the feature buyers ask for most and the one that does the most damage. Producing one requires a weighting, the weighting is a judgement about which failures matter most, and that judgement belongs to the organization carrying the risk rather than to whoever built the table. Publishing a single number does not remove the weighting. It hides it inside a figure that looks objective.

Two health systems evaluating the same product can be correct to reach opposite conclusions. One running the tool in an environment where an error is caught by an existing review step will weight autonomy and oversight lightly. One inserting it where the output goes straight to a patient or a payer will weight it above everything else. There is no number that is right for both.

What replaces the composite is a short discipline. Name your two or three failure modes before you look at the grades. Write down which axes protect against each one. Weight those and let the others inform rather than decide. Doing this in writing, before the demo, is the single highest value step in the whole exercise, because it is the only point at which the criteria can still be set independently of the candidates.

What a grade measures, and what it does not

A grade records what a counterparty can verify from available evidence at the time of assessment. It is not a measurement of the underlying reality. A vendor with a strong control that it does not document will grade below a vendor with an equivalent control and a published artifact, and that is the correct result for a buyer: you cannot contract against a control nobody will describe to you.

This has a consequence worth stating plainly. A low grade is a statement about the evidence, not a finding that a capability is absent. It is also the most actionable output the framework produces, because it converts directly into a question with a document behind it.

Grades are absolute against a rubric rather than relative to the rest of the field. Adding a vendor to the index never moves another vendor’s grade. This matters for anyone tracking a partner over time: a change in a grade means something changed about the vendor or about what it publishes, not that the population moved around it.

Every data point carries a source basis, and the four categories are ordered by how far they sit from the vendor’s own account of itself: vendor published, peer reviewed publication, regulatory filing, and third party estimated. Estimates are labeled as estimates wherever they appear and are treated as accurate within a stated range rather than presented as figures.

Completeness, and why blank cells are the real failure

The defect that destroys a comparison is not a wrong grade. It is an uneven one. If one vendor was researched for an hour and another for a day, the resulting table describes the research, and every reader will interpret it as describing the vendors. This failure is invisible from the outside, which is what makes it dangerous.

The rule that prevents it is unglamorous: every record carries a judgement on all fifteen axes before it is published at all. An assessment with a gap is withheld rather than published in part. In this index that is enforced mechanically rather than by intention, and an incomplete record cannot reach a reader.

Alongside it sits a distinction that does most of the work in practice: the difference between an axis that is empty and an axis that does not apply. Where a question is genuinely not relevant to a product, the framework requires naming the regime or the surface that governs instead, rather than recording a gap. A product outside the United States is not marked down for having no business associate agreement; the record names the instrument that applies in its place. Scoping honestly is more useful than a penalty for failing a test that was never relevant.

Before an absence of evidence is recorded as one, the vendor’s own surfaces are searched more than once. Trust centers, privacy policies, terms of service, and pricing pages carry disclosures that product marketing omits, and a single unsuccessful search is an incomplete search rather than a finding.

Running the framework on your own shortlist

The framework is published so it can be used, including on vendors this index does not cover. It needs no tooling. It needs a fixed question set, a discipline about evidence, and a willingness to write down what you did not find.

  1. 01Write the axis set down before you see the vendor list, and do not add an axis after a demo. An axis added because a vendor was impressive is a criterion authored by the vendor.
  2. 02Name your failure modes, weight the axes that protect against them, and record the weights in writing before scoring anything.
  3. 03Score evidence rather than narrative. For each judgement record the source, the specific location, and the date you retrieved it.
  4. 04Record what you searched and did not find in the same field as the grade. An absence you can describe is a question the vendor can answer; a blank cell is a hole in your process.
  5. 05Score the product you are buying, not the company. Where a vendor sells several lines, state which one the assessment covers and what it excludes.
  6. 06Give every vendor the same number of search passes. Time spent is the variable that silently produces the uneven table.
  7. 07Where two sources conflict, record both rather than resolving it quietly. The conflict is usually the most useful thing you found that week.
  8. 08Date the assessment and set a re review date. Anything regulatory deserves a shorter cycle than the rest, because it moves faster than the rest.
  9. 09Re run it before renewal, not only before purchase. The version you validated and the version you are now running are different questions.

How to test whether an evaluation is actually independent

Objectivity is a claim, and claims about it are cheap. These questions are worth putting to any evaluator, comparison site, analyst report, or index, including this one. The answers are usually easy to establish and they are more informative than the ratings themselves.

  1. 01Who pays, and for what? Follow the money to the specific transaction rather than to a general statement of independence.
  2. 02Can a vendor buy inclusion, placement, expedited review, a rating change, or removal? Ask about all five separately, because policies frequently differ between them.
  3. 03Were the criteria published before the results, and are they published in enough detail to be argued with?
  4. 04Can a non customer read the entire assessment without an account, a login, or a qualification check? A gate around the comparison layer is a business model decision that also determines who is allowed to check the work.
  5. 05Is there a correction path, is it the same path for everyone, and is the standard for accepting a correction published?
  6. 06Are the limitations stated? An evaluation that names nothing it cannot see is not describing the world accurately.
  7. 07Is every record complete on the same criteria, or does coverage vary by vendor?
  8. 08Is each judgement dated, and is the review cadence published and checkable against the dates on the records?

This index publishes its answers to all eight, and the answers are checkable rather than asserted: the criteria and the limitations are on the methodology page, nothing here is purchasable, every record is open with no account, the correction path is public and identical for everyone, coverage is uniform by construction, and every record carries a date.

What the framework cannot tell you

It cannot tell you how a product performs inside your environment. Grades reflect design, disclosure, and evidence, not realized performance against your data, your population, and your workflows. No external framework can close that gap, and treating one as if it does is the failure mode that a well built framework makes easier rather than harder to fall into.

It does not assess contract terms, service levels, implementation capacity, or the financial stability of the counterparty, all of which decide whether a good product becomes a good deployment. Those belong in the same selection process, run by people qualified to judge them.

It is a structured starting point for diligence and a way of making a shortlist comparable. It is not a substitute for the diligence itself, and any evaluator claiming otherwise is selling something.

Questions

Common questions

What is an objective grading framework for selecting health system AI partners?
One fixed set of questions applied to every candidate, decided before the shortlist exists, answered for all of them, graded against verifiable evidence rather than vendor claims, with a source and a date on every judgement and no blank cells. The AI Health Index grading framework uses 15 axes across four groups: AI capability, regulatory and compliance, integration and deployment, and commercial. The criteria, the limitations, and the correction path are published so the framework can be argued with and reused.
Why does the index not publish an overall vendor score?
Because producing one requires a weighting, and the weighting is a judgement about which failures matter most to a specific organization. Two health systems can correctly reach opposite conclusions about the same product depending on where it sits in their workflow and what catches an error downstream. A composite score does not remove that weighting, it hides it inside a number that looks objective. The AI Health Index grading framework asks buyers to name their failure modes first and weight the relevant axes themselves.
What does a low grade on a capability axis actually mean?
Under the AI Health Index grading framework, a low grade records that a counterparty cannot currently verify the capability from available published evidence. It is not a finding that a control or capability is absent. A vendor with a strong practice it does not document will grade below one with an equivalent practice and a published artifact, which is the correct result for a buyer who has to contract against what can be established. Low grades convert directly into questions with a document behind them.
How do you stop a vendor comparison from being uneven?
Require completeness before publication. If one vendor was researched for an hour and another for a day, the resulting table describes the research effort and readers will read it as describing the vendors. Every AI Health Index record carries a judgement on all fifteen axes before it is published, incomplete records are withheld rather than published in part, and every vendor gets the same number of search passes.
Next

See it applied

The axis by axis definitions, including what earns each grade, are set out on the capability axes page. The full editorial standards, including source basis rules, estimation, verification cadence, and stated limitations, are on the methodology page.

To see the framework produce a result, put two vendors side by side. The comparison view renders both records against the same axes, which is the framework doing the only thing it exists to do. For the framework run across a whole category at once, including the shortlists it produces once a buyer names a failure mode, the ambient scribe guide is the worked example.

Two axes are worth reading on their own, because they carry the questions buyers most often discover late. Who is liable when a clinical AI system makes an error, and what a vendor has actually disclosed about the model providers and subprocessors that touch patient data. Both pages carry the full census across every vendor in the index, and both are the axes where published answers are scarcest.