← All issues

The AI Health Index Brief

August 23 to August 29, 2026 · Published August 30, 2026

The week in one line

The ambient scribe finished writing the note and went looking for more to do. This week it prepared the visit, drafted the prescription, and proposed the billing code, while on the administrative side the agents stopped sorting the queue and started clearing it. The work is moving from software that describes it to software that does it, and the review step in front of the signature is now the part worth reading closely.

This issue covers August 23 to August 29: 20 entries across 20 vendors, 17 Verified at source and three Partially Verified from trade press. Funding rounds, valuations, and awards are not logged, here or anywhere on this index.

Theme 1: The scribe left the note

FDB commercially deployed FDB Script Agent, which converts ambient encounter dialogue directly into structured prescriptions. The agent applies FDB’s drug terminology codification and clinical validation to what it hears, so the output is a coded medication order rather than free text, and it queues that order for clinician review rather than transmitting it. It launched first on the Tebra platform. Abridge shipped its August release with Pre Visit Summaries that synthesize a patient’s history, active problems, and recent changes before the visit, every detail linked back to its source in the chart, and Pre Admission Summaries that do the same for inpatient handoff, pulling together the emergency department course, labs, vitals, imaging, and prior documentation. The release also adds transparent level of service recommendations for coding, and lets clinicians claim AMA PRA Category 1 CME credit for reviewing eligible topics at the point of care.

Suki launched Suki Dictation, a standalone clinical dictation product built natively inside Epic and Meditech and sold separately from its ambient scribe, so an organization can buy dictation without committing to full ambient documentation. Recorded Partially Verified, trade press. And Nabla shipped the August 25 release of its Core API, adding an optional specialty field and an encounter date field to note generation so a caller can tell Nabla what kind of encounter it is transcribing instead of leaving the model to infer it, plus an audio chunk acknowledgement protocol and strict endpoint path prefixes, which is a breaking change for anyone relying on loose paths.

A year ago the job description was one sentence: listen to the visit and write the note. This week’s entries extend it in both directions. Upstream, Abridge’s summaries are the pre visit reading a clinician used to do standing up, four minutes before the patient walked in, and they arrive with each line linked to the chart. Downstream, FDB’s agent turns the same audio into a coded order. Each step away from the note moves the product into a part of the chart where a mistake is not a bad paragraph. A misheard dose is a different class of error from a misheard adjective, and FDB’s design says so: a coded order that waits for a signature rather than a transmission that has already left the building.

Our read

The scribe category is turning into the operating layer for the encounter, and the useful question changes with it. A scribe is judged on the accuracy of a paragraph. An order entry agent is judged on what the clinician sees before signing and on what happens when it mishears. Nabla’s specialty and encounter date fields are the unglamorous plumbing of that, a note shaped to the service producing it rather than a generic one. And Abridge found a way to make the review step worth something to the person doing it, which every vendor in this theme should study, because review steps that reward nobody tend to get skipped.

Buyer question

For any ambient tool that now proposes an order, a code, or a summary: what exactly does the clinician see before signing, is the specialty and encounter type passed by the system or inferred from the audio, does every line of a summary link to its source in the chart, and what is written to the record when the clinician overrides the suggestion.

Theme 2: The agents picked up the phone

Waystar launched agentic capabilities inside its AltitudeAI platform, moving from software that scores and routes work to agents that carry it out: claim resolution, conversational performance intelligence, and agentic clinical documentation that reads the full medical record and pre populates correction requests with the supporting clinical context already attached. Waystar reports roughly a 40 percent reduction in manual correction workload in early deployments, a vendor figure. Medallion shipped three capabilities at once: an outreach agent that contacts stalled providers by phone, text, and email to chase profile completion, provider profile auto fill from NPPES, CAQH, and uploaded documents, and self serve webhooks that push enrollment and credentialing status changes into an organization’s own systems over API.

Luma Health shipped its Summer 2026 release, extending its Operational AI from pre visit preparation onto the clinic floor: self check in by kiosk or QR code, a Queue Manager that routes arrivals automatically, and financial administration tooling for facility managed card readers and transaction reporting. OmniMD launched an AI Revenue Intelligence product with remote patient monitoring built in rather than alongside, covering onboarding, vital sign streaming, clinician summaries, and a billing workflow that creates the note natively in the EHR; Partially Verified, trade press. And Hyro published the architecture behind its patient scheduling agent, disclosing that it orchestrates seven distinct smaller models, each with one job, such as extracting date and time constraints from what the patient said, querying slot availability, or answering general questions.

The verbs changed. For a decade the contribution of administrative software was ordering: score it, route it, prioritize it, and a person did the rest. This week’s entries use different verbs, resolve, chase, check in, populate, and the test of a real shift is whether the vendor names what the agent completes without a person. Waystar, Medallion, and Luma all do. Credentialing coordinators have spent whole careers leaving voicemails for providers; Medallion has built something that leaves them tirelessly and without sighing. And seven small models with one job each is a sound engineering choice that also happens to describe every well run front desk.

Our read

A 40 percent reduction in manual corrections is the right kind of number, and the figure that would settle it is one no launch post publishes: whether the corrections an agent assembles are accepted by payers at the same rate as the ones a person writes. Hyro’s disclosure is the model for the category. A buyer who knows the topology knows where an error would come from and whether fixing it means retraining everything or adjusting one small part. A vendor willing to describe how the thing is built is easier to evaluate than one that will only describe what it achieved.

Buyer question

What does the agent complete without a person, what is the acceptance rate of agent authored work against human authored work, how does an outreach agent identify itself to the people it calls, and what is the escalation path when it cannot reach someone or is told no.

Theme 3: The evidence arrived in the right order

Caris Life Sciences published a study in npj Precision Oncology reporting that its AI guided therapy selection predicts longer survival in pancreatic cancer, a deliberately hard test case with short survival and few targetable alterations. The work validates Caris AI Insights specifically rather than the profiling platform in general. Brainomix published an implementation study of Brainomix 360 Stroke in a high volume stroke system that already used routine perfusion imaging, finding that the platform improved the efficiency of the acute stroke pathway even in a setting that was not imaging constrained to begin with. GeneDx published results from Seattle Children’s supporting hospital wide adoption of rapid genomic sequencing rather than the usual narrow deployment in one unit, evaluating clinical and operational impact together. And Mercy and Aidoc released a white paper commissioned by AVIA reporting more than 50,000 new clinical findings surfaced across the Mercy system in five months, with the associated changes in speed of diagnosis and time to treatment; Partially Verified, trade press.

Four evidence entries in one week, and they happen to land in descending order of weight: a survival endpoint in a peer reviewed journal, a pathway measured in an operating service, an institution wide operational study, and a commissioned white paper from a named health system. Each answers a different buyer question and none answers the others. Survival is the endpoint nobody argues about, and it is also the one that takes years to collect, which is why most evidence in this category measures something quicker. Brainomix measured the thing hospital leaders actually purchase, the pathway, in the setting with the least headroom for a tool to add value, so a positive result there travels further than one from an under resourced site.

The caveats are printed on each entry and they are the useful part. Predicting longer survival in a retrospective cohort is not the same as showing that following the recommendation caused it, and the design decides which claim is supported. Fifty thousand new findings depends entirely on what counts as a new finding. AVIA commissioning the Mercy work makes it more independent than a vendor case study and less independent than peer review, which is roughly where its weight should sit.

Our read

The ladder matters because each rung is priced differently at the committee. A survival study gets an oncology service through utilization review; a pathway study gets a stroke director a budget line; an institution wide study gets a genomics program past the ward it started on; a commissioned white paper gets a meeting. Read the design before the headline number. On every one of these entries the design is on the first page.

Market notes

Tempus received FDA 510(k) clearance for Tempus ECG PH, software that reads a standard 12 lead electrocardiogram and identifies signs associated with pulmonary hypertension, a condition characteristically diagnosed late. It needs no new hardware and no new test, which is the good news, and the downstream question is what happens after a positive flag, because confirming pulmonary hypertension requires right heart catheterization and a screening tool that raises referral volume without a matching diagnostic pathway creates a bottleneck rather than removing one. Stratipath Breast received CE marking under the EU In Vitro Diagnostic Regulation, a substantially higher bar than the directive it replaced, requiring clinical evidence, demonstrated analytical validity, and an ongoing post market surveillance commitment rather than largely self declared conformity. Second issue running with a European approval in the log, and worth repeating that IVDR certification and FDA clearance cover different claims and are not interchangeable.

Paige launched a tool that screens 505 genes directly from H and E stained pathology slides, predicting a genomic profile from routine morphology and flagging biomarkers for follow up, positioned as a triage layer that decides which cases warrant full sequencing rather than as a replacement for it. The slide was already on the bench; the expensive part was everything that came after it. The question to ask is which of the 505 have been validated against sequencing ground truth and at what sensitivity, because a screen that misses a targetable mutation carries a different risk than one that over refers. Insilico Medicine convened O3DC, an open consortium on benchmark quality in AI drug discovery, and published a live catalogue of the benchmarks the field uses to claim performance, including each one’s known caveats and biases. It is useful and it is not disinterested, since the convening vendor is scored on those benchmarks too, and a map of the measurement problems is worth having from whoever draws it.

Three vendors published dated release notes in one week, which in an index where roughly one vendor in ten publishes any counts as a parade. Commure version 2026.3.1 adds a widget that starts or resumes a scribe session from the iOS Home Screen or Lock Screen and moves AI Studio into the main tab bar. Scribe adoption is decided in the four seconds before the patient sits down, so a lock screen button may be the most consequential scribe release of the week. Eleos Health version 10.0.17 lets the assistant populate dropdown fields, so it now completes structured behavioral health records rather than only free text, and tightens note visibility so notes surface only to the providers they are relevant to. CarePilot’s August 28 release adds organization wide note attestation settings, automated questionnaire scoring at intake, preventive screening and ENT codes in billing suggestions, and an account lockout after five failed attempts, which is exactly the kind of control that appears on security questionnaires and rarely in a public release note.

What the week says about the category

Three themes, one pattern. The scribes moved into the order queue and the coding suggestion, the administrative agents moved from sorting work to finishing it, and the evidence entries came with their designs printed on the front. What ties them together is that the products in this log are now competing on the step after the model output: what the clinician sees before signing, whether the agent’s work is accepted, how the thing is built, which endpoint the study measured. Those are all things a buyer can inspect and an index can grade, which is the good news, and it is a different category from the one that competed on the accuracy of a paragraph.

Index Answer

Which healthcare AI vendors accept responsibility when their AI is wrong?

Almost none in writing, and this week’s log makes the question sharper. Of the 554 vendors the AI Health Index has assessed on AI Liability and Recourse as of August 31, 2026, the number that publish all three things the top grade requires, an error rate with its method and denominator, a plain statement of what the system does not do, and a warranty, indemnity or remediation commitment written down, is 1. Another 112 sit one band below, close on some of the three and short on at least one. The remaining 441, 200 at C and 241 at D, sit in the bottom two grades on the question of whether any warranty, indemnity or remediation commitment is stated anywhere a buyer can read it before a sales conversation.

That matters this week because the products in the log moved closer to the consequence. An ambient tool that drafts a coded prescription, an agent that assembles a claim correction from the medical record, and a cleared algorithm that flags pulmonary hypertension from an ordinary ECG all produce output a person signs and an institution answers for. The AI Health Index grades this axis on what the vendor publishes, not on what it says in a deal, so a vendor can ship an order entry agent and still sit in the bottom band on who carries the cost when it mishears a dose. Most of the field does exactly that, and it is the weakest of the fifteen axes the index applies, by a wide margin.

The reason nobody else answers this question is that it is a question about absence. Every vendor describes its own posture and every contract allocates liability somewhere, but the person carrying the clinical risk is usually not a party to the contract, and no vendor has an incentive to publish where the whole field stands. The practical move is the same one that works for subprocessors: ask for the artifact rather than the posture. Ask for the limitation of liability clause, the indemnity language, the carve outs, and the recourse process as documents, and ask whether any of it reaches the patient or stops at the health system. A vendor graded B here can usually produce those in a day. The difference between B and D is mostly whether anyone at the vendor was willing to write it down.

Full grades, the axis definition and what separates each band are on the AI Liability and Recourse page.

The AI Health Index Brief is published weekly by AI Health Index, an independent reference for evaluating AI vendors in healthcare. No vendor pays for inclusion, placement, or rating. Compare any indexed vendors by capability at Compare and read the evaluation standards at Methodology.