← All issues

The AI Health Index Brief

August 9 to August 15, 2026 · Published August 16, 2026

The week in one line

A systematic review graded a clinical reference tool across eleven studies, an external study compared commercial bone age systems head to head, the FDA admitted a drug titrating AI to its newest pilot, and an insurer priced fall prevention AI into a liability premium. Vendors used to publish their own evidence. This week other people published it for them.

Theme 1: Someone else did the grading

A systematic review published in npj Digital Medicine evaluated OpenEvidence across eleven studies of its clinical question answering. The findings: the platform consistently produced evidence supported responses without fabricating citations, performed best in structured, guideline based settings, and varied in accuracy on complex scenarios. Then the finding worth the price of admission: the system tended to reinforce rather than alter existing clinical decisions.

Sit with that one. From inside a workflow, a tool that confirms a good plan and a tool that flatters a bad one produce the identical experience: the clinician asks, the tool agrees, the day moves on. No accuracy percentage distinguishes those two situations, because the failure is not in the answer. It is in the fact that the answer never collides with anything.

The comparisons arrived too. A retrospective study in Diagnostics externally validated several CE certified bone age systems, including Milvue’s TechCare Kids, and found no significant accuracy differences among them for 90 percent of the clinically relevant pediatric cohort. The finding is parity, and parity published by a third party is worth more to a buyer than superiority published by a vendor. It converts a marketing question, which one is best, into a procurement question: which one fits your PACS, your price, and your support model.

The platforms joined in. Ubie’s Smart Support, which autonomously determines clinical urgency, appropriate specialty, and visit type from patient calls, chats, and portal messages at intake, was qualified on the Mayo Clinic Platform, an independent review of the algorithm’s performance before a health system ever pilots it. The entry is recorded Partially Verified in the log.

And the vendor reported tier kept producing numbers of its own. Creyos reported outcomes from Claremedica, a value based primary care group serving 34,000 Medicare Advantage patients across more than 35 sites: replacing the paper MMSE with its digital cognitive assessment cut average screening time from roughly 12 minutes to six, with more than 90 percent of newly identified dementia cases caught at the early or mild stage. And Ochsner Health reported results from Paradigm Health’s trial recruitment platform across its 47 hospitals and 370 health centers: 41 percent more screening capacity, 3.6 fold more patients identified for future trial eligibility, and a 75 percent reduction in manual review by research coordinators. Both are recorded Partially Verified: real institutions, real numbers, reported by the parties with the strongest interest in them.

Our read

An evidence market only develops around a category once it has enough published claims to be worth auditing, and healthcare AI just crossed that line. This week’s log holds a review of studies, a comparison across competing vendors, a platform qualification, and two named system deployments, each attesting something different. The most useful finding of the week was behavioral rather than statistical. A reference tool that rarely changes the decision is either confirming good judgment or reinforcing bad judgment, and nothing on a benchmark separates the two.

Buyer question

For every evidence claim in a vendor deck, ask who ran the study, who paid for it, and whether anyone outside the company has compared the product against a competitor on the same data. This week produced examples at every tier. The deck will present them in the same font.

Theme 2: The regulator and the insurer both placed bets

Cadence announced that its HypertensionOS software was selected by the FDA as the second participant in the TEMPO for Digital Health Devices Pilot. It is a prescription software medical device with AI assisted functionality that supports clinician supervised medication initiation and titration for Stage 2 hypertension, operating under FDA enforcement discretion while generating real world evidence for the agency. The entry is recorded Partially Verified.

Last issue, clearance became a process: the FDA approved a plan for changing a product, not just the product. TEMPO is the next species in that genus. It is not a clearance at all. It is an admission: the agency letting an AI that helps adjust blood pressure medication operate inside a bounded pilot, in exchange for the evidence it produces. The species census from last issue grows by one, and the word “FDA” on a slide now spans at least three meanings.

The other bet came from an underwriter. VirtuSense launched VSTOne Go, a mobile version of its fall prevention platform for post acute and skilled nursing settings, using infrared sensors rather than cameras to predict unassisted bed and chair exits 30 to 65 seconds before they happen. The detail that matters sits at the end of the announcement: through the CareAgents specialty insurance program, eligible operators adopting the technology qualify for general liability premium credits.

An insurer discounting premiums for a specific AI deployment is the actuarial version of a citation. It is the first external validator in this log that expresses its opinion in dollars per year, and it is refreshingly narrow: the credit prices expected fall claims, nothing more. Actuaries are difficult to impress with adjectives.

Our read

Two institutions that price risk professionally built mechanisms for healthcare AI in the same week, and neither mechanism is a clearance. A pilot admission attests that the agency wants evidence a product can safely generate. A premium credit attests that an underwriter expects fewer claims. Both are real endorsements with real consequences attached, and both are narrower than the sentence a sales deck will build from them.

Buyer question

For a pilot, ask what happens when it ends: does the product then require the clearance it currently operates without, and what is your exit if it does not receive one. For an insurance credit, ask what evidence the underwriter actually reviewed and whether the credit survives your own incident history.

Theme 3: The front door got agents

Assort Health took Referrals to general availability, an agent that runs the inbound referral process end to end: it extracts the referral from EHR feeds and faxes, verifies patient eligibility, applies specialty specific criteria to select the right provider, then contacts the patient and books the appointment directly in Epic or athenahealth. The category calls the problem referral leakage. The patient experiences it as never hearing back.

Cedar evolved its single billing agent into the Kora Platform, a suite of agents spanning autonomous inbound voice, proactive outbound voice, and two way text, handling billing questions, payment processing, and Medicaid enrollment while carrying conversation history across interactions so patients do not repeat themselves. An AI agent now walks patients through Medicaid enrollment, which is among the most consequential paperwork an American can fill out.

Ubie’s Smart Support, qualified above, does the same job at the very first touchpoint, deciding urgency, specialty, and visit type before a human sees the request. And GeneDx opened a direct to family path for pediatric exome testing: families of children with developmental delay, intellectual disability, or epilepsy can initiate testing from the website, with licensed virtual clinicians reviewing histories, ordering the test, and delivering results, against specialty waits that can exceed a year.

Underneath all of it, the parts are becoming standard. Corti added three prebuilt Experts to its Agent Library, a DrugBank module for structured medication and interaction queries, a PubMed module for direct research evidence, and an adaptive clinical Interviewing module, each identifier grounded and exposed through the Model Context Protocol, the connector standard the broader agent industry has settled on. A day later it shipped an ICD-10-UK Coding Expert, which is what a substrate looks like when it starts localizing.

Our read

The clinical encounter has governance around it: committees, validation, sign off. The front door increasingly does not, and the front door is where the consequential administrative decisions now happen: how urgent you are, which specialist you get, how soon, what you owe, whether you end up enrolled in coverage. A misrouted referral or a mis triaged call is a clinical outcome produced by an administrative system no clinical committee ever reviewed.

Buyer question

For every access layer agent, ask who approved the matching criteria and urgency thresholds it applies, what triggers escalation to a human, and how its decisions are audited after the fact rather than just monitored in real time.

Market notes

The journals kept working the diagnostic bench. Medical IP published in PLOS ONE, validating MEDIP PRO’s segmentation of orbital structures on CT to grade thyroid eye disease activity, reaching an area under the curve of 0.938 when combined with clinical parameters. Valar Labs published in Urologic Oncology, validating an image only prognostic biomarker for muscle invasive bladder cancer that risk stratifies patients from standard hematoxylin and eosin slides, with no separate molecular test required. And Techcyte published in the Journal of Clinical Microbiology, its stool ova and parasite detection reaching 96.62 percent positive agreement and 93.33 percent negative agreement alongside operator review. All three Verified, all peer reviewed: the routine hum of the machine this index exists to record.

The plumbing advanced in the settings that usually get enterprise software last. mdhub became a vetted partner on the athenahealth Marketplace with a native athenaOne integration, connecting its admissions, documentation, and scheduling automation into behavioral health workflows without a parallel system. And Inspiren launched a bidirectional integration with Extended Care Professional, so assisted living census changes, room transfers, and level of care updates flow into its monitoring platform without double entry. Behavioral health and assisted living both received native integrations in the same week, which counts as a small correction to a long standing ordering.

The action item: run the disagreement audit

The npj Digital Medicine finding generalizes beyond one product, so treat it as a test for whichever clinical reference AI your organization uses. Pull the last twenty complex cases where a clinician consulted the tool, and count how often the answer changed the plan.

A low number is not automatically damning. Good clinicians mostly ask questions they can already answer. But if the number is zero, you are paying for confirmation, and the review suggests confirmation is exactly what these tools drift toward. Ask the vendor whether they measure disagreement at all, because a reference tool earns its keep precisely on the cases where it disagrees and turns out to be right.

Record the rate, and repeat the audit quarterly. It is one afternoon of work, and it measures the only thing the benchmarks cannot.

The AI Health Index Brief is published weekly by AI Health Index, an independent reference for evaluating AI vendors in healthcare. No vendor pays for inclusion, placement, or rating. Compare any indexed vendors by capability at Compare and read the evaluation standards at Methodology.