The AI Health Index Brief
September 19 to September 26, 2026 · Published September 27, 2026
The week in one line
GRAIL took its cancer blood test before an FDA panel and came away with a split decision, in the same week two large trials of it landed in Nature Medicine. Doximity published a clinical AI exam and Spring Health opened a safety rubric for public comment. InsightRX, meanwhile, announced an accuracy gain of two points, a number too small for marketing and therefore probably true.
This issue covers September 19 to 26. The log holds 29 entries across 25 vendors, and every one of them is Verified at source. Product changes led with 11 entries, and clinical evidence, model changes and EHR integration followed with four each. GRAIL alone accounts for three entries in two days.
Galleri went before the panel
GRAIL had the kind of week most healthcare AI vendors never have. On September 22, Nature Medicine published two studies of its Galleri multi cancer early detection blood test. On September 23, an FDA advisory panel voted on its application for approval.
PATHFINDER 2 is the registrational US study, with 35,878 people enrolled and 32,007 analyzable for performance. The test found cancer in 0.54 percent of them. When it flagged a signal, the person had cancer 60.3 percent of the time, and specificity was 99.64 percent.
Sensitivity is where it gets harder. Across all cancers, the test caught 39.3 percent of those diagnosed during the screening episode. For a prespecified group of 12 cancer types it caught 69.8 percent. It named the likely tissue of origin correctly 91.3 percent of the time, and 0.6 percent of participants had an invasive procedure, with no serious study related adverse events reported.
The NHS Galleri trial in England adds three annual screening rounds. Specificity held between 99.50 and 99.60 percent. Positive predictive value fell each year, from 58.0 to 50.4 to 45.8 percent.
That decline is what repeat screening usually looks like. The first round finds cancers that had been building for years, and later rounds find only the new ones, so each positive result carries a little less certainty. It is also the number a health system would actually live with after year one.
Then the panel. The ten voting members were unanimous that the test is safe. They split 6 to 4 on whether it is effective, and voted 7 to 2, with one abstention, that its benefits outweigh its risks. The FDA is not bound by the vote, and GRAIL expects a final decision in the coming months.
Safe, arguably effective and probably worth it is not a slogan anyone would choose. It is, however, the most carefully argued verdict this log has recorded.
Our readThe last issue said that under Medicare’s new payment model, evidence had become the invoice. This week it became a public hearing. The useful thing about the split is that the evidence was strong enough to argue over, with trials large enough and numbers specific enough for ten experts to disagree about what they mean. Most of this category never publishes enough to be argued with. The disagreement here is about sensitivity, and that is the number every screening tool in this index will eventually be asked for.
The exam went public
Doximity published Bedside Bench, an open source benchmark for clinical AI. It holds 500 cases across ten sub benchmarks, including drug safety, guideline adherence, hallucination and false premises, health equity and diagnostic safety.
The training half and the grading rubrics are public on Hugging Face, and the test half is held out. Bedside Bench is one of the domain benchmarks in Fireworks’ Specialized Intelligence Index, and Doximity states that Doximity Ask ranked first against frontier models when Fireworks graded it.
Writing an exam and then topping it invites an obvious question. A held out test split and an outside grader are the right answers to it, and they are why this counts as a benchmark rather than a brochure.
Spring Health added a Harm From Others rubric to VERA-MH, its open source benchmark for mental health AI safety. It scores how an AI system responds when an adult describes a risk of physical or sexual violence from another person, using 100 personas across five areas. The rubric is open for 60 days of public comment.
A comment period is an unusual thing for a vendor to offer. It turns a safety claim into a standard other people can shape. It is also, it should be said, a very effective way to be the one who wrote the standard.
Vendors also got more candid about their own results. Ambience changed its speech recognition pipeline so a reconciliation model can revise earlier transcript segments as more audio arrives. On a 226 encounter evaluation across 28 specialties, keyword error fell 16 percent. Ambience also published the cost, a median processing cycle about 40 percent longer.
InsightRX added a new adult vancomycin dosing model to the engine that picks models in InsightRX Nova, fit on a sample drawn from 549,171 patients across 304 US organizations. It reports prediction accuracy rising from 66 to 68 percent and bias halving, from 1.8 to 0.9 percent. Nobody writes a press release about two points unless the two points are real.
Deciphex launched CipherX, which pairs pathology foundation models with a layer that translates their output into named tissue signatures that pathologists have validated. It reports negative predictive values of 99.85 percent for adenocarcinoma and 98.76 percent for melanoma in production, with every case still signed out by a pathologist.
Cohere Health reported that Cohere Capture, which finds quality measure gaps in data a health plan already holds, improved Stars performance by an average of 15 percent per measure. And Atropos Health said its evidence network now covers more than 330 million US patients, with queries up to 50 times faster and 30 times cheaper than traditional analytics. No method for that comparison is published, which is the difference between a figure and a finding.
Our readThe published number is getting more honest in two directions. Benchmarks now arrive with held out splits, outside graders and comment periods, and vendors are printing small gains next to real costs. The part worth watching is who writes the tests. An open benchmark is a gift to the field and a positioning move at the same time, because whoever defines the exam defines what good looks like. Expect more vendors to publish one, and expect each to test for the thing its author does best.
The clinician’s edit outranked the transcript
Corti changed how Guided Generation writes a note. Facts a clinician edited, added or discarded now reach the model labeled by type. An edited fact beats the transcript where the two disagree, and a discarded fact stays out of the note even when the recording mentions it.
Before this, edits and discards were not reliably reflected, which is a problem every scribe has to solve eventually. The recording knows what was said. Only the clinician knows what was meant, and the note has to side with the clinician.
Indica Labs gained voice control inside its HALO AP pathology platform through a Voicebrook integration. A pathologist can now navigate slides, order stains and run AI analyses by voice. One spoken command pulls results such as ER and PR percentages into the matching CAP cancer protocol checklist.
Enzo Health shipped Handoff Summary, a pre visit screen that gathers the patient’s context, the alerts for that visit and, for Enzo EHR customers, every care plan change since the last one. Alcidion’s Miya Precision 7.11.0 added a native inpatient admissions workflow with two way integration to the patient administration system, so clinicians no longer have to open it separately.
Alcidion also added a three tier access model for patient records. A clinician gets full access, no access, or break glass access, which requires a justification entered with a personal PIN and leaves an audit trail. Break glass is an old idea in hospital IT. Seeing it arrive in the same release as more automation is the right order of operations.
Our readThe documentation tools are learning whose word is final. For two years the pitch was that the AI listens so the clinician does not have to type. The next pitch is that the clinician’s correction is the most important input the system receives, and that it sticks. It is a less dazzling claim and a far more useful one, because a note is only trusted when its author knows the edit will hold.
The front desk got a test suite
Syllable release 26.9.22 added Experiments, which splits live call traffic across weighted variants of a voice agent and compares the results session by session. It is an A/B test for the voice on the other end of a patient’s phone call.
Artera added automated tests that check an agent’s responses before a change goes live. Its agents can also now listen for the voicemail beep and start speaking when recording begins, which puts them ahead of a meaningful share of human callers.
Artera’s Intake Hub now matches patient forms to appointments using the EHR’s own event ID, so renaming an appointment type no longer breaks form assignment. The release notes say groundwork is in place to write form answers back to the EHR, the same direction the front desk agents in the last issue took.
Relatient extended its self scheduling and call center tools to behavioral health providers, citing one customer where 60 percent of self booked appointments came from new patients. Upheal added voice input to its assistant, so clinicians can speak requests rather than type them.
Our readVoice agents are being managed like software releases, with variants, tests before deployment and results compared by session. That matters more here than in most software. The phone is often a patient’s first contact with a health system, and until recently the voice on it could change whenever a vendor shipped, with nobody measuring whether the change helped.
The lab bench took an API
Insilico Medicine made Model Context Protocol servers available across its Pharma.AI software. Outside AI agents can now connect directly to its biology, chemistry and biologics engines and run multistep discovery workflows. Drug discovery software has grown a door for other companies’ agents, a few months after enterprise software did.
Veeva announced Study Builder Agent, which configures EDC forms, visits and edit checks directly from a study protocol, reusing an organization’s own standards. Veeva says it can bring study configuration down to as little as one day. It is planned for early adopter availability in December, is included in EDC with no additional license, and is delivered as a Claude Cowork plugin.
The distribution choice is the story. A clinical trials platform chose to ship its agent inside somebody else’s assistant rather than ask users to open one more tab.
Edison Scientific’s Kosmos can now read licensed full text from the Nature portfolio when it generates hypotheses or evaluates targets. Each article links to its Version of Record, so outputs reflect corrections, retractions and editors’ notes. An AI that checks for retractions is ahead of a surprising number of human citations.
Iambic released Enchant v3, the model behind its discovery platform, which it states has 41 billion parameters and covers more than 6,000 molecular properties across 16 data modalities. Helix launched Research Workspaces, a governed research environment over more than 550,000 clinicogenomic records, built on Databricks.
Our readThe research tools are converging on the same idea as the clinical ones. Meet the scientist inside the tool they already use, and bring the provenance along. A connector that flags a retracted paper is worth more than one that surfaces ten more papers.
Market notes
Aidoc received FDA Breakthrough Device Designation for a chest radiograph triage product built on its CARE foundation model, covering pneumomediastinum and lobar or lung collapse. It is Aidoc’s third breakthrough designation in just over a year, and the 510(k) is still under review. Coming the week after DeepHealth cleared a foundation model for chest radiographs, it is a fair signal of where radiology AI’s next round of submissions is headed.
ClosedLoop, a healthcare data science and predictive analytics company, no longer sells a healthcare product. Its domain now offers a workspace for agentic software development, and no acquirer of the healthcare business has been announced. Losing a vendor to the general agent market is an exit this category has not had to think about much. It may have to.
Cynerio’s medical device security now ships inside Axonius for Healthcare, and its own domain redirects there. Zus Health refreshed the code sets it uses to tag specially regulated records. Contraceptive management records now carry the same sensitive label as abortion and gender affirming care under California AB 352, and data already stored was retagged.
What the week says about the category
Twenty nine entries, all verified, and the theme is exposure. GRAIL put its test in front of a federal panel. Doximity and Spring Health put their benchmarks in front of the public. Ambience and InsightRX put their tradeoffs and their small gains in front of anyone who reads release notes.
Corti let the clinician’s correction overrule the recording. Syllable and Artera started testing their voice agents before patients hear them. In each case something that used to be private or assumed became visible and checkable.
That is the direction healthcare AI has to travel, and this week a surprising number of vendors chose to travel it in public. The category is still mostly claims. The claims are starting to arrive with their homework attached.
Which healthcare AI vendors publish bias evaluations and model cards?
Very few, and the gap sits between saying and showing. Of the 545 vendors the AI Health Index has assessed on AI Governance and Bias Disclosure as of September 25, 2026, 12 grade A. An A requires something a reader can retrieve and check. That means a bias or fairness evaluation with a stated method, performance compared across the groups the product serves, or an independent audit of how the model behaves.
Another 114 grade B. They publish a governance framework with a named process behind it, such as certification to an AI management standard, or evaluation material written for a customer’s review committee. That is real work. It describes how the vendor governs, though, rather than how the model performs for different patients.
The remaining 419, 316 at C and 103 at D, publish principles with no evaluation behind them, or nothing at all. A responsible AI page grades C, and C is where most of the index sits.
This week showed what the missing artifact looks like when someone builds it. Doximity’s benchmark carries a health equity sub benchmark and a held out test split. Spring Health published a safety rubric and invited the public to argue with it. Both are evaluations that could have come out badly, which is the property that makes an evaluation worth reading.
The question that separates the bands is short. Ask for the subgroup results, the population they were measured on, and who ran the analysis. A vendor graded A here can usually send all three the same day, because they already exist.
Full grades, the axis definition and what separates each band are on the AI Governance and Bias Disclosure page.
The AI Health Index Brief is published weekly by AI Health Index, an independent reference for evaluating AI vendors in healthcare. No vendor pays for inclusion, placement, or rating. Compare any indexed vendors by capability at Compare and read the evaluation standards at Methodology.