Layer Health
Layer Health is a Massachusetts Institute of Technology spin out applying large language models to clinical chart review, founded in 2023 and based in Boston. Its co founders are David Sontag, an MIT professor of machine learning in healthcare who serves as chief executive, with Monica Agrawal, Luke Murray and Divya Gopinath, alongside Steven Horng, an emergency physician and clinical informatician at Harvard Medical School.
The first product is called Distill. Rather than summarising a chart for a clinician about to walk into a room, it performs structured abstraction: reading unstructured notes, labs and imaging across a patient's record and extracting the specific fields a downstream process needs, with each extracted field linked back to the evidence in the record that produced it. The stated applications are clinical registry submission, quality measurement, curation of real world evidence, clinical documentation improvement and revenue cycle. The company states the platform works without requiring labelled training data, which is the usual bottleneck in this work.
Registry abstraction is the wedge. Hospitals employ nurse abstractors to read charts manually and submit data to national registries in cardiovascular disease, oncology, surgery and other areas, and the work is slow and expensive. Named deployments include White Plains Hospital, Froedtert, and Intermountain Health, which took a strategic investment through Intermountain Ventures and committed to a multi year deployment across 33 hospitals starting with stroke, bariatric surgery and cardiovascular registries. The company also works with the American Cancer Society on abstraction of patient data drawn from across the United States for research.
It raised a 4 million dollar seed round in November 2023 from GV, General Catalyst and Inception Health, followed by a 21 million dollar Series A in March 2025 led by Define Ventures with GV and Flare Capital Partners.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
Large language models reading unstructured clinical text are the entire product. There is no workflow layer, network asset or transport business underneath that would function without them, and the task being automated, reading a chart and judging what it says, is irreducibly a language problem.
The claim that the platform works without requiring labelled training data is the sharpest expression of this. Labelling is the cost that has historically confined clinical natural language processing to narrow, expensively curated tasks, and a system that generalises without it is making a claim about model capability rather than about data assembly.
The oversight architecture is unusual and better than most, because it puts validation in the customer's hands before the system does any real work. Both published deployments describe the same pattern: the hospital's own clinical data management team or abstractors first validate the model's accuracy against their own standard, and only then is it deployed. At Intermountain that validation is described as a prerequisite to ensure the system meets the performance standard needed for real registry reporting.
That is a meaningfully different posture from asking a buyer to accept vendor published accuracy. It also fits the work, since registry abstraction has a gold standard, namely what a trained human abstractor would have recorded, which most clinical AI tasks lack. Held at B because no confidence threshold, exception routing rule or ongoing monitoring cadence after go live is published.
There is a genuine tension on this record and the grade reflects the product rather than the founders. The research lineage is as strong as anything in this index, with a machine learning professor as chief executive and a founding team from MIT, Harvard, Microsoft and Google, and that community publishes heavily.
The product does not. No architecture, no model description beyond large language models, no accuracy figure, no evaluation methodology and no peer reviewed assessment of the platform was located. The company states that its artificial intelligence performs better than a human at complex chart abstraction, which is a strong and testable claim carrying no published denominator, task definition or comparison protocol. Re verify, because a team with this background may well have published work not surfaced in this pass.
One control is real and one headline claim says less than it appears to. The control is field level evidence linking: an abstracted value carries a pointer to the note it came from, so it can be checked, and a fabricated one has nowhere to point. That is a structural property rather than a policy, and for an abstraction product it is the right one, because the characteristic failure is a value that looks plausible and has no source.
An audited attestation and stated compliance give a reasonable baseline around it. The headline claim needs reading carefully. The company states that the platform works without labelled training data, which is a claim about not needing human annotation, and it leaves entirely open what the system was trained or tuned on instead and whether customer chart data contributes.
No labelled data is not no training data, and a reader who takes the first for the second has been misled by their own inference rather than by the sentence. Nothing published states retention, de identification, or the boundary between one health system's data and another's, and for a vendor holding longitudinal records across several large systems that boundary is the thing to establish in the contract rather than to assume from the architecture. Ask what the models were trained on, whether customer data contributes, and how customers are separated.
Adoption is real and named rather than anonymous: White Plains Hospital, Froedtert, Intermountain Health across 33 hospitals, and the American Cancer Society for research abstraction. All major registries are stated to be supported, including cardiovascular, oncology and national surgery.
What is missing is any published result. No accuracy figure, no comparison against human abstractors on a defined task, no error rate by registry or field type. The pre deployment validations run by customers are exactly the studies that would answer this, and their outcomes are not published, which is a shame because they are being conducted anyway. Deployment breadth is not accepted here as a substitute for evidence of benefit.
SOC 2 Type II certification and stated compliance with United States health privacy law give a reasonable baseline, and the field level evidence linking is itself a safety feature: an abstracted value that carries a pointer to the note it came from can be checked, and a fabricated one has nowhere to point.
The unaddressed question is what the models learn from. The company's headline technical claim is that the platform works without labelled training data, which is a claim about not needing human annotation and leaves open what the system was trained or tuned on instead, and whether customer chart data contributes. Nothing published on retention, de identification or the boundary between one health system's data and another's. For a vendor holding longitudinal records across several large systems, that boundary is the thing to establish in the contract.
Compliance with United States health privacy law is asserted in company announcements alongside the security certification. No body certifies HIPAA compliance, so that assertion reports intention rather than an audited finding, and it is graded accordingly.
No published business associate agreement posture, notice or compliance documentation was located. The customer base of large academic and regional health systems implies these terms are negotiated and in place, since no organisation of that size deploys across 33 hospitals without them, but the public record does not show it.
SOC 2 Type II certification is stated in company announcements with the report type named rather than left as an unqualified claim, which is the disclosure most vendors skip and it earns the grade.
Held below A because that is the whole of it. No trust centre or security portal was located, no penetration testing statement, no ISO 27001 or HITRUST certification, and no indication of whether the report can be requested by a prospective buyer. For a company selling into large health systems whose vendor risk processes will demand the report, making its availability visible would be a cheap improvement.
No clearance and none needed. Abstracting a chart into a registry field is a data operation, not a diagnostic or treatment recommendation, and it sits outside device regulation.
The boundary is worth watching rather than dismissing. The chief executive has stated that a forthcoming module will address real time clinical decision support to automate clinical care pathways. A system that reads a chart and tells a clinician what to do next is a materially different regulatory object from one that reads a chart and files a registry field, and the same underlying technology can cross that line without the company changing anything a buyer would notice.
Nothing published on subgroup performance, model monitoring or how errors are detected after deployment, and the failure mode here is distinctive enough to name precisely.
Registry data does not stay in the registry. It feeds national quality measurement, public hospital rankings, reimbursement adjustments and downstream research, including the American Cancer Society work drawing on patients nationwide. If abstraction accuracy varies with documentation style, note length, specialty, language or site, then the error does not surface as a wrong answer to a clinician who can correct it. It is written into a permanent research and quality record and everything computed from it inherits the distortion silently. That makes unpublished performance a more consequential omission here than in a product a human reviews in real time.
There is a genuine tension on this record and the grade reflects the product rather than the founders. The research lineage is as strong as anything in this index, with a machine learning academic as chief executive and a founding team drawn from leading research universities and technology companies, and that community publishes heavily as a matter of professional practice. The product does not.
Two passes located no architecture, no model description beyond a general category, no accuracy figure, no evaluation methodology and no peer reviewed assessment of the platform, and no warranty, indemnity or remediation commitment.
One claim is strong and testable and carries none of what would make it testable: the company states that its artificial intelligence performs better than a human at complex chart abstraction, with no published denominator, task definition or comparison protocol. Better than which humans, at which abstraction tasks, measured how, against what reference.
A team from this background knows exactly how such a comparison is designed, which makes the absence a choice about publication rather than a limitation of capability, and it is the sort of claim that would be straightforward for them to substantiate. Re verify at refresh, because work by this team may exist that this pass did not surface. Ask for the comparison protocol, the abstraction tasks tested, and the error rate on each.
The deployments demonstrate the capability even though the mechanism is thinly described. Running across 33 hospitals in a single system, and separately at two other health systems and a national research organisation, requires real access to charts at scale rather than a document upload workflow.
Delivery is described as application programming interface access or modules integrated with the record, with plug in support for major registries and a standalone abstraction workflow as an alternative. No named integration, interface standard or write back mechanism was located, and registry submission itself involves a second set of interfaces to the registry bodies which is not described at all.
Not described. No statement of hosting model, region, on premise or private cloud option, or retention schedule was located.
The question matters more than usual for this product because the workload is bulk historical chart reading rather than one patient at a time. Whatever the architecture, a large volume of longitudinal records from named health systems is being processed somewhere, and the public record does not say where or for how long it persists.
Nothing published. No price, no unit of sale, and no indication whether the model is per chart abstracted, per registry, per hospital or per seat, which matters because those scale very differently for a system running 33 hospitals against several registries.
The comparison a buyer actually needs is against the current cost of the nurse abstractors doing this work manually, which is knowable inside any hospital, so this is a category where a published per chart rate would be unusually easy to evaluate and unusually persuasive. A third party roundup describes pricing as enterprise and undisclosed, consistent with the company's own silence.
Broad within a narrow function. The buyer is a health system quality, registry or clinical data management team rather than a clinician, and the same platform is stated to support all major registries including cardiovascular, oncology and national surgery, with named early work in stroke, bariatric surgery and cardiovascular disease.
That spread is meaningful because each registry has its own definitions, inclusion rules and field specifications, so covering several is closer to covering several products than to configuring one. A second and different setting exists in research, where a national organisation uses the same abstraction against patient data drawn from institutions across the country. Nothing addresses ambulatory practices or the point of care, which is a stated future direction rather than a current capability.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
—
|
Not published. Enterprise sale to health systems and research organisations. | Not published. Compliance with United States health privacy law is asserted in company announcements but no agreement posture was located. | Not published. Both named deployments describe a customer run accuracy validation phase before go live, which is a real project cost in the customer's own staff time whether or not the vendor charges for it. | Third Party Estimated |
Nothing is published by the vendor. A third party roundup describes the pricing as enterprise and not publicly disclosed, priced for health systems and large service lines, which is consistent with the company's own silence. The unit of sale is the open question and it matters more here than usual: per chart abstracted, per registry, per hospital and per seat scale very differently for a customer running several registries across 33 hospitals, and a buyer cannot tell from public material which curve they are on.
This is also a category where the alternative cost is unusually knowable, since the work is currently done by employed nurse abstractors whose salaries the hospital already tracks, so a published per chart rate would be straightforward for a buyer to evaluate. One structural note recorded neutrally: Intermountain Health is both a named customer and, through its venture arm, an investor, so that reference is not fully arm's length.