Truveta
Truveta is a collective of 30 United States health systems that pooled their patient records into one research dataset and own the company that runs it. Founded in 2020 and based in Bellevue, Washington, led by co founder and chief executive Terry Myerson, it holds de identified electronic health record data representing more than 120 million patients. Members include Providence, Advocate Health, Trinity Health, Northwell Health, CommonSpirit Health, Tenet Healthcare, Henry Ford Health, Ochsner Health, MedStar Health, Memorial Hermann and twenty others.
The artificial intelligence is the Truveta Language Model, described as a large language multi modal model that normalises billions of data points from thirty different record estates into a single consistent structure, running on Microsoft Azure. Microsoft has invested in the company. Normalisation is the hard part of this business: the same diagnosis, drug and laboratory result are recorded differently in every system, and research is only possible once they agree.
The company positions the result as regulatory grade safety and effectiveness data that can replace slow and expensive clinical trials and registries, and published work using it spans cardiovascular, neurology, oncology and metabolic research, including studies of GLP-1 medicines and methodological comparisons between claims and record data.
In January 2025 members launched the Truveta Genome Project with the Regeneron Genetics Center and Illumina, which will sequence the exomes of an initial ten million volunteers and is described as more than ten times the scale of previous efforts, with explicit representation across ancestries, ethnicities, sex and social drivers of health as a design goal. Consent is obtained at the point of care to use leftover biospecimens from routine laboratory tests, which are sent for sequencing and linked to de identified records, with remaining material stored for future multi omics work. Health systems, Regeneron and Illumina invested 320 million dollars in preferred equity at a valuation above one billion dollars, Regeneron contributing 119.5 million and Illumina 20 million.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The model does work that nothing else can do at this scale. Thirty health systems record the same diagnosis, medicine and laboratory result in thirty different ways, and research is impossible until those agree. The Truveta Language Model exists to normalise billions of such data points into one consistent structure, and it is named and described rather than implied.
Held at B because the decisive asset is the collective. Thirty health systems agreeing to contribute their records into a shared entity is a governance and contracting achievement, not a technical one, and a competitor with a better model and no members would have nothing to normalise. This is the moat is the network case in its strongest form: the model makes the asset usable, the asset is why the model matters.
Low autonomy and appropriately so. Researchers pose questions and run studies; the platform prepares and serves data. Nothing acts on a patient and no clinical decision is automated.
The consequential automated act is normalisation itself. A model deciding that two differently recorded things are the same thing, at a scale no person could review, determines what every downstream study sees. That judgement is invisible in the results and there is no described mechanism for a researcher to inspect or challenge it for a specific variable.
Above the norm for a data business. The model is named, characterised as large language and multi modal, its job is stated plainly as normalisation, and the infrastructure it runs on is identified. Scale is quantified at more than 120 million patients across thirty named health systems rather than described as vast.
The company also publishes methodological work rather than only findings, including comparisons of what claims data and record data each capture, which is the kind of material that helps a reader understand the limits of the dataset.
What is absent is any accuracy measure for the normalisation. No published validation of how often the model maps a term correctly, no error rate by variable or by contributing system, and no description of how disagreements between systems are resolved. For research resting on that step, it is the number that matters most.
The process is described and consented rather than assumed, which puts this above almost every data business in the index, and the hardest unanswered question in the index sits inside it. What is stated: records are de identified, the genomics programme uses leftover biospecimens from routine laboratory tests rather than fresh collection, consent is obtained at the point of care, sequencing is performed on anonymised material, and results are linked back to de identified records.
Contributing health systems are named and the scale is quantified rather than described as vast. Using leftover specimens is a materially better position than collecting new ones, and point of care consent is a real mechanism rather than an assumed authorisation. What is not addressed is that genomic data resists de identification in a way ordinary record data does not.
A sequence is itself an identifier, and linkage attacks against anonymised genomic datasets are established in the scientific literature. A programme intending to sequence millions of exomes and store the remaining material for future multi omic work is assembling the most re identifiable dataset in medicine, and consent to unspecified future use is the hardest consent to obtain meaningfully, because neither party can describe what is being agreed to.
Nothing public describes the controls against re identification, the withdrawal mechanism, or what happens to stored material if the organisation changes hands. Ask all three, and ask what the consent form actually says about future use.
Real research output rather than customer counts. Published studies span cardiovascular, neurology, oncology and metabolic medicine, including work on GLP-1 medicines, and the company publishes methodological comparisons that examine its own data's limitations.
Held at B rather than A for two reasons. The evidence is that the data supports research, not that the platform's own machinery is accurate: no validation of the normalisation model is published. And the central claim, that this data is regulatory grade and can replace slow and expensive trials and registries, is a strong proposition in a genuinely contested area. Regulators have accepted real world evidence for specific purposes, principally label expansions and safety, and not as a general substitute for randomisation. The claim should be read as a direction of travel rather than a settled position.
More disclosed process than most, and the hardest unanswered question in this index sits here.
What is stated: records are de identified, the genomics work uses leftover biospecimens from routine laboratory tests rather than fresh collection, consent is obtained at the point of care, sequencing is performed on anonymised material, and results are linked back to de identified records.
What is not addressed is that genomic data resists de identification in a way ordinary record data does not. A sequence is itself an identifier, and linkage attacks against anonymised genomic datasets are established in the scientific literature. A programme intending to sequence ten million exomes and store the remaining material for future multi omics work is creating the most re identifiable dataset in medicine, and no public material located here describes the controls against that, the withdrawal mechanism, or what the consent covers when the future use is by definition unspecified.
Graded B because the process is described and consented rather than assumed. The questions above should be asked directly.
Graded on an honest basis, and the frame genuinely differs here. Once data is de identified to the standard health privacy law sets, that law no longer governs it, which is the legal basis on which datasets like this exist.
So the operative controls are the member agreements between thirty health systems and the entity they jointly own, plus the consent obtained for the genomics programme, rather than a business associate arrangement. None of those documents is public. A health system considering membership is negotiating something closer to a joint venture than a vendor contract.
Recorded honestly and provisionally: the dedicated trust and security search this index requires was not run in this pass, and no attestation was encountered incidentally.
The concentration is worth noting without alarm. One entity holds de identified records for more than 120 million people and intends to hold linked exome sequences for ten million of them, on a single cloud platform whose provider is also an investor. That is among the largest single accumulations of health data in existence and its protections were not retrieved.
No device pathway applies and none is claimed. This is research infrastructure.
The regulatory context is nonetheless central to the business, because the company's positioning rests on producing evidence regulators will accept. Federal frameworks for using real world evidence in regulatory decisions do exist and have been used for specific purposes, and the standard for what counts is set by the regulator rather than by the data holder. A reader should treat regulatory grade as a description of the intended standard rather than a status conferred by anyone.
The genomics programme sits under human subjects research oversight, consent and institutional review at the contributing sites, which is a different regime again.
Representativeness is an explicit design objective here rather than an afterthought, which is rare enough in this index to credit.
The genomics programme states comprehensive representation across ancestries, ethnicities, sex and social drivers of health as a purpose, at more than ten times the scale of previous efforts, and the thirty member systems span a deliberately broad cross section of the country. This index has repeatedly recorded vendors unable to report performance by ancestry because the data to do it does not exist; this is an attempt to build that data.
What is not published is any evaluation of the model that makes it usable. If normalisation performs less well on records from smaller or less standardised systems, or on documentation styles associated with particular populations, that bias enters every study built on the dataset and is invisible in the output. No error rate by contributing system, no audit of what fails to map, and no statement of how the resulting dataset's composition compares to the national population.
The model is named, characterised as large language and multimodal, and its job is stated plainly as normalisation, which is more than most data businesses disclose and which tells a reader exactly where the artificial intelligence sits in the pipeline.
The company also publishes methodological work rather than only findings, including comparisons of what claims data and record data each capture, and that kind of material is unusually useful because it helps a reader understand the limits of the dataset rather than its size. Publishing the ways your own data is incomplete is a form of limitation disclosure. What is absent is the number the whole enterprise rests on.
No published validation states how often the model maps a term correctly, no error rate is given by variable or by contributing system, and nothing describes how disagreements between systems are resolved when the same concept is recorded differently in each.
Normalisation is the step every downstream study inherits: a systematic mapping error does not produce a visible failure, it produces a study with a quietly wrong denominator, and the researchers using it have no way to detect that from the output. With more than a hundred systems' worth of variation being reconciled, per system error rates would differ. No warranty, indemnity or remediation commitment was located. Ask for normalisation accuracy by variable and by contributing system, and how conflicts are resolved.
Reconciling thirty separate record estates into one queryable dataset is the entire product, and it is a deeper interoperability problem than any live integration in this index: not moving one patient's data between two systems, but making thirty systems' vocabularies mean the same thing.
Held at B rather than A because the direction is one way and periodic. Data flows out of member systems into a research environment; nothing flows back into care, and no interface standard, refresh cadence or latency figure is published. A reader should not confuse this with a platform that operates inside the clinical workflow.
Explicit rather than inferred, which is why it clears C. The platform runs on Microsoft Azure, named as the exclusive cloud provider for the genomics programme, and Microsoft is also an investor in the company.
That relationship deserves stating plainly rather than being buried: the cloud provider holding the data has an equity interest in the entity that holds it. Held at B because no region, retention schedule or customer controlled option is published, and because biospecimen storage adds a physical residency question that cloud architecture does not answer.
No pricing is published for data access, which is standard in this segment, and the ownership structure makes the commercial question unusual enough to spell out.
Customers are life sciences, public health and academic organisations buying access to a dataset contributed by thirty health systems that also own the company and, in seventeen cases, invested in its equity at a valuation above one billion dollars. Revenue from a pharmaceutical customer therefore flows toward the same institutions that supplied the underlying records.
That is a coherent structure and a genuinely different one from a data broker. The questions to ask are what a member health system pays or receives, whether academic access is priced differently from commercial, and whether the patients whose records and biospecimens constitute the asset participate in any of it.
Very broad and confined to research. Thirty health systems spanning a deliberately wide cross section of the United States contribute records for more than 120 million patients, and published work covers cardiovascular, neurology, oncology and metabolic medicine, so the dataset is specialty agnostic by construction.
The genomics programme extends it into a second dimension entirely, linking phenotype to sequence at a scale that no single institution could reach. Held at B because none of it touches care delivery: this is infrastructure for research and regulatory evidence, with nothing operating at the bedside, and coverage is United States only.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
—
|
Not published. Data access sold to life sciences, public health and academic organisations; the contributing health systems are also the owners. | Not the governing instrument. Data is de identified, which places it outside health privacy law; the operative documents are the member agreements between thirty health systems and the entity they jointly own, plus the consent obtained for the genomics programme. None is public. | Not applicable to research customers. For a health system, joining is closer to entering a joint venture than to buying software, and the terms are not public. | Third Party Estimated |
No pricing is published for data access, which is standard in this segment, and the ownership structure makes the commercial question unusual enough to spell out. Customers are life sciences, public health and academic organisations buying access to a dataset contributed by thirty health systems that also own the company, seventeen of which invested alongside Regeneron and Illumina in 320 million dollars of preferred equity at a valuation above one billion dollars.
Revenue from a pharmaceutical customer therefore flows back toward the same institutions that supplied the underlying records. That is a coherent structure and a genuinely different one from a data broker, and it should be understood rather than assumed. Four questions. What a member health system pays for membership and what it receives, since both directions exist.
Whether academic and public health access is priced differently from commercial access, because the mission framing implies it should be. What a life sciences customer is actually buying, meaning cohort level access, a study, or a licence to the underlying data. And whether the patients whose records and leftover biospecimens constitute the entire asset participate in any of the value, which is the question this structure raises and does not answer.