insitro
Machine learning driven drug discovery and development company founded in 2018 by Daphne Koller, self described as the AI therapeutics company built on causal biology. The approach converges in house generated multimodal cellular data (induced pluripotent stem cells, genome editing, high content cellular phenotyping) with high content human cohort data, using machine learning to identify genetic drivers and prioritize targets rather than to generate molecules alone.
The POSH platform was validated in Nature Communications in December 2025, with the company's stated finding that self supervised models trained on unbiased cellular morphology can reconstruct gene function and causal relationships without being told what to look for. In January 2026 insitro acquired CombinAbleAI and launched TherML, a modality agnostic design platform that uses a physics informed optimization engine pretrained on over 100,000 molecular dynamics surrogates for complex biologics including multispecific antibodies and T cell engagers, and proprietary Quantitative Adaptive Libraries to map chemical space for small molecules.
Programs are concentrated in metabolic disease and neuroscience. The company has raised approximately 800 million dollars including approximately 150 million from non dilutive pharma partnerships, with named collaborations including Bristol Myers Squibb, Eli Lilly and Genomics England.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
Machine learning is the mechanism rather than an analysis layer over conventional biology. The company describes itself as built on causal biology, generating an integrated multimodal corpus of human and cellular data specifically so that models can learn from it, and the stated finding from the POSH platform is that self supervised models trained on unbiased cellular morphology reconstruct gene function and causal relationships without being told what to look for, which is a claim only a model can make. The January 2026 TherML platform extends model driven design across modalities. The wet lab exists to produce training data and to test model output.
Oversight is structural and appropriate to the stage: model generated hypotheses about genetic drivers and targets are tested against purpose built in vitro disease models using induced pluripotent stem cells, genome editing and high content phenotyping, so predictions meet experimental evidence before any programme advances.
Held at B because no disclosure was located on how target prioritization decisions are made between model output and committed programme spend, or what confidence is required to advance, so the human decision points are inferred from the platform description rather than documented.
Core method submitted to peer review rather than described in marketing terms, which is the standard this index applies. The POSH platform was validated in Nature Communications in December 2025 with named authors including the founder, so the approach can be read and contested independently.
Architectural detail elsewhere is unusually specific: the TherML biologics engine is described as physics informed and pretrained on over 100,000 molecular dynamics surrogates to predict protein structure and flexibility, and the small molecule path uses proprietary Quantitative Adaptive Libraries to densely map chemical space and generate high resolution local training data. Terminology is used precisely rather than decoratively.
The strongest human data pattern found in this category, and it is worth stating as a general principle rather than a company fact: the models moved to the data instead of the data moving to the vendor. Rather than extracting a national dataset, the company deployed its embedding search capability inside a secure national research environment and made it available to that organisation's research partners there, applied to almost one hundred and fifty thousand whole genomes with linked phenotypic data from rare disease and cancer patients plus histopathology images.
Nothing left the environment. That inverts the arrangement this index usually finds, where a vendor's value depends on assembling a corpus in its own estate, and it is the pattern a national or institutional data holder should ask for by name. Held below the top grade because the good pattern rests on one partner's governance rather than on a published commitment by this company.
No vendor side policy on consent provenance, retention, permitted secondary use or model training rights was located, so a different partner with weaker governance would not necessarily get the same arrangement, and nothing obliges the company to repeat it.
Ask whether the in environment pattern is the company's standing position or a term that partner required, what the vendor retains from such a deployment, and whether model representations derived inside a partner's environment can leave it.
Platform evidence is genuine and peer reviewed while clinical evidence is still absent, and buyers should hold those separately. The POSH validation in Nature Communications is real third party reviewed evidence that the method does what is claimed at the biology level, which places this above peers whose platform claims rest on press releases.
Commercial validation is also meaningful: approximately 150 million dollars of the roughly 800 million raised is non dilutive money from pharmaceutical partnerships including Bristol Myers Squibb and Eli Lilly, meaning partners paid for delivered work. What was not located is any programme with human data, and coverage frames the current period as the test of whether the data flywheel translates into clinical impact. No molecule from the platform with efficacy in patients was found in this review.
This axis genuinely applies here, unlike the chemistry platforms in this category, because the work runs on human cohort data at scale. The Genomics England collaboration is the most instructive disclosure in the category on data handling and it points the right way: rather than extracting a national dataset, insitro deployed its embedding search capability inside the secure Genomics England Research Environment and made it available to that organization's research partners there, so the models moved to the data instead of the data moving to the vendor.
Applied to almost 150,000 whole genomes with linked phenotypic data from NHS rare disease and cancer patients plus histopathology images, that architecture is the strongest human data pattern found in this category. Held at B rather than A because no vendor side policy on consent provenance, retention, permitted secondary use or model training rights was located, so the good pattern rests on one partner's governance rather than on a published commitment.
Not applicable as framed. The most significant human data relationship located runs through Genomics England under United Kingdom research governance and a secure trusted research environment rather than under HIPAA, and pharmaceutical partners contract as research collaborators rather than as covered entities. Buyers should note that the absence of a HIPAA posture here reflects the shape of the business, not an unmet obligation.
No attestation was located and no trust centre or security page exists. The company publishes a privacy policy; its scope was not examined in this review, and under this index's reading a website privacy instrument is not a data processing disclosure until its subject matter is checked.
Two retrieval notes belong here. A general query pairing the company name with certification terms returned nothing about this company at all, only compliance vendor marketing, which is a retrieval failure rather than evidence of absence. And this company's platform names collide with unrelated consumer and enterprise products, so results must be checked against its own domain before anything is credited.
One piece of real evidence sits outside the usual sources and it is the strongest thing on this row. The company deployed its embedding search capability inside Genomics England's secure research environment, for use by that organisation's own research partners, rather than having the data moved out to it. Admission into a national genomics service's controlled environment is not a certificate, but it is a demanding assessment by a third party with strong incentives and statutory duties, and a vendor that has passed it has been examined in a way most of this category has not.
What that does not establish is the security of this company's own estate, where its proprietary models, partner collaboration data and internally generated cellular and genomic datasets live. A buyer should ask what the company holds in its own name, whether the controls that satisfied the trusted research environment apply across the wider platform or only to that engagement, and what its position is on independent attestation.
The discovery platform is not a regulated device and is not presented as one, correctly. At asset level no cleared IND, registered clinical trial or human dosing was located in this review, with programmes in metabolic disease and neuroscience described as advancing toward the clinic. Regulatory standing is therefore materially behind category peers that have dosed patients, and a buyer should treat clinical readiness claims as forward looking.
No governance framework, model card or cohort composition disclosure was located, and a search directed at the scientific literature rather than product pages did not change that. The same approach overturned absence findings for several peers in this category, so this is a tested absence.
What the search did establish is that the concern is not this index's invention. Ancestry imbalance in human genetics is a first order, extensively documented and actively researched problem, with published quantification of how heavily discovery cohorts skew toward European ancestry, how poorly risk prediction transfers between populations, and why effect sizes and linkage structures differ across groups. A company whose entire proposition is discovering genetic drivers of disease operates inside that problem rather than adjacent to it.
The exposure here is compounded because it enters from two directions. Induced pluripotent stem cell lines carry the genetic background of their donors, so a cellular platform inherits whoever those donors were. And the human cohorts the company works with have their own documented ancestry composition that does not mirror global populations. A driver absent from the cohort cannot be found, and a target validated in one ancestral background may not generalise to another.
One piece of the company's own language should be read carefully rather than credited. Describing its cellular imaging as capturing unbiased morphology refers to unbiased feature selection, meaning the system does not decide in advance which visual features matter. It says nothing about unbiased sampling of who the cells came from. Those are different claims and the first does not imply the second. Ask for the ancestry composition of the cell lines and cohorts behind each programme, and whether any target has been checked for transferability.
The core method was submitted to peer review rather than described in marketing terms, which is the standard this axis applies. The platform was validated in a peer reviewed journal with named authors including the founder, so the approach can be read, reproduced and contested by people with no commercial relationship to the company, and a customer evaluating it has something to argue with rather than a claim to accept.
Architectural detail elsewhere is unusually specific and uses its terminology precisely rather than decoratively: the biologics engine is described as physics informed and pretrained on more than one hundred thousand molecular dynamics surrogates to predict protein structure and flexibility, and the small molecule path uses adaptive libraries to densely map chemical space and generate high resolution local training data.
A reader who knows the field can tell what is being claimed and where it would break. Held below the top grade because nothing attaches to the claim commercially. No warranty, indemnity, service level or remediation commitment was located, and no published error characteristic exists for the deployed system as distinct from the validated method, which are different artefacts once a method becomes a product. Ask what the deployed platform's performance is against the published validation, and how often it is re measured.
Not applicable. This is a discovery platform with no provider workflow surface and no EHR touchpoint. Integration work is directed at multimodal research data, including the embedding search capability delivered into the Genomics England Research Environment for use by that organization's research partners.
Better disclosed than the partnership only norm in this category because one deployment pattern is public and unusual: capability delivered into a partner's secure research environment, as with the embedding search engine made available inside the Genomics England Research Environment to that organization's research network. That is compute moving to the data rather than the reverse.
Held at B because this is one disclosed arrangement rather than a published deployment option, no general tenancy, hosting or regional residency terms were located, and the core discovery platform is operated internally rather than deployed to customers.
Capital structure is disclosed with useful precision, at approximately 800 million dollars raised including approximately 150 million dollars of non dilutive money from pharmaceutical partnerships, and the split matters because non dilutive partner funding is a harder signal than venture capital. Named collaborations include Bristol Myers Squibb, Eli Lilly and Genomics England.
Coverage notes explicitly that deal terms are not uniform across partnerships, varying by therapeutic area, target novelty and collaborator strategy, which is honest but leaves per programme economics undisclosed. No rate card exists and none would apply, since the commercial surface is multi target discovery partnership.
Deliberately concentrated on indication and newly broad on modality, which pulls in opposite directions. Disclosed programmes focus on metabolic disease and neuroscience, a narrow therapeutic footprint by category standards. Modality coverage widened materially with the January 2026 TherML launch, which handles complex biologics including multispecific antibodies and T cell engagers alongside small molecules within one platform.
Buyers outside metabolic disease and neuroscience should treat applicability as unproven, since the platform's causal biology approach is tied to the specific cellular disease models the company has built.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not applicable
|
Multi target discovery partnership with upfront payments, research funding and milestones. No software licence is sold. | — | Not applicable. | Vendor Published |
No price exists because the commercial surface is partnership rather than product. Capital disclosure is more useful than most here: approximately 800 million dollars raised in total, of which approximately 150 million dollars is non dilutive money from pharmaceutical partnerships, and that split is the informative number since non dilutive partner funding indicates work delivered rather than belief purchased.
Named collaborations include Bristol Myers Squibb, Eli Lilly and Genomics England. Coverage states explicitly that deal terms are not uniform across partnerships, varying by therapeutic area, target novelty and collaborator strategy, so no representative structure can be inferred from any single agreement.