Genesis Molecular AI
Stanford spinout building GEMS (Genesis Exploration of Molecular Space), a small molecule discovery platform that integrates language models, diffusion models and physics based machine learning simulations to generate novel molecules and predict properties including potency, selectivity, ADMET and pharmacokinetics. The stated focus is targets that are biologically well validated but considered undruggable because the chemistry is difficult, including modelling how molecules bind to flexible or otherwise difficult proteins.
The company runs both an internal pipeline and partner programmes, staffed by forward deployed engineers and drug hunters who work inside partner teams, with the stated intent that every partnered and internal programme stress tests GEMS. Disclosed collaborations: Genentech; Eli Lilly, reported at up to 670 million dollars with 20 million upfront; Gilead, 35 million dollars upfront across three targets with an option to nominate more at a predetermined per target fee; and Incyte, initiated 2025 and expanded 2026, with 150 million dollars in total upfront consideration including a 40 million dollar equity investment.
The 2026 Incyte expansion also involves Incyte sharing significant experimental data for use in training GEMS. Total capital raised exceeds 280 million dollars, including a 200 million dollar Series B co led by Andreessen Horowitz with Fidelity, BlackRock and NVIDIA's venture arm participating. The company renamed from Genesis Therapeutics to Genesis Molecular AI in late 2025, and its published email addresses still resolve to the former genesistherapeutics.ai domain while the site is at genesis.ml.
In October 2025 it released Pearl, a generative diffusion model for protein and ligand structure prediction presented as a core component of GEMS and described as trained on physics based synthetic data proprietary to the company.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
GEMS is the whole proposition. Four major pharmaceutical companies have paid specifically for access to it rather than for chemistry services, and the disclosed architecture is model led throughout, integrating language models, diffusion models and physics based machine learning simulations to generate molecules and predict potency, selectivity and ADMET. There is no instrument business, no compound library franchise and no services arm underneath. The company's stated purpose for the platform, reaching targets that are biologically well validated but chemically intractable, is a claim only a generative model can make good on.
The disclosed oversight structure is human embedded rather than threshold based: forward deployed engineers and drug hunters work inside partner discovery teams, and the company states that those same researchers stress test GEMS predictions on every partnered and internal programme. That places experienced medicinal chemists between model output and synthesis commitment by design. Held at B because no confidence thresholds, prediction reliability bands or documented approval gates were located, so the check is organizational rather than specified.
More specific than marketing language but not independently checkable. The disclosed method names real architectural components rather than gesturing at AI generally: language models, diffusion models and physical machine learning simulations combined in one system, with stated prediction targets across potency, selectivity, ADMET and pharmacokinetics, and an origin in published Stanford research on modelling flexible protein binding.
Held at B because no peer reviewed methods paper establishing GEMS performance was located in this review, and no model cards or published benchmarks were found, so the accuracy claims rest on company communications and partner behaviour.
The components are named and the partner confidentiality question is unusually live here. Language models, diffusion models and physical simulation are identified as the constituents of the system, so a partner knows what stages their chemistry passes through even though no supplier is named for any of them, and the clinical form of this axis does not reach the workflow, which operates on molecular structures and assay data with no patient records involved.
The obligation that does exist covers partners' undisclosed targets and compound information, which in this lane is the sensitive material: a target a partner has not disclosed publicly is a strategic asset, and knowing which company is working on which target is commercially valuable to a competitor before a single molecule exists.
That obligation is made more live by a disclosed data sharing arrangement, which means material moves between parties by design rather than incidentally, and a partner should establish precisely what moves, to whom, and whether anything derived from their chemistry can inform work for others. Nothing else is enumerated: no hosting arrangement, no sub processor list, no retention position, and no statement of what happens to partner derived model artefacts at the end of an engagement. Ask what the sharing arrangement covers, what is excluded from it, and whether the exclusion is architectural or contractual.
Partner validation is genuine and unusually deep for a private company, with four major pharmaceutical collaborations from Genentech, Eli Lilly, Gilead and Incyte, and Incyte extending its deal in 2026 after seeing the platform work. Sophisticated buyers repeating business is a real signal. It is not clinical evidence.
No molecule from this platform with human data was located, the internal pipeline is described as preclinical, and the 2023 statement that the company was approaching an inflection point with its first candidates entering the clinic was not matched in this review by a disclosed IND, registered trial or dosed patient. Applying the index precedent that commercial traction does not substitute for evidence of benefit, this sits at C until an asset reaches the clinic or platform performance is published.
Not applicable in the provider sense and rated accordingly rather than penalized. The platform operates on molecular structures and assay data with no patient records in the workflow. The stewardship obligation that does exist covers partners' undisclosed target and compound information, and it is unusually live here because of the data sharing arrangement noted on the governance axis.
Not applicable. Counterparties are pharmaceutical R&D organizations entering multi target discovery collaborations, not covered entities transferring protected health information.
Converted from Not Rated after a second search. No SOC 2, ISO 27001 or equivalent attestation was located and no trust centre was found. The company is private, so the annual report route that supplies a cybersecurity governance disclosure for the listed vendors in this category does not exist here.
The exposure is concentrated rather than diffuse, and one partner arrangement makes it heavier than target structures alone. Named collaborations include Genentech, Eli Lilly, Gilead and Incyte, so a small number of large pharmaceutical organisations have placed undisclosed target selections inside the environment. Under the expanded Incyte agreement the company also receives significant experimental data from Incyte for model training, which means a partner's own generated data rather than merely its target list sits in the vendor's systems and is used to improve a shared asset. A security review here would run entirely through the collaboration agreement, and the training data provision should be reviewed alongside it rather than treated as a commercial term.
A practical diligence point specific to this company. It renamed from Genesis Therapeutics to Genesis Molecular AI, its site is now at genesis.ml, and its published email addresses still resolve to the former genesistherapeutics.ai domain. Any attestation, certificate or contract, if one exists, could be held under either name. Search both before concluding that nothing exists.
The company publishes a website privacy policy. Nothing describing the platform environment, tenancy, access control, or the handling of partner supplied training data was located.
The platform is not a regulated device and is not presented as one, correctly. At asset level no cleared IND, registered clinical trial or dosed patient was located for any Genesis originated molecule, with the internal pipeline described as preclinical and partner assets remaining under the partner's control and not publicly attributed. Regulatory standing is therefore materially behind category peers with molecules in humans, and buyers should read the company's clinical timing statements as forward looking.
Converted from Not Rated after a second search directed at publications and platform releases. Material was found and the reasoning sharpens considerably, but nothing located characterises where the models fail, so the grade stands.
Note first that the company now trades as Genesis Molecular AI, having renamed from Genesis Therapeutics in late 2025.
Retained from the prior assessment and now confirmed on the company's own current site rather than at second hand: the expanded Incyte agreement involves Incyte supplying significant experimental data for training GEMS. A partner's proprietary experimental data improving a model that then serves the vendor's other partners and its own internal pipeline is a real structural issue in this category, and no disclosure was located on what boundaries apply, whether improvements are ring fenced, or what becomes of the trained model if the collaboration ends. It is to the company's credit that the arrangement is public at all, which is what makes it a concrete thing to ask about.
The newer material sharpens the chemical representativeness question. In October 2025 the company released Pearl, a generative diffusion model for protein and ligand structure prediction. Its founder's public framing of that release is the finding here: he identifies generalisation failure as the defining weakness of competing cofolding models, saying they often fail to truly generalise and sometimes produce obvious physical errors. That makes generalisation the company's own chosen ground. No generalisation analysis of its own models was located. The performance claims are vendor generated, announced by press release, and the evaluations behind them are not identified, which is the unfalsifiable posture this index flags elsewhere in this category. Naming a failure mode as a competitor's weakness while not reporting your own performance on it is a construction worth watching for generally.
Pearl is described as trained on physics based synthetic data proprietary to the company. That relocates the representativeness question rather than escaping it. A model trained on simulation output inherits the simulator's assumptions and whichever regions of chemical space it reproduces well, and no characterisation of that coverage was located.
A public code repository organisation exists under the former company name, but the visible repositories are predominantly forks of upstream open source tools including a molecular simulation toolkit and a machine learning framework. A repository list is not a code release. Check whether the repositories are forks before crediting one.
Caution for anyone re running this search. The acronym GEMS collides with at least two unrelated things in adjacent fields, one of them an academic graph neural network for binding affinity prediction published in a Nature portfolio journal with an open repository and a released leakage corrected training dataset. That artefact is precisely the kind of evidence this axis rewards and it belongs to a different group entirely.
The architecture is named in real components rather than gestured at, which is the basis for the grade. Language models, diffusion models and physical machine learning simulations are described as combined in one system, with prediction targets stated individually across potency, selectivity, absorption and disposition properties, and pharmacokinetics, and an origin in published academic research on modelling flexible protein binding.
Naming the prediction targets separately matters, because those are different problems with different data availability and different failure modes: potency prediction has abundant training data and disposition properties do not, so a system that performs well on one may perform poorly on another, and a single platform claim would conceal that. Held at C because nothing is measured.
No peer reviewed methods paper establishing platform performance was located, no model cards or published benchmarks were found, and no warranty, indemnity or remediation commitment attaches, so the accuracy claims rest on company communications and on partner behaviour.
That second form of evidence deserves a word, because it recurs in this lane: sophisticated partners continuing to sign is a market signal and not a performance figure, since a deal reflects an assessment made privately under terms nobody outside can see, and it is compatible with the platform working, with it being cheap to try, and with the partner hedging. Ask for accuracy per prediction target.
Not applicable. This is a preclinical discovery platform with no provider workflow surface and no EHR touchpoint.
No software is deployed to a customer and no tenancy, hosting or residency terms were located, so the axis does not apply in its usual form. The disclosed model is an unusual hybrid worth noting: rather than shipping software or keeping everything in house, the company embeds forward deployed engineers and drug hunters inside partner teams, so people move to the partner while the platform stays with the vendor. That resolves some working proximity concerns without resolving where partner data ultimately sits.
Deal structure is disclosed with more granularity than most private companies in this category offer, and much of it on the company's own site rather than only in trade coverage: Gilead at 35 million dollars upfront across three initial targets with an option to nominate more at a predetermined per target fee; Eli Lilly reported at up to 670 million dollars with 20 million upfront; Incyte at 150 million dollars in total upfront consideration across the 2025 collaboration and 2026 expansion, with the 40 million dollar equity component separated out. Total raised exceeds 280 million dollars with the investor syndicate named. What is absent is any rate card, which is expected here, and per programme economics beyond the upfront figures.
Concentrated on both dimensions that matter. Modality is small molecules only, with no biologics capability disclosed. Therapeutic focus is oncology for the internal pipeline, described as several preclinical programmes, plus small molecule programmes against well validated immunology and autoimmune targets where biologics have shown efficacy but oral options do not exist, which is a coherent and specific thesis rather than broad coverage. Partner programmes extend the applied range but against targets the partners select and do not disclose. Buyers outside small molecule work should treat this platform as out of scope.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published
|
Multi target discovery collaboration with upfront payments, per target fees, milestones and royalties. Equity investment has featured in at least one deal. | — | Not published. The engagement model includes forward deployed engineers and drug hunters working inside partner teams, a staffing cost presumably reflected in research funding rather than billed separately. | Vendor Published |
No rate card exists and none would apply, since the commercial surface is multi target discovery collaboration. Deal structures are disclosed with more granularity than most private companies in this category offer: Gilead at 35 million dollars upfront across three initial targets with an option to nominate additional targets at a predetermined per target fee; Eli Lilly reported at up to 670 million dollars with 20 million upfront; Incyte at 150 million dollars in total upfront consideration across the initial 2025 collaboration and its 2026 expansion, with the 40 million dollar equity component separated out; and Genentech undisclosed.
One non monetary term in the 2026 Incyte expansion is worth as much attention as the money: Incyte is sharing significant experimental data for use in training GEMS. Buyers negotiating here should treat data contribution as a priced element of the deal rather than a courtesy.