Genesis Molecular AI
Stanford spinout, now trading as Genesis Molecular AI, building GEMS, a small molecule discovery platform its own chemists and its partners' programs run on. GEMS combines generative models for molecule ideation, multitask models for absorption, distribution, metabolism and excretion prediction, and agents that orchestrate routine steps while chemists interrogate predictions and decide what to synthesize.
Its foundation model, Pearl, released in October 2025, predicts the three dimensional structure of protein ligand complexes, is trained on physics based synthetic data proprietary to the company alongside public data, and can be fine tuned on a program's own data or conditioned on what a chemist already knows about a target.
In June 2026 the company reported that its Pearl system reached a 78 percent success rate on the OpenBind Consortium's public benchmark across 802 complexes in a zero shot setting, using production settings and a template structure postdating the model's training cutoff, ahead of the six cofolding models and the docking methods that benchmark evaluated. The stated focus is targets that are biologically validated but hard to drug because the chemistry is difficult.
Disclosed collaborations: Genentech; Eli Lilly, reported at up to $670 million with $20 million upfront; Gilead, $35 million upfront across three targets with an option to nominate more at a predetermined per target fee; and Incyte, initiated in 2025 and expanded in 2026 with $150 million in total upfront consideration including a $40 million equity investment, the expansion also providing for Incyte to share significant experimental data for training the platform. Total capital raised exceeds $280 million. The company renamed from Genesis Therapeutics in late 2025 and its site is now at genesis.ml.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The models are the product. GEMS is described as the operating system the company's own chemists run their whole design workflow on: models generate candidate molecules, Pearl predicts the three dimensional structure of the protein ligand complex, multitask models predict absorption, distribution, metabolism and excretion properties, and agents orchestrate the routine steps so chemists decide what to make.
The company describes itself as building world models at molecular scale and states that the platform powers both its own programs and its pharmaceutical partners'. Remove the models and there is no platform, no partnership offer and no pipeline.
The oversight structure is described and the threshold is missing. The platform's own account of the workflow puts the chemist at the decision: models generate candidates and predict properties, agents orchestrate the routine steps and surface trade-offs, and the chemist interrogates the predictions, visualizes the structures, filters and sorts, and decides what to synthesize. The company describes agents running design workflows continuously alongside its drug hunters rather than in place of them, and every prediction is followed by synthesis and assay before it counts.
What is not published is where the system should not be trusted: no confidence measure is described with a value at which a prediction is set aside, no hit rate for generated molecules is given, and nothing states which steps agents complete without a person. Ask what confidence accompanies a predicted structure or property, and which agent actions require sign off.
The models are named and the approach described, without version or update discipline. Pearl is named as the foundation model for protein ligand cofolding, stated to be trained on physics based synthetic data proprietary to the company alongside public data, and able to be fine tuned with program specific data and conditioned on what a chemist already knows about a target; GEMS is the platform around it, combining generative models for molecule ideation, multitask models for absorption, distribution, metabolism and excretion prediction, and agents that orchestrate steps. The company publishes a model paper, a benchmark evaluation describing its inference time scaling and its physics and AI based pose ranking, and work on model distillation.
What is not published is version discipline: the company writes of a "production version" and says it has continued improving the model since the paper, with no version identifier, release note or statement of how a partner learns the system changed under a running program. Ask for version identifiers and release notes, and for what the synthetic training data consists of.
A structurally short chain: the models are the company's own and are named. Pearl and the GEMS models are built and run by Genesis, trained on proprietary physics based synthetic data and public data, and the platform is operated by the company rather than delivered to partners, so no outside model provider is named between a partner's target and a prediction.
What is published about data flowing the other way is unusual and belongs here: the 2026 expansion of one pharmaceutical collaboration includes that partner sharing significant experimental data for use in training the platform. That is a stated term rather than an inference, and it is the question every other partner should ask, because nothing published says whether their data is treated the same way, what governs it, or what happens to a model trained on it when a collaboration ends. No hosting provider or subprocessor list is published either. Ask for the subprocessor list, the hosting arrangement, and the training and termination position on partner data.
Published research short of independent validation, with named partners and no candidate in the clinic. The evidence that can be checked is computational: the Pearl system's 78 percent success rate on 802 complexes from an outside consortium's public benchmark, reported against the six cofolding models and the docking methods that benchmark evaluated, in a zero shot setting with production settings and a post cutoff template. The company also published the Pearl model itself in October 2025 and follow on work on model distillation.
Adoption is named and dated: collaborations with Genentech, Eli Lilly, Gilead and Incyte, the last initiated in 2025 and expanded in 2026. No molecule designed on the platform is reported in a clinical study, and the benchmark result measures structure prediction rather than a program outcome, so what the platform does to a discovery timeline is not evidenced. Ask which partnered programs have advanced and to what stage, and for the benchmark repeated on a partner's targets.
Nothing is published about how a partner's data moves through the platform. The company's site carries no privacy policy, no terms of use and no data handling statement of any kind; the footer offers navigation and a contact page and nothing else.
The one training fact on the record runs the other way: the 2026 expansion of a pharmaceutical collaboration includes that partner sharing significant experimental data for use in training the platform. That is a negotiated term for one partner, not a statement of what happens to anyone else's structures, assay results or program data, and nothing published states retention, permitted use, segregation between programs, or what becomes of a fine tuned program specific model when a collaboration ends, which matters because the platform's own description offers fine tuning on program specific data. Patient data does not enter the workflow. Safety engineering is not addressed. Ask for those terms in writing before any data moves.
No statement of status and no privacy document at all. The site publishes no privacy policy, no terms of use and no data handling statement, so there is nothing that reaches the product and nothing that reaches a website visitor either. No business associate status is stated and no scope position explains whether protected health information could enter the platform.
The work is molecular design rather than patient data processing, which makes a scope statement easy to write and its absence the finding. Ask for a written position on protected health information and for the data terms that govern a collaboration.
Nothing is published. No SOC 2 report, ISO 27001 certificate, penetration testing statement, security page, trust center, terms of use or privacy policy appears anywhere on the estate; the site is a product and news site with a contact form.
The partners named on that site place proprietary targets and, in at least one case, significant experimental data into the platform for training, and the company's engineers work inside partner teams. For a business of that shape, publishing nothing about security controls is the finding rather than an oversight to excuse. Ask what the company holds in its own name, how partner data is separated and protected, and whether it will complete a partner's security review.
No device claim is made and the platform is scoped accordingly. GEMS is discovery software used inside the company and in partner programs, not a product presented as regulated, and no clearance is claimed or apparently required. Molecules designed on it carry their own regulatory positions, held by whoever develops them; the company discloses no clinical stage asset of its own on the pages read, so there is no regulatory activity on this record to match against its marketing.
A published evaluation of model behavior with a stated method, short of results across the groups the product serves. The company evaluated its Pearl system on an outside consortium's public benchmark and published the conditions in detail: three settings (zero shot, pocket conditioned with three named anchor residues described as realistic prior knowledge rather than oracle information, and redocking against a fixed crystallographic conformation), production hyperparameters with no benchmark specific tuning, a denominator of 802 complexes, and a note that the template structure postdates the training cutoff so the result reflects generalization. Reporting the conditions that would flatter a model and separating them from the honest one is the disclosure this axis asks for.
What is absent is the breakdown: the benchmark covers a single target, and nothing is published on how the models perform across target classes, protein families or chemical series, which is where a structure prediction model trained partly on synthetic physics data would be expected to vary. There is no model card, no governance framework and no statement of intended and unsuitable use. Ask for performance by target class and for a statement of where Pearl should not be relied on.
A measured success rate for the product's core prediction is published with its method, its denominator and outside baselines. In June 2026 the company evaluated its Pearl system against the OpenBind Consortium's public structure and affinity benchmark, a curated set whose analysis covers 802 follow-on complexes against a single well characterized target, and reported a 78 percent success rate on the consortium's joint criteria (pose within 2 angstroms, a physical validity check, an interaction similarity threshold, best of 25) in a zero shot setting where the system receives only the protein sequence, the ligand's two dimensional structure and an apo template, with no binding site knowledge and no target specific tuning. It states that production hyperparameters were used with no benchmark specific tuning and that the template structure postdates the model's training cutoff, so the result is not memorization, and it reports the same comparison at a stricter one angstrom threshold against the six cofolding models and the docking methods the consortium evaluated.
The evaluation is the company's own run against an outside benchmark rather than an independent assessment of its system, and nothing attaches commercially: no warranty, remediation or fee consequence is published, and the company publishes no terms at all. Ask for the same figures on a partner's own targets, and what a collaboration commits to when predictions miss.
No clinical record integration exists: GEMS is an internal design platform used by the company's chemists and by its partners' programs, and no clinical record system is involved. For a pharma partner the record systems that matter are its research systems: electronic lab notebooks, laboratory information systems, compound registries, research data platforms and cloud data environments.
None is named as connected and no integration is claimed; the model the company describes for working with a partner is people rather than interfaces, with its engineers and drug hunters embedded in the partner's team. That is no integration evidence on this axis, which for a collaboration business records an absence rather than a defect. A partner should ask how compound data, assay results and predictions move between the two organizations.
An internal platform, with location implied rather than committed. GEMS is not delivered to partners as software: the company's chemists run their design workflow on it, and a partnership is staffed by the company's engineers and drug hunters working inside the partner's team, with the partner's programs run on the company's platform from Burlingame, California.
Nothing states where the platform runs, whether outside cloud or compute providers are used for a system training foundation models on synthetic physics data, where a partner's program data rests, or how one partner's material is separated from another's and from the company's own pipeline. Ask where the platform and any program specific fine tuned models run and are stored.
No price is published, and the deal architecture is disclosed in more detail than most private companies in this category offer. Four collaborations are described with their terms: one at $35 million upfront across three initial targets with an option to nominate further targets at a predetermined per target fee, one reported at up to $670 million with $20 million upfront, one at $150 million in total upfront consideration across an initial agreement and its 2026 expansion including a $40 million equity investment, and one undisclosed. The engagement model is also described, with the company's engineers and drug hunters working inside partner teams.
A per target fee that exists but carries no figure does not let a buyer size the cost, and the headline values are what particular partners negotiated rather than a price anyone can apply. One non monetary term deserves the same attention as the money: the 2026 expansion includes the partner sharing significant experimental data for training the platform, which is consideration in kind. Ask for the per target fee, and treat data contribution as a priced element rather than a courtesy.
Coverage is stated clearly with little validating it yet. The scope is named precisely: small and medium size molecule drug discovery, aimed at targets that are biologically well validated but hard to drug because the chemistry is difficult, including modeling how molecules bind to flexible or otherwise awkward proteins, with the platform covering generation, structure prediction, potency and selectivity and a range of absorption, distribution, metabolism and excretion predictions. Biologics are outside it.
What validates that scope is thin: the published benchmark work covers a single target, and the partnered and internal programs are named as collaborations without the target classes or therapeutic areas they cover being public. Ask which target classes the platform has been run on in production, and for performance on the hard cases the positioning rests on.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published
|
Multi target discovery collaboration with upfront payments, a predetermined per target fee for additional targets, milestones and royalties. Equity investment and partner data contribution have featured in announced deals. | — | Not published. The engagement model includes forward deployed engineers and drug hunters working inside partner teams, a staffing cost presumably reflected in research funding rather than billed separately. | Vendor Published |
No price is published, and the deal architecture is disclosed with more granularity than most private companies in this category offer: $35 million upfront across three initial targets with an option to nominate additional targets at a predetermined per target fee; a second collaboration reported at up to $670 million with $20 million upfront; a third at $150 million in total upfront consideration across an initial 2025 agreement and its 2026 expansion, with a $40 million equity component separated out; and a fourth undisclosed.
The engagement model is staffed rather than licensed, with the company's engineers and drug hunters working inside partner teams. The per target fee exists but carries no figure, so a buyer cannot size a program before contact. One non monetary term deserves the same attention as the money: the 2026 expansion provides for the partner to share significant experimental data for use in training the platform, which is consideration in kind. Ask for the per target fee, and treat data contribution as a priced element of the deal rather than a courtesy.