Drug Discovery AI
C

Chai Discovery

AI foundation models for molecular design, licensed as software to pharmaceutical and life sciences R&D organizations. The models predict and reprogram interactions between biochemical molecules: Chai-1 (2024) for structure prediction, Chai-2 (2025) for fully de novo antibody design, and Chai-3 (2026), which the company describes as roughly doubling the success rate of its predecessor.

On Chai-2, the company reports designing all complementarity determining regions from a target and epitope prompt alone, with hit rates reported between roughly 16 and 20 percent depending on the source, against sub 0.1 percent rates the company attributes to prior methods; validation covered approximately 50 antibody targets with fewer than 20 designs tested per target. Models are reported in production at Eli Lilly, Pfizer, and Novartis, with a collaboration announced with argenx. Unlike vendors that use AI to build their own drug pipeline, Chai's product is the model itself, including custom versions trained on a customer's proprietary data.

AI Health Index verifiedJuly 28, 2026
Compare Chai Discovery with other vendors
Founded
2024
Headquarters
San Francisco, California
Categories
drug-discovery
Indexed Products
Chai-1, Chai-2, Chai-3
Buyer Segments
Pharma / Life Sciences
Assessment

Capability Axes

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read

AI Capability
AA on AI CentralityThe artificial intelligence is the product. Remove the model and there is nothing left to sell.
Vendor Published

The models are the product. Chai licenses foundation models for molecular structure prediction and de novo antibody design as software to pharmaceutical R&D organizations, including custom versions trained on a customer's proprietary data. There is no non AI version of this offering.

BB on Autonomy and Oversight ModelThe oversight structure is described and one part is missing, commonly the threshold at which the system stops or what happens after it is wrong.
Vendor Published

Substantively answered, and answered by a workflow rather than by a policy document.

The model proposes and the laboratory disposes. Published validation describes designing a bounded number of candidates per target, on the order of twenty or fewer, and testing all of them experimentally before anything proceeds. The company reports the resulting hit rate openly, on the order of sixteen to twenty percent for antibodies, which is an explicit statement that most designs fail and that failure is the expected condition. An oversight model in which every output is experimentally falsified before use, with the failure rate published, is more concrete than most oversight descriptions in this index, and the company does not claim autonomous action anywhere.

The acceptable use policy adds a second layer by presupposing safety filters in the product and prohibiting attempts to circumvent them, and access gating means the company also controls who can run the model at all.

Held at B for two reasons. The oversight is inherent in how the product is currently used rather than a control the company describes, specifies or commits to, and nothing states what the safety filters actually refuse. More importantly the company's stated long term objective is to generate biologics ready for regulatory submission in a single computational pass, which aims squarely at removing the experimental step that presently constitutes the entire oversight mechanism. Nothing located says what would replace it. A buyer should ask that question directly, and should also ask who reviews output from a custom model trained on their own data and deployed in their own programmes.

AA on Model and Technology TransparencyWhat is under the hood is named: proprietary or adapted foundation models identified, training data characterised, and versioning and update practice published so a buyer knows when the system changed.
Vendor Published

Unusually specific for this index. Named, versioned models with a public release history (Chai-1, Chai-2, Chai-3), stated task scope per model, and technical results published in preprints that describe methodology and validation design. Chai-1 was released with open weights for non commercial use. Model versioning and update practice are visible rather than inferred, which is what this axis asks for.

AA on Model Supply Chain DisclosureEvery party is enumerated by name including the model layer. A public subprocessor list naming the model provider, with the retention and training terms that govern data once it arrives, is the canonical artefact.
Vendor Published

The model layer here is more inspectable than anywhere else in this index, because part of it can be downloaded and run. Models are named and versioned with a public release history, task scope is stated per model, technical results appear in preprints describing methodology and validation design, and one model was released with open weights for non commercial use.

Open weights answer this axis in a way no disclosure statement can: a customer does not have to be told what the model is, they can hold it. Versioning and update practice are visible rather than inferred, which matters because a named model with no version history tells a buyer nothing about what changed under them between deployments.

The residual sits on the hosted side rather than the published side and it is the question a discovery organisation should settle before uploading anything. The company licenses custom models trained on a customer's proprietary data and operates a hosted application into which customers enter target and epitope selections.

Nothing located states whether customer inputs or customer training data improve the general models, whether a custom model remains exclusive to the customer that funded it, what retention applies, or what happens to a customer specific model when the agreement ends. The published privacy policy does not reach any of it, since its subject is website visitors. Ask all four in writing.

BB on Clinical and Operational EvidenceNamed deployments with dated outcome figures and enough method to test them, or published research short of independent validation.
Vendor Published

Quantified wet lab validation with stated design: approximately 50 antibody targets, fewer than 20 designs tested per target, with reported hit rates between roughly 16 and 20 percent depending on source, and reported binders for 5 of 5 miniprotein targets. Evidence is reported in preprints rather than peer reviewed journals, and the figure varies between the company's announcement and the preprint, so the record states a range. Deployment at named pharmaceutical customers is commercial validation, not clinical evidence; no candidate from these models has published clinical results.

CC on AI Safety and PHI StewardshipGeneral assurances of privacy and security that do not answer the questions artificial intelligence raises: what is retained, what reaches a model, and what happens to it there.
Vendor Published

The safety half of this axis is answered better than almost anywhere in this category. The stewardship half is not answered at all, and the stewardship question here is real.

On safety, the company publishes a binding acceptable use policy prohibiting a defined set of harmful applications, presupposing safety filters in the product, and reaching derivatives as well as outputs. It gates access to its most capable models and states that it prioritises work with clear health benefit. That is a genuine safety posture rather than a disclaimer.

On stewardship, protected health information is not the material at issue. This is preclinical design software whose inputs are protein sequences, structures and target specifications, so the clinical framing of this axis does not reach it. The domain equivalent is what happens to the sensitive material a customer does entrust to it, and that is where the record is silent. The company licenses custom models trained on a customer's proprietary data, and operates a hosted application into which customers enter target and epitope selections. Nothing located states whether customer inputs or customer training data are used to improve general models, whether a custom model remains exclusive to the customer that funded it, what retention applies, or what happens to a customer specific model when the agreement ends.

Those are the questions a discovery organisation should put in writing before uploading anything, and they are the ones this axis would grade upward if answered. The published privacy policy does not reach them; its subject is website visitors.

Regulatory and Compliance
BB on HIPAA and BAA PostureBusiness associate status is stated and supported by a substantive privacy document, with the agreement or its scope not fully published. For a vendor outside the United States, an equivalent regime documented to this depth grades here.
Vendor Published

A scoping determination and the question closes. The platform designs antibodies and protein binders from target and structural specifications. It does not receive, create, maintain or transmit protected health information on behalf of a covered entity, it sits entirely in preclinical discovery, and it has no treatment, payment or healthcare operations relationship. The health privacy rule does not reach this vendor in its current configuration, so the absence of a business associate agreement is a correct consequence rather than a gap.

One pathway would change that and it should be checked rather than assumed away. The company trains custom models on customer proprietary data. If a customer intended to supply patient derived material, for example sequences or neoepitopes originating from identifiable individuals, the arrangement would need examining on its own terms, and the governing instrument might be an institutional review board approval and a research authorisation rather than a business associate agreement. Establish which framework applies before any patient derived material is transferred, because the two answer to different authorities.

Note also that the published privacy policy covers website visitors. It is not a data processing disclosure for the model service, and it should not be read as one.

CC on Security Certifications and Trust CenterControls are described with an outside check behind them, such as independent penetration testing on a stated cadence, but no attestation against a recognised framework.
Vendor Published

No SOC 2, ISO 27001 or equivalent attestation was located and no trust centre was found. The company is private, so the annual report route that supplies a cybersecurity governance disclosure for the listed vendors in this category does not exist here.

The legal instruments are better constructed than most in this lane and it is worth saying so, because it distinguishes a drafted estate from a template. Website terms, a privacy policy, an acceptable use policy and a separate community licence for distributed model weights are published as distinct documents with defined scope, and the model terms are properly separated from the website terms rather than a single policy being asked to cover both. None of them is a security attestation.

The gap is sharper here than for a vendor that only ships weights. The company operates a hosted application through which pharmaceutical customers submit target and epitope selections, and trains custom models on customer proprietary data. That is a multi tenant service holding the competitive crown jewels of several named large pharmaceutical organisations, with no published attestation, no stated encryption or key management position, no penetration test summary and no subprocessor list. Ask for all of it, and ask specifically how one customer's proprietary training data and resulting model are isolated from another's.

Search caution. At least three unrelated companies operate under the name Chai, including a consumer chatbot business incorporated in Delaware with a Palo Alto address, which publishes detailed data processing, retention and residency terms that are easily mistaken for this vendor's. Confirm the domain resolves to chaidiscovery.com before crediting anything.

BB on FDA and Regulatory StatusThe pathway is stated and in progress, or a clearance is named without the vintage and scope a buyer needs to match it to the product on offer.
Vendor Published

A scoping determination and the question closes cleanly. This is preclinical molecular design software. It does not diagnose, does not inform the treatment of an identified patient, and is not offered as a device. No pathway attaches and none is claimed.

The regulatory position sits entirely downstream and with someone else. Molecules designed on this platform would proceed as biologics through investigational and marketing pathways sponsored by the pharmaceutical customer that develops them. The company is a tool supplier, not a sponsor, and no candidate designed on these models has been reported in clinical study.

One phrase deserves care in a sales conversation. The company describes a long term ambition to produce biologics ready for regulatory submission in a single computational pass. Readiness for submission is determined by a regulator assessing a sponsor's dossier, not by a design tool, and the phrase describes an engineering goal rather than a regulatory status. Do not read it as one, and confirm that any regulatory language used in a pitch attaches to the customer's own programme.

AA on AI Governance and Bias DisclosureA bias or fairness evaluation with a stated method, subgroup performance, or an independent audit of model behaviour.
Vendor Published

The most complete governance posture located in this category, and the only one backed by a costly action rather than a statement.

The company publishes an acceptable use policy that identifies prohibited uses of its models, their outputs and any derivatives. It is contractually binding rather than aspirational: it is incorporated into both the community licence under which model weights are distributed and the web application terms, amendments bind existing users after a notice period, and violation permits suspension or termination of access. It prohibits providing instructions for synthesising or accessing illegal articles, prohibits violating any person's privacy rights, and prohibits attempts to override or circumvent safety filters. That last prohibition is the informative one, because it presupposes that safety filters exist in the product. Most published use policies in this index restrict the user without evidencing any control in the software.

Access is governed by a stated responsible deployment policy under which the newer models are released selectively, prioritising work with clear health and societal benefit and restricting designs the company judges unsafe or ethically concerning. The action behind that statement is what carries weight here. The first generation structure prediction model was released with open weights. The de novo design models that followed were not: there is no public download and no open API, and access is granted on application. A company that gives up the distribution advantage of open release for its most capable model has done something, not merely said something, and this index has found very few vendors in any lane that have.

On technical validity the record is also strong. Reported results carry denominators rather than headline rates, including the proportion of targets that yielded no clean candidate. A known limitation is acknowledged in the company's own materials, namely that modelling the flexible loop regions of antibodies remains a bottleneck relative to more rigid scaffolds. Negative controls are published, including designs that discriminate a single residue mutation in a cancer peptide while avoiding the unmutated counterpart, which reports what the model correctly declines to bind rather than only what it binds.

What is missing, and a buyer should ask for it. There is no model card for any released model. The responsible deployment policy is named but its criteria and decision process are not published, so an applicant cannot know what is restricted or who decides. No position has been published on the 2025 finding that protein design tools could generate variants of proteins of concern evading nucleic acid synthesis screening, which evaluated the tool class rather than this company. And representativeness is entirely unaddressed: nothing describes which target classes, species or structural regimes the models handle poorly.

BB on AI Liability and RecourseA published falsifiable commitment, or a real correction route for the affected person. A published error rate with its method and denominator grades here, and so does a jurisdiction whose law gives the patient an enforceable right to correct an inaccurate record.
Peer Reviewed Publication

Two published commitments carry this and the first is the strongest form of verifiability available anywhere. Releasing a model with open weights means a customer can benchmark it themselves against their own targets rather than relying on the vendor's published figures, so the performance claim is not merely falsifiable in principle but testable by the party who bears the cost of it being wrong.

Preprints describing methodology and validation design sit alongside, so the approach can be read and contested by people with no commercial relationship to the company. The second is a binding acceptable use policy prohibiting a defined set of harmful applications, presupposing safety filters in the product and reaching derivatives as well as outputs, with access to the most capable models gated and a stated priority on work with clear health benefit.

In a domain where the misuse case is designing something dangerous rather than misdiagnosing a patient, a policy that reaches derivatives is the substantive control, and it is a genuine posture rather than a disclaimer. Held below the top grade because nothing attaches to being wrong. No accuracy commitment, service level or remediation applies to a design output, and in preclinical work a wrong prediction costs a synthesis campaign and months rather than a patient. Ask what the hosted models' performance is against the open one, and what the vendor commits to on a failed design.

Integration and Deployment
BB on EHR and Interoperability DepthNamed systems with read access or one directional writing, or standards support with named deployments behind it.
Vendor Published

A scoping determination, and the domain equivalent is disclosed even though it is restrictive.

This is preclinical design software. It does not sit in a clinical workflow, does not read or write a patient record, and no electronic health record integration is claimed or would be meaningful. Grading the absence of an integration the product is not intended to have would misdescribe it.

The equivalent question for a model vendor is how the model reaches a customer's existing research pipeline, and the company's position is clear and public even though it differs by product. The structure prediction model is distributed with downloadable weights under a community licence through a public repository, so it can be embedded directly into a customer's own stack. The de novo design models are not: there is no public API and no download, access runs through the vendor's own web application, and integration is therefore whatever a negotiated agreement provides.

That is a disclosed position rather than a gap, which is why this sits at B. But a buyer planning to build the design models into an automated pipeline should confirm that programmatic access exists at all, since the public position is a portal rather than an interface.

CC on Deployment Model and Data ResidencyA single hosted option with location implied rather than committed.
Vendor Published

The company operates two entirely different deployment positions and a buyer's exposure depends on which product is licensed, so the question should be split.

The first generation structure prediction model was released with downloadable weights under a community licence, which means it can run inside the customer's own environment with nothing transiting to the vendor. That is the strongest position available on this axis and it is disclosed.

The de novo design models are the opposite. There is no public download and no open API, and access is provided through a web application the company operates. For a customer, the input to that application is a target and epitope specification, which in a discovery organisation is among the most commercially sensitive information it holds, and by construction it transits to and is processed in the vendor's environment. Nothing located states where that environment sits, in which regions, whether tenancy is shared or dedicated, how long inputs are retained, or where a custom model trained on a customer's proprietary data is trained and hosted.

Graded on the current products rather than the earlier one. A buyer licensing the design models should press for the deployment topology, retention terms and residency of the hosted service, and should not accept the availability of downloadable weights for a different model as an answer.

Commercial
CC on Commercial TransparencyNo price is published and the posture is discoverable: a buyer can establish how the product is sold and what drives the cost before contacting the vendor. Most of the index sits here.
Vendor Published

No public pricing. Contact the vendor. Commercial terms are licensing agreements with pharmaceutical customers, in at least one case including a custom model trained on customer data; financial terms of announced agreements were not disclosed.

AA on Setting and Specialty CoverageWhere the product is validated to operate is named and supported, settings and specialties both, whether the coverage is broad or deliberately narrow.
Vendor Published

Precisely bounded and honestly stated: preclinical molecular design, principally antibodies and miniproteins, for pharmaceutical and life sciences R&D. The company does not claim clinical or downstream development capability.

Tracked Since Listing

What Changed

Material product, regulatory, evidence and commercial changes at Chai Discovery, each verified against a live source and tagged to the capability axis it bears on. Funding rounds and awards are not product changes and are not logged.

Jul 12, 2026Model / architecturePartially verified

Chai Discovery deployed Chai-3, a next generation antibody design model that the company describes as roughly doubling the success rate of Chai-2 and binding more tightly to harder targets. Date recorded here is the date of the trade coverage confirming deployment; the company has not published a precise release date for Chai-3, and this record does not assert one.

Bears on: Model and Technology TransparencySource
Our read on this change →Tracked since Jul 2026
Comparisons

Compared With

Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Contact the vendor
Model licensing agreements; custom models trained on customer data Vendor Published

Model licensing agreements with pharmaceutical and life sciences organizations, in at least one announced case including early access to a next generation model and a custom model trained on the customer's proprietary data. Financial terms of announced agreements were not disclosed.