EvolutionaryScale
AI research lab building protein language models, founded by researchers who emerged from Meta's Fundamental AI Research unit where they had developed ESM1, the first transformer language model for proteins. Its flagship model ESM3 reasons simultaneously over the sequence, structure and function of proteins in a single all to all architecture, meaning structure and function annotations can be supplied as inputs rather than only produced as outputs, which lets a scientist prompt the model toward a protein with specified properties.
ESM3 was trained on 2.78 billion protein sequences at 98 billion parameters using more compute than any previously disclosed biological model. Its headline demonstration was esmGFP, a novel green fluorescent protein distant enough from natural variants that the company describes it as equivalent to simulating hundreds of millions of years of evolution. ESM3 was published in Science in January 2025.
The model family comes in three sizes reached through the company's Forge API and partner platforms including AWS SageMaker and NVIDIA BioNeMo as an NIM microservice, while ESM3-open, a deliberately smaller model, has weights and source code on GitHub under a non commercial licence. The company states all models are built and deployed under a responsible development framework, and the Science paper records an external review of the risks and benefits of releasing ESM3-open. Raised 142 million dollars in seed funding led by Lux Capital, Nat Friedman and Daniel Gross with NVIDIA and Amazon participating.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
There is no other business. The company is a foundation model lab whose entire output is protein language models, with a direct lineage from ESM1, the first transformer language model for proteins, through ESM2 to ESM3 at 98 billion parameters trained on 2.78 billion sequences. No wet lab franchise, no pipeline, no instrument line sits underneath it.
The all to all architecture is the specific technical claim: because structure and function annotations can be supplied as inputs rather than only emitted as outputs, the model can be prompted toward a specified protein, which is a generative capability rather than a prediction service.
The tool posture limits the autonomy question by design: the model responds to prompts and produces candidate proteins that the user must then express and characterise experimentally, so nothing acts without a scientist deciding to build it. The company's own headline result followed that path, with esmGFP generated and then functionally characterised rather than asserted.
Held at B because no confidence calibration, published failure rates or guidance on where outputs should not be trusted was located, which matters more for a general purpose model than for a narrow one because users span many applications the lab cannot anticipate.
Among the strongest in the index. ESM3 was published in Science in January 2025, so the method has passed external peer review at a top general science venue rather than being described in a technical report. Training scale is stated concretely at 2.78 billion sequences and 98 billion parameters, model sizes are documented, and ESM3-open has both weights and source code on GitHub under a non commercial licence, meaning outside researchers can run a version of the system, probe where it breaks and publish disagreement. That combination of peer review plus inspectable artefact is the standard this index treats as separating a demonstrated capability from an asserted one.
The model layer is documented to an unusual standard and the customer data terms sit with somebody else, which is the finding on this record. On the model, training scale, parameter counts and model sizes are published and an open version is obtainable under a non commercial licence, so the substrate is inspectable rather than merely described, and the clinical form of this axis does not reach the workflow since the models operate on protein sequence, structure and function data with no patient records involved.
The stewardship question that does arise concerns customer proprietary sequence data used in fine tuning, and it is addressed through the cloud platform agreements rather than by a published vendor policy. That distinction deserves naming as a general point.
When a model is distributed through a cloud marketplace, the data terms a customer actually operates under may be the marketplace's rather than the vendor's, which means the party making the commitment is not the party providing the model, and a customer with a question about fine tuning data may find each pointing at the other.
A partner should establish which entity it is contracting with for the data terms, whether the vendor has any obligation under them, and what happens to a fine tuned model at the end of the arrangement. Ask that, plus whether partner sequences inform the base model.
Capability evidence is strong and peer reviewed, therapeutic evidence does not exist. The esmGFP result is a genuine demonstration published in Science: a functional fluorescent protein distant enough from natural variants to constitute novel design rather than interpolation. Distribution is real, with the model family reachable through the company's Forge API and through AWS and NVIDIA platforms.
What is absent is any drug: no therapeutic programme, no clinical asset and no disclosed downstream molecule that reached development. Reach claims describing availability to large numbers of researchers and to most of the largest pharmaceutical companies describe distribution potential through cloud partners rather than measured adoption, and should not be read as evidence of use.
Not applicable in the provider sense and rated accordingly rather than penalized. The models operate on protein sequence, structure and function data with no patient records involved. The stewardship question that does arise is customer proprietary sequence data used in fine tuning, which is addressed through the cloud platform agreements rather than by a published vendor policy.
Not applicable. Users are research organizations across pharmaceutical, biotechnology, materials and academic settings accessing models through an API or cloud marketplace, not covered entities transferring protected health information.
No attestation was located for this company and no trust centre exists. A retrieval caveat is essential here: the company's access platform shares its name with an unrelated enterprise product that does hold ISO 27001 and SOC 2, and searches conflate them. Those certifications belong to the other product. Nothing found in this review attests to this company's own controls.
What does exist is a fuller legal instrument set than most research organisations publish: terms of use, a privacy policy, an acceptable use policy, and separate commercial and non commercial licence agreements governing model access and weight distribution.
One clause in those terms is worth flagging even though it is a governance control rather than a security one, because it is rare. Users must disclose any sequence of concern use of the models and must use them consistently with the purpose they described when applying for access, and the company reserves the right to limit access in response. That is a dual use biosecurity control written into binding terms rather than asserted in a values statement, and for a generative protein design model it is the control that matters most.
The distribution model creates the usual inheritance problem. The models are reachable through major cloud marketplaces and accelerated computing microservice platforms, and those platforms carry their own extensive certifications. Those are the platforms' certifications, covering the infrastructure the models run on, not this company's controls over its own systems, its access approval process or the prompts and sequences customers submit. Ask what the company holds in its own name, and ask where submitted sequences are retained, since a customer's target sequence is the asset it is least willing to disclose.
Nothing to assess rather than something assessed poorly. The models are research tools, not regulated medical devices, and the company maintains no therapeutic pipeline of its own. Regulatory standing for anything designed with these models sits entirely with the organization that develops it.
The strongest governance disclosure located anywhere in this category, and it addresses the specific risk that most protein design companies leave unspoken. The company states that all models are built and deployed under a responsible development framework, and the Science paper records two distinct external inputs: experts who gave feedback on the approach to responsible development, and experts who participated in a review of the risks and benefits of releasing ESM3-open.
The release decision itself reflects that process, since the openly published model is a deliberately smaller and, in the company's description, safer model rather than the frontier system. De novo protein design carries real dual use considerations, and this is the only vendor in this category found to have run a documented pre release risk and benefit assessment with outside experts rather than releasing and leaving the question to others. Held at A on that basis while noting the framework itself was not retrieved in full, and that no analysis of performance variation across protein families or data sparse regions was located.
Peer review at a top general science venue plus an inspectable artefact is the combination this index treats as the standard, and both are present. The core method was published in a leading general science journal rather than described in a technical report, so it passed external review by people with no stake in the company, and training scale is stated concretely at billions of sequences and tens of billions of parameters with model sizes documented.
An open version carries both weights and source under a non commercial licence, which means outside researchers can run a version of the system, find where it breaks and publish the disagreement. That last property is what makes the published claims durable: a paper can be wrong and stand uncorrected for years, while a released model gets stress tested by people looking for failure. Held below the top grade because nothing attaches commercially.
No warranty, indemnity, service level or remediation commitment was located, and the open version is licensed for non commercial use, so a commercial partner cannot independently verify the configuration they would actually deploy and is relying on the published account of a related system. The gap between an open research release and a commercial offering is where an evaluation quietly stops applying. Ask what differs between the open version and the commercial one, and what performance the deployed configuration achieves on the design tasks a partner cares about.
Not applicable in the provider sense, with no EHR touchpoint or clinical workflow surface. The research equivalent is unusually strong: models are reachable through a first party API and through third party enterprise platforms, so they integrate into an existing computational biology stack rather than requiring a separate environment.
The widest set of access routes in this category, and the only one where a user can hold the weights. Four surfaces are disclosed: the company's own Forge API, AWS through SageMaker with further availability stated on additional AWS services, NVIDIA BioNeMo including delivery as an NIM microservice under an NVIDIA AI Enterprise licence, and ESM3-open with weights and source on GitHub under a non commercial licence. Customers can fine tune on their own data within their cloud environment.
Buyers should note the tiering: the frontier models are reachable through API and partner platforms while the model that can be run freely on your own infrastructure is the smaller open one, so the residency answer differs by tier rather than being uniform.
Licensing posture is published even though price is not, which is the useful half for most buyers evaluating a model provider. The non commercial licence attached to ESM3-open is a public term, the tiering between open and frontier models is stated, and API access was opened as a free limited time preview through Forge with an application process.
Funding is disclosed at 142 million dollars in seed capital with investors named including Lux Capital, Nat Friedman, Daniel Gross, NVIDIA and Amazon. No commercial rate card, enterprise pricing or usage terms for the frontier models were located, and pricing through the cloud marketplaces was not established in this review.
Deliberately general rather than therapeutic. Because the model reasons over proteins as a class rather than over a disease area, disclosed applicability spans drug discovery, materials science, carbon capture and biofuels, and the demonstration result was a fluorescent protein used as a laboratory reagent rather than a medicine.
That breadth is genuine and is the point of a foundation model, but buyers should read it correctly: wide applicability is not the same as validated performance in any particular application, and no per domain performance data was located.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Free tier available (non commercial)
|
Tiered model access: open weights under a non commercial licence, plus API and cloud marketplace access to the frontier model family | — | Not published. Fine tuning on customer data is performed within the customer's own cloud environment via partner platforms. | Vendor Published |
Licensing posture is published even though price is not, which is the more useful half for a model provider. ESM3-open carries a stated non commercial licence with weights and source on GitHub, the tiering between the open model and the frontier family is stated plainly, and API access through Forge was opened as a free limited time preview with an application process.
Frontier models are reached through Forge or through partner platforms, including AWS SageMaker and NVIDIA BioNeMo delivered as an NIM microservice under an NVIDIA AI Enterprise licence, so some commercial terms sit with the cloud provider rather than with the vendor. No enterprise rate card, usage pricing or marketplace pricing was located. Funding is disclosed at 142 million dollars in seed capital led by Lux Capital, Nat Friedman and Daniel Gross with NVIDIA and Amazon participating.