Profluent
Protein design company built on the ProGen family of protein language models, originating in a Salesforce research project that first demonstrated large language models could generate functional proteins, published in Nature Biotechnology in 2023. ProGen3 is a family of billion parameter generative models trained on more than 3.4 billion protein sequences using a sparse architecture the company reports delivers a fourfold speedup, and it generates full length novel proteins or redesigns specific domains of an existing protein.
The company's defining result is OpenCRISPR-1, described as the first CRISPR gene editor designed from scratch by AI and published in Nature in 2025: from generated candidates, 48 sequences were functionally characterised in human cells, and the lead showed comparable on target editing to SpCas9 at 55.7 percent against 48.3 percent while cutting off target editing by roughly 95 percent, at 0.32 percent against 6.1 percent, sitting 403 mutations from SpCas9 and 182 from any natural protein in the CRISPR-Cas Atlas.
OpenCRISPR-1 was released for free licensing, which also sidesteps the licence payments attached to existing CRISPR patent families, and the company reports tens of thousands of downloads within a day and use across academic, pharmaceutical and commercial operations. Further work includes E1, described as the first retrieval augmented model for protein engineering, Protein2PAM for programming the DNA motifs an editor recognises, and OpenAntibodies covering single shot antibody design against 20 drug targets.
Commercial paths are asset licensing, collaboration, or early access to the API, with named relationships including Revvity, Corteva Agriscience, the Rett Syndrome Research Trust and Integrated DNA Technologies. Total funding is 150 million dollars including a 106 million dollar Series B co led by Altimeter Capital and Bezos Expeditions.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The models are the company and the proof is unusually direct. ProGen3 is a family of billion parameter generative language models trained on more than 3.4 billion protein sequences, and the company's founding scientific claim, published in Nature Biotechnology in 2023, was that large language models can generate functional proteins at all.
OpenCRISPR-1 closes the argument: a working genome editor whose sequence sits 182 mutations from any natural protein in the CRISPR-Cas Atlas could not have been arrived at by screening nature, so the model is not accelerating a search, it is producing something that was not there to find.
The disclosed workflow puts experimental characterisation directly after generation rather than treating model output as a result. For OpenCRISPR the company generated candidates, selected 48 sequences and characterised them functionally in human cells, then reported both on target and off target activity against the incumbent enzyme, which is the correct shape for validating a designed molecule.
Held at B because no confidence thresholds, generation to hit ratios across programmes, or published guidance on where the models are unreliable were located, so the oversight is demonstrated in one flagship case rather than specified as a standing practice.
Peer reviewed twice at the highest level and specific about architecture. The foundational demonstration that language models generate functional proteins appeared in Nature Biotechnology in 2023, and the AI designed CRISPR system appeared in Nature in 2025, so both the method and its most consequential output have been externally reviewed.
ProGen3 is described concretely, as billion parameter models over 3.4 billion sequences using a sparse architecture reported to give a fourfold speedup without performance loss, accompanied by a preprint the company describes as the first wet lab evidence of scaling benefits in biological design. Subsequent work is named and characterised rather than gestured at, including E1 as a retrieval augmented model for protein engineering and Protein2PAM for programming editor recognition motifs.
The model line is named and published, which answers more of this axis than most of the lane, and the chain question here is a different kind from the one this axis usually asks. Individual models are named and characterised rather than presented as an undifferentiated platform, with training corpus scale and architecture stated, so a reader can identify what produced a given output.
The clinical form of this axis does not reach the workflow, which operates on protein and nucleic acid sequence data with no patient records involved. What replaces it is a question about the outputs rather than the inputs. This company designs functional proteins and gene editing systems, so the material that leaves the platform is itself capable of biological effect, and the relevant chain question is who can obtain a design, under what terms, and what constrains onward transfer.
Nothing published describes access controls over designed sequences, screening applied before release, or what a partner may do with a design afterwards. That is a different axis of concern from patient data and it is the one this vendor's outputs actually raise. On the ordinary enumeration there is nothing: no hosting arrangement, no sub processor list, and no position on whether partner supplied sequences improve models used for others. Ask what governs access to and onward transfer of designs, and the partner data position.
Molecular evidence is excellent, clinical evidence is absent, and buyers should hold those apart. OpenCRISPR-1 was benchmarked head to head against SpCas9, the standard tool it would replace, with on target editing of 55.7 percent against 48.3 percent and off target editing of 0.32 percent against 6.1 percent, published in Nature.
Comparing a designed molecule directly against the incumbent on the axis where the incumbent is weakest, and publishing the numbers, is a stronger form of evidence than most in this category offer. Operational adoption is named rather than anonymous, with commercial relationships including Revvity, Corteva Agriscience, the Rett Syndrome Research Trust and Integrated DNA Technologies.
What does not exist is a therapeutic in humans; the company states bringing an AI designed therapeutic to patients as a goal, which places it behind category peers with clinical assets. Download and licence request figures are company reported.
Not applicable in the provider sense and rated accordingly rather than penalized. The models operate on protein and nucleic acid sequence data with no patient records in the workflow. The distinct safety question raised by this vendor's output is biological rather than informational and is treated on the governance axis.
Not applicable. Counterparties are pharmaceutical, agricultural, diagnostic and biomanufacturing organizations licensing molecules or model access, not covered entities transferring protected health information.
No attestation was located and no trust centre exists. The company does publish terms of use and a privacy policy, but their scope is narrower than their titles suggest and a reviewer should read the definitions before crediting them.
Both documents define their subject as the website, its subdomains, and the outputs of one named computational server. That covers a visitor and a user of that single tool. It does not reach the licensing and partnership business, where a counterparty's target sequences and design constraints would be the sensitive material, and no separate terms governing that relationship were located. A published privacy policy is not evidence of a position on customer research data unless its scope says so.
One clause in the privacy policy is worth raising in diligence: the company reserves the right to dispose of any data at its discretion without notice. That is unusual phrasing, it runs the opposite way from the retention commitments buyers normally seek, and for a user who has submitted work to a hosted tool the practical question is whether their submissions and outputs are covered by it.
One credit belongs on the record even though it sits closer to governance than to security. The openly released gene editor is distributed under a licence carrying an explicit ethical use obligation, which is a deliberate dual use control on a capability that plainly warrants one. Note the direction of that control: it places an obligation on the licensee. It is not a statement about what the company does to secure its own systems, and the generative models themselves are held closely rather than released, so the assets a buyer would ask about are internal and undescribed.
Nothing to assess rather than something assessed poorly. The models and the designed molecules are research and licensing assets rather than regulated products, and the company maintains no disclosed clinical pipeline of its own. Regulatory standing for any therapeutic built on OpenCRISPR or another licensed asset sits with the licensee that develops it, which is a structural consequence of the open licensing model.
The disclosure gap here is specific and consequential, and it should be stated with the mitigating facts attached. Profluent designed a functional human genome editor with a generative model and released it openly for free licensing, and no accompanying dual use risk assessment, release governance framework or access policy was located in this review.
That is the same question EvolutionaryScale answered in this category by commissioning an external review of the risks and benefits before releasing an open model, which makes the contrast a direct one rather than a general complaint.
Two facts cut the other way and belong in any fair reading: OpenCRISPR-1 reduced off target editing by roughly 95 percent against SpCas9, so on the axis most relevant to gene editing safety the designed molecule is an improvement on the incumbent, and free licensing has a genuine access argument since it bypasses the licence payments attached to existing CRISPR patent families.
The separate domain bias question, whether generation quality falls away in protein families with sparse training coverage, is partially addressed by the company's scaling law work but no per family breakdown was located.
Peer reviewed twice at the highest level, and the second review is the one that matters most. The foundational demonstration that language models can generate functional proteins appeared in a leading biotechnology journal, and the designed gene editing system appeared in a leading general science journal, so both the method and its most consequential output have been externally reviewed.
Reviewing the output rather than only the method is rare and important here: a method paper establishes that an approach can work, while a paper on the resulting molecule establishes that this particular thing does what was claimed, and only the second protects a partner relying on a specific design.
The architecture is described concretely rather than gestured at, with model sizes over a stated corpus using a sparse design reported to give a fourfold speedup without performance loss, and a preprint the company describes as the first wet laboratory evidence of scaling benefits in biological design. Subsequent work is named and characterised individually rather than presented as a platform.
Held below the top grade because nothing attaches commercially: no warranty, indemnity, service level or remediation commitment was located, and a peer reviewed design says nothing about the success rate of designs made for a partner, which is the number a partner is buying. Ask for the prospective success rate on partner programmes with a denominator, and what happens when a designed protein fails in the laboratory.
Not applicable. This is a molecular design platform with no provider workflow surface and no EHR touchpoint.
Three access routes are stated openly, which is more than most in this category: licensing a designed molecule asset outright, collaborating on new proteins, or joining an early access programme for the models themselves via API. OpenCRISPR-1 is distributed as an openly licensed sequence, which is the least restrictive delivery model in the category since the recipient simply holds the molecule. Held at B rather than A because model access is an early access API rather than a generally available product, and no on premise option, tenancy terms or data residency commitments were located.
The commercial structure is unusually legible even though no price is published. Three routes are named plainly, being asset licensing, collaboration and early access to models, and the free licensing of OpenCRISPR-1 is itself a published commercial term with a stated rationale.
Commercial counterparties are named across sectors rather than described generically, including Revvity, Corteva Agriscience, the Rett Syndrome Research Trust and Integrated DNA Technologies, which lets a buyer see what kind of organization actually transacts here. Funding is disclosed at 150 million dollars total including a 106 million dollar Series B co led by Altimeter Capital and Bezos Expeditions. No rate card, licence fee schedule or deal economics were located.
Broad by design and demonstrated across genuinely different protein classes rather than asserted. Disclosed target markets span therapeutics, diagnostics, agriculture and biomanufacturing, and the underlying model family has produced or been applied to enzymes, antibodies, gene editors and peptides, with OpenAntibodies covering single shot design against 20 drug targets and OpenCRISPR covering genome editing.
The named commercial relationships mirror that spread, from a life sciences tools company to an agricultural science company to a rare disease research trust. Modality is proteins rather than small molecules, which is the real boundary.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Free for OpenCRISPR-1; other terms not published
|
Three routes: molecule asset licensing, discovery collaboration, and early access API. OpenCRISPR-1 is licensed free of charge. | — | Not published. | Vendor Published |
The commercial structure is stated openly in three named routes, which is more legible than most of this category: license a designed molecule asset outright, collaborate on new proteins, or join the early access programme for the models via API.
The most notable term is that OpenCRISPR-1 is licensed free of charge, which the company frames as opening access and which in practice also sidesteps the licence payments attached to existing CRISPR patent families, so the cost comparison for a buyer is against paying to use conventional Cas9 rather than against another AI vendor. Commercial counterparties are named across sectors including Revvity, Corteva Agriscience, the Rett Syndrome Research Trust and Integrated DNA Technologies.
No rate card, licence fee schedule, milestone structure or royalty terms were located. Total funding is 150 million dollars including a 106 million dollar Series B co led by Altimeter Capital and Bezos Expeditions.