Owkin
Tech bio company combining biological large language models, multimodal patient data, and agentic software. The Owkin K co pilot has two environments: K Navigator, an agentic research environment free to academic researchers that accelerates literature review across 26.5 million articles and 19 biomedical databases and explores spatial multiomic patient data, and K Pro, an enterprise co pilot that uses a single orchestrator to select and combine specialized biological AI skills across drug discovery and development. Both are powered by Owkin Zero, a fine tuned biological reasoning model.
The company reports K Pro accelerating internal drug target identification from more than 12 months to roughly 3 months, validated through collaborations with AstraZeneca, Bristol Myers Squibb, and Sanofi, including a three year AstraZeneca licensing agreement to build biopharma agents. Owkin operates a group of entities spanning a biology foundation model (Bioptimus), diagnostics (Waiv), and a clinical stage drug program (Epkin); this record covers the software platform. Its CE-IVD marked pathology diagnostics, RlapsRisk BC and MSIntuit CRC, are indexed separately as Owkin Dx. Founded 2016 by Thomas Clozel, MD and Gilles Wainrib, PhD.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The models are the product. Owkin Zero is a fine tuned biological reasoning model, and the K co pilot orchestrates specialized biological AI skills over multimodal patient data. Nothing in the offering functions without them.
The autonomy claim is explicit and the oversight design is not described.
Both products are presented as agentic. The enterprise co pilot is described as using a single orchestrator that dynamically selects and combines specialised biological skills across discovery and development, and the researcher environment is described as an agentic research environment. Nothing located states what the orchestrator decides without asking, whether a researcher approves or can inspect the chain of steps it selected, whether outputs are traceable to the underlying article or dataset that produced them, or what happens when a step is wrong.
Two things sit on the credit side and neither closes the question. The company's ethics page names the return of a person to the loop during validation as a concern it holds, which is more than most competitors say. And the federated training infrastructure is genuinely auditable, since a distributed ledger records and attributes each computation. That auditability belongs to the training system, not to the co pilot a researcher actually queries.
The mitigating context is real: the audience is researchers, the output feeds hypothesis generation, and laboratory work is itself a check on a wrong answer. But that is the index inferring a safeguard from the use case rather than the vendor describing one. Asserting agentic capability across the whole of discovery and development while describing no oversight mechanism invites a buyer to assume a structure that has not been shown to exist.
Specific and architecturally described rather than gestured at: a named reasoning model (Owkin Zero), a stated orchestration approach in which a single intelligent orchestrator dynamically selects and combines specialized skills, and disclosed data foundations spanning 19 biomedical databases, 26.5 million articles, and spatial transcriptomics from the MOSAIC initiative. The related entity structure, including the Bioptimus foundation model, is also disclosed rather than obscured.
The model layer, the corpus and the corporate structure are all named, which is the full set this axis asks for and which almost nothing else in the index provides together. The reasoning model is named. The orchestration approach is described, with a single orchestrator dynamically selecting and combining specialised skills, so a buyer knows the shape of what runs rather than only that something does.
The data foundations are enumerated by source, spanning nineteen biomedical databases, more than twenty six million articles and spatial transcriptomics from a named research initiative, which answers the corpus half properly rather than with an adjective.
And the related entity structure, including the affiliated foundation model company, is disclosed rather than obscured, which is the parent and affiliate question this index has to ask of every subsidiary and which is almost never answered without prompting. One residual should be obtained rather than assumed, and it is the enterprise half rather than the research half.
Nothing located states what happens to material an enterprise customer puts into the co pilot, whether those inputs contribute to training the reasoning model, or whether a model tuned for one customer stays exclusive to them. Ask for the enterprise input terms specifically, since the published instrument governs the research network rather than the product.
Commercial validation is strong and named: collaborations with AstraZeneca, Bristol Myers Squibb, and Sanofi, including a three year AstraZeneca licensing agreement specifically to build biopharma agents. Multi year licensing by a top tier pharma is meaningful third party validation of the platform.
Held back from A because the headline performance claim, target identification compressed from more than 12 months to roughly 3, is vendor reported without published methodology or a stated comparison baseline.
One of the stronger stewardship records in this category, and it rests on an instrument most vendors here do not publish at all.
The company maintains a dedicated patient information page addressed to the people whose data it processes rather than to its customers. It states which data protection regimes it holds itself to, naming the General Data Protection Regulation, its local implementations across the European Economic Area, the United Kingdom Data Protection Act 2018, the United Kingdom General Data Protection Regulation and the Swiss data protection act. It states that records are preselected against study criteria by each network member, and that pseudonymisation, de identification, or full anonymisation where possible is performed by the network member before anything is shared with the company or its partners. De identification at the source institution, rather than by the vendor on receipt, is a materially stronger control than the reverse.
The purpose limitation is genuine rather than decorative, which is worth stating because this index frequently finds the opposite. Reuse is confined to secondary research supporting public health, the quality and safety of healthcare, medicinal products or devices, and the sharing of scientific knowledge through publication. That is a narrowing clause, not the common formulation that permits a vendor to improve its services and thereby quietly permits model training.
Held at B on two grounds. The instrument governs the research network. It does not address what happens to material an enterprise customer puts into the co pilot, or whether customer inputs contribute to training the reasoning model, which is the question an enterprise buyer needs answered and which nothing located addresses. And the qualifier that anonymisation happens where possible is an honest concession that pseudonymised patient derived data does sometimes reach the company, without a description of the controls that then apply.
A genuine gap rather than a scoping determination, and the way it was established is worth recording.
The company publishes a patient facing data protection statement, which is exactly where a position on this regime would sit if one existed. That statement enumerates the regimes the company holds itself to and they are entirely European: the General Data Protection Regulation, its local implementations across the European Economic Area, the United Kingdom Data Protection Act 2018, the United Kingdom General Data Protection Regulation and the Swiss data protection act. The United States health privacy rule is not among them. No business associate agreement, no availability statement and no covered entity or business associate characterisation was located.
An enumerated list of regimes that omits yours is stronger evidence than silence. Silence usually means the question was never put. An enumeration means the vendor considered which regimes govern its processing and reached a conclusion that did not include yours.
This matters here because the scoping argument that excuses most vendors in this category does not apply. A company working only with chemistry and compound structures never touches patient data. This one does: it processes patient derived records supplied by member institutions, and it operates a United States entity. A United States institution contributing data, or a United States customer, should establish before contracting whether the company will execute a business associate agreement at all, which entity would sign it, and whether the arrangement is instead being structured as a research use under institutional review board oversight and an authorisation or waiver, which is a different legal instrument answering to a different authority.
A real certification, held and announced by the company in its own name, which is the strong pole of this axis and uncommon in this category.
The company states that it holds ISO 27001:2022 for information security management. That is a current version of a certifiable standard covering a management system rather than a marketing claim about practices, and it is asserted directly by the company rather than inherited from a cloud provider or borrowed from a corporate sibling.
Two things keep this at B rather than A. There is no SOC 2 report and no trust centre through which a counterparty can request the certificate, its statement of applicability, a penetration test summary or a subprocessor list, so a buyer must open a conversation to obtain anything. More importantly the certificate's scope is not published, and scope is the whole question for a company with this structure. This group spans the software platform indexed here, a separately indexed diagnostics business, a foundation model business and a clinical stage drug programme. A certification held at group level does not establish that every product estate sits inside the audited boundary, and this index has repeatedly found assurance published at a parent brand that on inspection covered one constituent business.
Ask for the certificate and read its scope statement. Confirm in writing that the product being licensed, and the infrastructure serving it, are named within the information security management system boundary.
A scoping determination rather than an absence finding, and the question closes cleanly. The products indexed on this record are research tools: an agentic research environment for academic researchers and an enterprise co pilot for biopharma discovery and development teams. They inform research decisions, not the diagnosis or treatment of an identified patient, so no device pathway attaches and none is claimed.
The important point on this axis is what does not transfer. The corporate group does hold genuine regulatory marks, and a reader who searches the company name will find them: the CE-IVD marked pathology diagnostics are real, and they carry the quality management certification that a diagnostics business requires. Those belong to the separately indexed diagnostics record. A regulatory mark authorises a specific product with specific labelling and intended use. It says nothing about a research co pilot sharing the parent brand, and the group also spans a foundation model business and a clinical stage drug programme, each with its own regulatory position.
Ask which legal entity contracts, and confirm that any regulatory claim made during a sales conversation names the product being purchased.
Among the better records in this category, earned through scientific and engineering practice rather than through a governance document.
The substantive evidence is threefold. Core methods are published in peer reviewed venues including a leading clinical journal and a major machine learning conference. The privacy infrastructure underpinning the federated approach has been open sourced and donated to a neutral open source foundation, which means the company no longer controls the code by which its own privacy claims can be checked. And the federated approach is itself a structural answer to the representativeness question this axis asks: training across multiple institutions without pooling data directly addresses the single institution skew that distorts models built on one site's population.
The company also maintains a dedicated ethics page, and it is worth being precise about what that page does. It names the right questions, including how inclusion decisions during model building are made, how bias is minimised, and how validation should account for heterogeneity and return a person to the loop. Naming those questions is better than the silence common in this category. But they are posed as questions rather than answered, and nothing on that page commits the company to a practice or reports a result.
Held at B rather than A for three reasons. There is no model card for the reasoning model, no statement of intended and unsuitable use, and no subgroup or heterogeneity performance published for the co pilot products. The strongest evidence also relates to the earlier federated pathology research rather than to the products indexed here, so check that the evidence a seller cites is drawn from the business being purchased.
The instrument that earns this is addressed to the people whose data is processed rather than to the customers who pay, which is rare enough to be the headline. A dedicated patient information page states which data protection regimes the company holds itself to, naming European, United Kingdom and Swiss instruments, so the affected person can identify the rights they hold and the authority behind them. Two features make it substantive rather than decorative.
De identification happens at the source institution before anything is shared with the company or its partners, with records preselected against study criteria by each network member, and de identification performed by the holder rather than by the recipient is a materially stronger control than the reverse.
And the purpose limitation genuinely narrows: reuse is confined to secondary research supporting public health, the quality and safety of healthcare, medicinal products or devices, and the sharing of scientific knowledge through publication. That is a narrowing clause rather than the common formulation permitting a vendor to improve its services, which quietly permits model training. Held below the top grade on scope and candour.
The instrument governs the research network and does not reach what an enterprise customer puts into the co pilot. And the qualifier that anonymisation happens where possible is an honest concession that pseudonymised patient derived data does sometimes reach the company, with no description of the controls that then apply. Ask for those controls, and for the enterprise terms.
A scoping determination, and the domain equivalent is answered.
These are research and discovery products. They do not sit in a clinical workflow, do not write to a patient's chart, and no electronic health record integration is claimed, so the clinical interoperability question this axis normally asks does not reach them. Grading the absence of an integration that the product is not intended to have would misdescribe the vendor.
The meaningful equivalent for a research platform is whether it can reach institutional data where that data lives, and here the answer is substantive. The federated software is deployed inside participating medical centres' own infrastructure, which is the deepest form of institutional integration available in this category, and it was built specifically to compute across sites that cannot pool data. On the research corpus side the co pilot is documented as spanning nineteen biomedical databases, a literature corpus of roughly 26.5 million articles, and spatial transcriptomic data from a named multi institution initiative.
What is not documented is the integration surface a buyer would need to operationalise it: whether an interface exists to connect a customer's own internal data, what formats are accepted, and whether results can be pushed into a customer's existing research systems rather than read inside the vendor's environment.
Substantively answered for the research network, and unusually well, though with a real gap on the co pilot itself.
The company's federated architecture is the answer and it is documented rather than asserted. Training runs on datasets that remain inside each medical centre's own infrastructure, with only models and non sensitive metadata moving between the company and its partners across a secure network. Computation is orchestrated through a distributed ledger so that each operation is traced and attributable. This was demonstrated in peer reviewed work training models across four hospitals without the underlying pathology data leaving hospital firewalls, and the underlying software has been open sourced and placed under a neutral foundation.
That is a stronger position than a residency commitment, because it removes the transfer rather than governing it. Buyers should note precisely what it covers. It describes the federated research infrastructure. It does not describe where the enterprise co pilot itself runs, in which regions its tenancy sits, where a customer's own uploaded material is held, or which inference infrastructure serves the reasoning model. No residency regions are stated for the co pilot products. Ask for the deployment topology of the specific product being licensed rather than accepting the federated architecture as an answer to both questions.
Structure is disclosed even though amounts are not, and the two sided model is itself informative: K Navigator is free to academic researchers in beta with a waitlist and a stated six month evaluation period, while K Pro is enterprise licensed through demo, exemplified by a three year AstraZeneca agreement. A researcher can determine access terms without a sales call; a biopharma buyer cannot determine cost.
Clearly stated across two distinct audiences: academic researchers via K Navigator and biopharma decision makers via K Pro, spanning discovery through development. Held back from A because the company also spans a foundation model business, a diagnostics business, and a clinical stage drug program, and the boundaries between the software platform indexed here and those adjacent entities are not always drawn crisply in its own materials.
What Changed
Material product, regulatory, evidence and commercial changes at Owkin, each verified against a live source and tagged to the capability axis it bears on. Funding rounds and awards are not product changes and are not logged.
Owkin received European regulatory approval for two first-in-class AI diagnostic solutions designed for breast cancer and colorectal cancer. The tools leverage multimodal patient data to assist in biomarker screening and outcome prediction.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Free for academic researchers (K Navigator); contact the vendor for K Pro
|
Free academic tier; enterprise licensing for biopharma | — | — | Vendor Published |
Two tiers with different commercial models, both disclosed in structure though not in amount. K Navigator is free to academic researchers, currently in beta with a waitlist and a stated six month evaluation period. K Pro is an enterprise platform sold through demo and licensing, exemplified by a three year AstraZeneca agreement. No rate card published for the enterprise tier.