Clinical Trials AI
O

Octozi

New York company applying agentic AI to clinical trial data operations, the unglamorous layer beneath drug development where trial data must be cleaned, reconciled, reviewed, and reported before a submission can reach regulators. Combines large language models with deterministic clinical algorithms under an explicitly human in the loop design, integrating with existing clinical systems rather than replacing them, and sells to pharmaceutical sponsors and contract research organizations.

A controlled study of 10 medical reviewers published August 2025 reported roughly sixfold higher data cleaning throughput, reviewer error rates falling from about 54.7 percent to 8.5 percent, and roughly fifteenfold fewer false positive queries. Raised a 3 million dollar seed round in July 2026 led by Surface Ventures, following earlier investment from the venture arm of Swiss pharmaceutical company Debiopharm.

AI Health Index verifiedJuly 28, 2026
Compare Octozi with other vendors
Founded
2024
Headquarters
New York, New York, United States
Website
octozi.com
Categories
clinical-trials-ai
Indexed Products
Octozi Clinical Intelligence Platform
Assessment

Capability Axes

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read

AI Capability
AA on AI CentralityThe artificial intelligence is the product. Remove the model and there is nothing left to sell.
Vendor Published

The agents are the product. The platform automates data cleaning, reconciliation, review, and reporting across trial datasets, and the company's stated positioning is explicitly against the alternative model: most tools in this space put trial data on a dashboard and leave analysis to clinical teams, whereas this performs the work. There is no data management services bureau underneath, which is the distinction separating it from the CRO outsourcing it displaces.

AA on Autonomy and Oversight ModelWhat the system may do and what it may not do are both published, with escalation thresholds, override paths and the conditions that route a case to a person.
Vendor Published

Human in the loop is stated as an architectural commitment rather than a disclaimer, and the founder articulates the division precisely: the human owns the data and stays in control while the model handles tasks that previously took weeks of manual effort. The hybrid design reinforces it, pairing large language models with DETERMINISTIC clinical algorithms, meaning the parts requiring reproducibility are not left to a probabilistic model.

That architectural choice is the same reasoning behind RAAPID and Mendel's neuro-symbolic approaches, and it is the appropriate posture for regulatory submissions where an unexplainable query is worthless.

BB on Model and Technology TransparencyThe approach or the suppliers are named without the version and update discipline behind them.
Vendor Published

The architecture is described specifically, large language models combined with deterministic clinical algorithms, and the reasoning for the split is published rather than asserted. Performance figures are unusually concrete and include the metric most vendors omit: false positive rate. The controlled study reports roughly sixfold throughput improvement, reviewer error rates falling from about 54.7 percent to 8.5 percent, and roughly fifteenfold fewer false positive queries. Publishing the baseline human error rate alongside the assisted rate is the honest framing. Model detail beyond the hybrid description is not published, and the study is small at 10 reviewers.

BB on Model Supply Chain DisclosureSubstantial partial disclosure, or a chain that is structurally short: an in house build, a cleared model that cannot be quietly swapped, or a deployment where the transfer does not occur at all. Naming only the hosting provider sits at the top of this band rather than in A.
Vendor Published

The two commitments that matter for this product type are both made explicitly. The company states it never uses customer private data to train or update its models, and that data is stored in siloed environments isolated from other customers' data, with control over access and insight into operations.

The training exclusion is the significant one and the commercial logic behind it is worth stating, because it explains why this vendor made a commitment others avoid: sponsors will not expose unblinded trial data to a vendor that might learn from it, so in this segment the exclusion is a condition of doing business rather than a courtesy.

That makes it more credible than the same sentence would be elsewhere and also means a buyer should confirm it contractually rather than relying on the incentive. Siloing addresses the companion question, since a platform serving many sponsors in the same therapeutic areas holds material that would be valuable to each of them about the others.

Held below the top grade because none of it is evidenced: it is stated as policy rather than supported by an attestation, audit or published processing agreement, and no model or model family, hosting arrangement or sub processor list was located, so a sponsor knows what the vendor promises not to do and not who else touches the data on the way. Ask for the attestation, the sub processor list, and retention on sponsor supplied trial data.

CC on Clinical and Operational EvidenceNamed customers, or vendor reported percentages with no method, denominator or reference standard. Scale of use is recorded here and is not treated as evidence of benefit.
Vendor Published

A published controlled study exists, dated August 2025 with 10 medical reviewers, reporting substantial throughput and error rate improvements, plus an accompanying economic analysis estimating more than 5 million dollars in savings for a representative Phase III oncology trial. That is more evidence than most seed stage companies offer.

But the sample is 10 reviewers, the economic analysis models a representative trial rather than measuring an actual one, and the study appears vendor conducted rather than independently run. Strategic investment from a pharmaceutical company's venture arm indicates industry validation of the problem. Early stage evidence, honestly presented, not yet independently replicated.

BB on AI Safety and PHI StewardshipCategorical commitments are published, such as no training on customer data, without the retention schedule or the safety engineering behind them.
Vendor Published

More specific than most vendors at this stage, and the commitments address the concerns that matter for a model operating on trial data. The company states it never uses customer private data to train or update its models, stores data in siloed environments isolated from other customer data, and provides control over data access and usage with insight into operations. The training data commitment is the significant one, since sponsors will not expose unblinded trial data to a vendor that might learn from it. Stated as policy rather than evidenced by attestation.

Regulatory and Compliance
BB on HIPAA and BAA PostureBusiness associate status is stated and supported by a substantive privacy document, with the agreement or its scope not fully published. For a vendor outside the United States, an equivalent regime documented to this depth grades here.
Vendor Published

Converted from Not Rated. The prior scoping was correct and there is now more to say than the absence of a business associate agreement.

The scoping first. Customers are pharmaceutical sponsors and contract research organisations, not providers. Trial data arriving from electronic data capture systems, safety databases, central laboratories and patient reported outcome platforms is held under subject identifiers within a study, not as a provider's medical records. The company is therefore not standing between a covered entity and its patients, and contracting runs through sponsor data agreements, transfer agreements and trial data handling frameworks rather than business associate terms. Grading the absence of a business associate agreement as a gap would misdescribe the relationship.

What lifts this above a bare scoping determination is that the company publishes the commitments that matter in the framework that does apply. It states that customer data is never used to train or update its models, which is an explicit exclusion rather than the far more common purpose clause that permits a vendor to improve its services and thereby quietly permits training. It states that data is held in siloed environments isolated from other customers' data, which is the tenancy answer a sponsor needs when a vendor serves competing sponsors. And it describes access controls built on zero trust, least privilege and strong authentication.

Held at B rather than A because none of that is contractual on the public record: no terms are published, no data processing agreement is offered for inspection, and there is no statement of retention or of what happens to a sponsor's data at the end of a study. Get the training exclusion and the isolation commitment into the agreement rather than relying on the website.

CC on Security Certifications and Trust CenterControls are described with an outside check behind them, such as independent penetration testing on a stated cadence, but no attestation against a recognised framework.
Vendor Published

Security architecture is described in concrete terms, zero trust, least privilege, and strong authentication, alongside siloed per customer data environments. That is more architectural detail than most vendors publish. However no SOC 2, ISO 27001, or equivalent third party attestation was located, and described controls are not audited controls. For a company handling pre-submission regulatory data, sponsors will require the attestation, so this is a gap a seed stage vendor will need to close.

CC on FDA and Regulatory StatusNo device claim is made and the product is scoped accordingly. Most administrative and operational products sit here and are not penalised for it, because this axis grades the appropriateness of the positioning rather than possession of a clearance.
Vendor Published

Converted from Not Rated. The prior note identified the right framework and the gap it names is still open after a second search, which is what holds this at C rather than higher.

No device pathway applies and none is claimed. The platform processes trial data for sponsors rather than diagnosing or treating anyone. That much is a clean scoping determination.

The regulatory exposure is real but indirect, and it is the part that matters. Outputs from this platform feed the datasets that support regulatory submissions, so the applicable expectations are data integrity and electronic records requirements rather than device clearance: that records are attributable, legible, contemporaneous, original and accurate, that electronic systems maintain secure computer generated audit trails, that changes do not obscure prior entries, and that the system has been validated for its intended use.

No statement of compliance with any of that was located, and the question is live now rather than prospective, because the company states the platform already supports Phase III trials. The published technical description does say the platform preserves source traceability across the systems it harmonises, which is directionally right and is not the same as a validation or electronic records position.

Four things to establish before a submission depends on it. Whether the system has undergone computer system validation and whether the documentation is available. Whether audit trails are computer generated, secure and independent of the user. How an automated query or correction is attributed, so a reviewer can tell which entries originated with a model. And what the sponsor receives if it needs to reconstruct the platform's reasoning during an inspection.

CC on AI Governance and Bias DisclosureResponsible artificial intelligence is committed to in policy language with no evaluation behind it. Most of the index sits here.
Peer Reviewed Publication

Converted from Not Rated. There is more published here than the prior note found, and the domain question it identified is still the right one and still unanswered.

What is published, and it is not nothing. The company released a technical paper describing a controlled study with ten experienced medical reviewers, naming the base model its own models are fine tuned from, and describing a hybrid design pairing language models with heuristic algorithms developed alongside expert reviewers. Crucially for this axis it reports precision alongside recall rather than a single accuracy figure: on an annotated evaluation set the discrepancy detector is reported at 97.5 percent recall and 77.2 percent precision, and in the reviewer study precision rose from 44.6 to 93.2 percent with assistance. Reporting the false positive side of the ledger, and quantifying over flagging, is the behaviour this index credits and most vendors avoid.

Two things keep it at C. The first is the limitation in how it was evaluated. The discrepancy detection figures come from an annotated synthetic dataset, and a discrepancy detector tested against synthetic discrepancies is being measured against its authors' own model of what goes wrong in trial data. That is a legitimate development step and it is not evidence of performance against the failures nobody anticipated, which are the ones that matter at database lock.

The second is the question the prior note framed correctly and which nothing addresses. The domain relevant risk here is not demographic bias but systematic query bias: whether flags concentrate disproportionately on particular sites, forms, therapeutic areas, source systems or investigator populations. An aggregate precision figure cannot answer that, because a detector can be well calibrated overall while over querying a subset of sites, and which discrepancies get investigated before lock shapes the dataset a regulator eventually sees. Ask for the flag rate broken down by site and by source system, and for whether that distribution is monitored over the life of a study.

BB on AI Liability and RecourseA published falsifiable commitment, or a real correction route for the affected person. A published error rate with its method and denominator grades here, and so does a jurisdiction whose law gives the patient an enforceable right to correct an inaccurate record.
Vendor Published

The controlled study behind this record does the one thing this index almost never sees: it publishes the human baseline alongside the assisted result. Reviewer error rates are reported falling from roughly fifty five per cent to under nine per cent, with a sixfold throughput improvement and roughly fifteenfold fewer false positive queries.

Naming the unassisted error rate matters more than the improvement ratio, because every vendor claiming to reduce errors is comparing against a baseline, and almost all of them leave it unstated so a reader cannot tell whether the gain is large or the starting point was poor. Here both numbers are visible, and the starting point is unflattering to the incumbent process rather than to the vendor, which is precisely the figure a vendor has no incentive to publish and every reason to.

Publishing the false positive reduction is the second unusual choice, since false positives are the cost side of any detection aid and the metric most often omitted. The architecture is also described specifically as language models combined with deterministic clinical algorithms, with the reasoning for the split published rather than asserted, which tells a reviewer which outputs are reproducible and which are generative.

Held below the top grade because the study is small at ten reviewers, model detail beyond the hybrid description is unpublished, and no warranty, indemnity or remediation commitment attaches. Ask for a larger replication and per query type error rates.

Integration and Deployment
BB on EHR and Interoperability DepthNamed systems with read access or one directional writing, or standards support with named deployments behind it.
Vendor Published

Integration with existing clinical systems is the stated design premise rather than a feature, with the company positioning the platform as an operational layer that works with the systems clinical teams already use rather than requiring migration. Seed funding is explicitly earmarked to deepen integrations with clinical systems used by pharmaceutical, biotech, and medtech companies. This is the clinical trial technology stack, electronic data capture and related systems, not the EHR. No named connector list was located, which is the natural gap at seed stage.

CC on Deployment Model and Data ResidencyA single hosted option with location implied rather than committed.
Vendor Published

Siloed per customer environments isolated from other customer data are stated, which is a meaningful architectural commitment for sponsors handling competitively sensitive trial data. Specific hosting, tenancy, and geographic residency terms are not published, and residency matters for multinational trials subject to differing data protection regimes.

Commercial
CC on Commercial TransparencyNo price is published and the posture is discoverable: a buyer can establish how the product is sold and what drives the cost before contacting the vendor. Most of the index sits here.
Vendor Published

No pricing is published, but the value case is quantified against a costable baseline: an economic analysis of a representative Phase III oncology trial estimated savings exceeding 5 million dollars per study, driven primarily by earlier database lock and compressed timelines. Sponsors can model that against their own trial economics, since delayed database lock has a known cost. The figure is vendor generated and models a representative rather than actual trial. Buyers should also weigh vendor durability, since this is a seed stage company with 3 million dollars raised operating in a market served by established CROs.

CC on Setting and Specialty CoverageCoverage is claimed broadly without specifics, or stated clearly with nothing validating it yet.
Vendor Published

Focused on the clinical trial data operations layer, spanning data cleaning, review, reconciliation, and reporting, sold to pharmaceutical sponsors, biotech, medtech, and contract research organizations. Deliberately narrow within clinical development, addressing the data management function rather than trial design, patient matching, or simulation, which distinguishes it from Unlearn, Mendel, QuantHealth, and the matching vendors indexed here. Published evidence centres on Phase III oncology, though the capability is not disease specific.

Comparisons

Compared With

Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Contact the vendor
Undisclosed. Sold to pharmaceutical sponsors, biotech, medtech, and contract research organizations; no rates published. Not disclosed. Customers are pharmaceutical sponsors and CROs rather than providers, so contracting runs through sponsor data agreements rather than HIPAA business associate terms. Not disclosed. Designed to integrate with existing clinical systems rather than requiring migration, with seed funding earmarked to deepen those integrations. Vendor Published

No pricing is published, but the value case is quantified against a baseline sponsors can cost themselves: an economic analysis of a representative Phase III oncology trial estimated savings exceeding 5 million dollars per study, driven primarily by earlier database lock and compressed timelines. Delayed database lock has a known cost to any sponsor, so that is a modellable comparison rather than an abstract efficiency claim. Two caveats.

The figure is vendor generated and models a representative rather than an actual trial, and the supporting controlled study covered 10 medical reviewers. Vendor durability is the other consideration: this is a seed stage company that raised 3 million dollars in July 2026, competing in a function currently served by established contract research organizations, and sponsors will likely require a SOC 2 or equivalent attestation that was not located.