Clinical Trials AI
O

Octozi

New York company applying agentic AI to clinical trial data operations, the unglamorous layer beneath drug development where trial data must be cleaned, reconciled, reviewed, and reported before a submission can reach regulators. Combines large language models with deterministic clinical algorithms under an explicitly human in the loop design, integrating with existing clinical systems rather than replacing them, and sells to pharmaceutical sponsors and contract research organizations. A controlled study of 10 medical reviewers published August 2025 reported roughly sixfold higher data cleaning throughput, reviewer error rates falling from about 54.7 percent to 8.5 percent, and roughly fifteenfold fewer false positive queries. Raised a 3 million dollar seed round in July 2026 led by Surface Ventures, following earlier investment from the venture arm of Swiss pharmaceutical company Debiopharm.

Last VerifiedJuly 21, 2026
Compare Octozi with other vendors
Founded
2024
Headquarters
New York, New York, United States
Website
octozi.com
Categories
clinical-trials-ai
Indexed Products
Octozi Clinical Intelligence Platform
Assessment

Capability Axes

AI Capability
AI Centrality
A
Vendor Published

The agents are the product. The platform automates data cleaning, reconciliation, review, and reporting across trial datasets, and the company's stated positioning is explicitly against the alternative model: most tools in this space put trial data on a dashboard and leave analysis to clinical teams, whereas this performs the work. There is no data management services bureau underneath, which is the distinction separating it from the CRO outsourcing it displaces.

Autonomy and Oversight Model
A
Vendor Published

Human in the loop is stated as an architectural commitment rather than a disclaimer, and the founder articulates the division precisely: the human owns the data and stays in control while the model handles tasks that previously took weeks of manual effort. The hybrid design reinforces it, pairing large language models with DETERMINISTIC clinical algorithms, meaning the parts requiring reproducibility are not left to a probabilistic model. That architectural choice is the same reasoning behind RAAPID and Mendel's neuro-symbolic approaches, and it is the appropriate posture for regulatory submissions where an unexplainable query is worthless.

Model and Technology Transparency
B
Vendor Published

The architecture is described specifically, large language models combined with deterministic clinical algorithms, and the reasoning for the split is published rather than asserted. Performance figures are unusually concrete and include the metric most vendors omit: false positive rate. The controlled study reports roughly sixfold throughput improvement, reviewer error rates falling from about 54.7 percent to 8.5 percent, and roughly fifteenfold fewer false positive queries. Publishing the baseline human error rate alongside the assisted rate is the honest framing. Model detail beyond the hybrid description is not published, and the study is small at 10 reviewers.

Clinical and Operational Evidence
C
Vendor Published

A published controlled study exists, dated August 2025 with 10 medical reviewers, reporting substantial throughput and error rate improvements, plus an accompanying economic analysis estimating more than 5 million dollars in savings for a representative Phase III oncology trial. That is more evidence than most seed stage companies offer. But the sample is 10 reviewers, the economic analysis models a representative trial rather than measuring an actual one, and the study appears vendor conducted rather than independently run. Strategic investment from a pharmaceutical company's venture arm indicates industry validation of the problem. Early stage evidence, honestly presented, not yet independently replicated.

AI Safety and PHI Stewardship
B
Vendor Published

More specific than most vendors at this stage, and the commitments address the concerns that matter for a model operating on trial data. The company states it never uses customer private data to train or update its models, stores data in siloed environments isolated from other customer data, and provides control over data access and usage with insight into operations. The training data commitment is the significant one, since sponsors will not expose unblinded trial data to a vendor that might learn from it. Stated as policy rather than evidenced by attestation.

Regulatory and Compliance
HIPAA and BAA Posture
Not rated

No explicit HIPAA or BAA commitment was located. Customers are pharmaceutical sponsors and contract research organizations rather than providers, so contracting typically runs through sponsor data agreements and clinical trial data handling frameworks rather than HIPAA business associate terms.

Security Certifications and Trust Center
C
Vendor Published

Security architecture is described in concrete terms, zero trust, least privilege, and strong authentication, alongside siloed per customer data environments. That is more architectural detail than most vendors publish. However no SOC 2, ISO 27001, or equivalent third party attestation was located, and described controls are not audited controls. For a company handling pre-submission regulatory data, sponsors will require the attestation, so this is a gap a seed stage vendor will need to close.

FDA and Regulatory Status
Not rated

No FDA device pathway applies and none is claimed, since the platform processes trial data for sponsors rather than diagnosing or treating patients. The regulatory relevance is indirect but real: outputs feed regulatory submissions, so the applicable framework is data integrity expectations such as ALCOA principles and 21 CFR Part 11 electronic records requirements rather than device clearance. No statement of compliance with those specific frameworks was located, which is worth asking about given the submission context.

AI Governance and Bias Disclosure
Not rated

No governance framework or bias analysis was located. The relevant risk in this domain is not demographic bias but systematic query bias: if automated data review flags certain data patterns disproportionately, it shapes which discrepancies get investigated before database lock, and nothing was located describing how that is monitored.

Integration and Deployment
EHR and Interoperability Depth
B
Vendor Published

Integration with existing clinical systems is the stated design premise rather than a feature, with the company positioning the platform as an operational layer that works with the systems clinical teams already use rather than requiring migration. Seed funding is explicitly earmarked to deepen integrations with clinical systems used by pharmaceutical, biotech, and medtech companies. This is the clinical trial technology stack, electronic data capture and related systems, not the EHR. No named connector list was located, which is the natural gap at seed stage.

Deployment Model and Data Residency
C
Vendor Published

Siloed per customer environments isolated from other customer data are stated, which is a meaningful architectural commitment for sponsors handling competitively sensitive trial data. Specific hosting, tenancy, and geographic residency terms are not published, and residency matters for multinational trials subject to differing data protection regimes.

Commercial
Commercial Transparency
C
Vendor Published

No pricing is published, but the value case is quantified against a costable baseline: an economic analysis of a representative Phase III oncology trial estimated savings exceeding 5 million dollars per study, driven primarily by earlier database lock and compressed timelines. Sponsors can model that against their own trial economics, since delayed database lock has a known cost. The figure is vendor generated and models a representative rather than actual trial. Buyers should also weigh vendor durability, since this is a seed stage company with 3 million dollars raised operating in a market served by established CROs.

Setting and Specialty Coverage
C
Vendor Published

Focused on the clinical trial data operations layer, spanning data cleaning, review, reconciliation, and reporting, sold to pharmaceutical sponsors, biotech, medtech, and contract research organizations. Deliberately narrow within clinical development, addressing the data management function rather than trial design, patient matching, or simulation, which distinguishes it from Unlearn, Mendel, QuantHealth, and the matching vendors indexed here. Published evidence centres on Phase III oncology, though the capability is not disease specific.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Contact the vendor
Undisclosed. Sold to pharmaceutical sponsors, biotech, medtech, and contract research organizations; no rates published. Not disclosed. Customers are pharmaceutical sponsors and CROs rather than providers, so contracting runs through sponsor data agreements rather than HIPAA business associate terms. Not disclosed. Designed to integrate with existing clinical systems rather than requiring migration, with seed funding earmarked to deepen those integrations. Vendor Published

No pricing is published, but the value case is quantified against a baseline sponsors can cost themselves: an economic analysis of a representative Phase III oncology trial estimated savings exceeding 5 million dollars per study, driven primarily by earlier database lock and compressed timelines. Delayed database lock has a known cost to any sponsor, so that is a modellable comparison rather than an abstract efficiency claim. Two caveats. The figure is vendor generated and models a representative rather than an actual trial, and the supporting controlled study covered 10 medical reviewers. Vendor durability is the other consideration: this is a seed stage company that raised 3 million dollars in July 2026, competing in a function currently served by established contract research organizations, and sponsors will likely require a SOC 2 or equivalent attestation that was not located.

AI Health Index

An independent reference for evaluating AI vendors in healthcare. No vendor pays for inclusion, placement, or rating.

Index Status
Last index update
July 21, 2026
The AI Health Index is an editorial reference, not a regulatory body. Vendor data is verified against published sources and public regulatory filings. Figures labeled “Estimated” have not been confirmed by the vendor. See the Methodology page for evaluation standards and limitations.
© 2026 AI Health Index
3801 N Capital of Texas Hwy, Ste E240 · Austin, TX 78746