Medidata
Indexed for the AI capabilities layered across the Medidata clinical trial platform rather than for the platform itself, which is treated as context under this index's product scoping rule. A Dassault Systemes brand, the underlying platform spans more than 38,000 trials and 12 million patients across roughly 2,300 customers and over one million registered users, anchored by Rave EDC.
The AI products indexed here: Clinical Data Studio, an AI data quality management workspace that integrates Medidata and non Medidata sources, identifies data issues and safety signals, and is reported by one named customer to deliver up to 80 percent faster data review; Medidata AI Study Build within Designer, which automates study construction; and an AI imaging capability introduced at ASCO 2026 using proprietary algorithms including automated text detection that the company reports makes protected health information redaction 32 percent faster.
Health Record Connect uses FHIR and health information exchanges to pull patient health records into trial data capture, reducing manual re entry at sites. The company reports its AI has supported more than 500 clinical studies over a decade, with more than 120 AI supported studies starting in 2025. Named customers for the AI products include Eisai.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The third case of this grade in the index alongside Elation and Waystar, and descriptive rather than critical. What a sponsor buys is the clinical trial platform, anchored by Rave EDC across more than 38,000 trials; the AI capabilities are woven into that platform and accelerate work inside it. The company's own framing, that it weaves intelligence into more solutions across a unified platform, correctly places the AI as a layer. No sponsor selects Medidata for the AI alone, which is precisely the comparison a buyer needs against a specialist such as Saama or Deep 6.
The capabilities differ in how much they decide, and none of the oversight design is published.
The three indexed capabilities sit at different points. Data quality management surfaces issues and signals for a reviewer to act on, which is assistive by construction. Automated study construction generates configuration a study team presumably reviews before a trial opens, though nothing states that it must. Automated redaction is the one that operates on the data itself and produces an output that downstream users consume as finished.
That third case is where the absence matters most. Redaction either happened correctly or it did not, and a reviewer looking at a redacted image cannot see what was missed. Unlike a flagged discrepancy, which announces itself and invites judgement, a silent redaction failure presents as a clean result. Nothing published states whether output is sampled or verified, at what rate, by whom, or what happens when a miss is found after images have been distributed.
For the study construction capability the equivalent question is whether generated configuration can reach a live study without a human approving it, and what the approval record looks like in a validated environment where every change must be attributable.
Nothing located answers any of this. Ask for the human checkpoints in each capability, whether any can be disabled or bypassed for throughput, and how an action taken by a model is attributed in the audit trail so a reviewer can distinguish it from one taken by a person.
Named capabilities, undescribed methods.
The products are individually named and their functions are clearly stated, which is more than many platform vendors manage: a data quality workspace that ingests both first party and third party sources, an automated study construction capability inside the design tool, an imaging capability, and a record connection service built on recognised interoperability standards. A buyer can tell what each is for.
What is not published is what any of them is. No model or method is named, no architecture is described, no training data is characterised, no versioning or update practice is documented, and no technical documentation was located for any capability. The strongest technical description found is that the imaging capability uses proprietary algorithms including automated text detection, which names a category rather than a method.
The reported figures follow the same pattern. A named customer reports up to 80 percent faster data review and the company reports 32 percent faster redaction. Both are plausible operational improvements. Neither has a published denominator, baseline or method, and a percentage improvement in speed says nothing about whether the output is correct.
For an axis about what a buyer can know about the technology they are licensing, that leaves the answer at the level of the brochure. Ask which models underlie each capability, whether any are third party or externally hosted, how versions are managed and communicated across a validated environment, and what evidence exists for output quality rather than throughput.
This vendor publishes the contract, which almost nothing else in this index does and which is the artifact this axis most wants. A data processing exhibit to the services agreement is public, including standard contractual clauses for international transfer, so a counterparty can read the processing terms before entering a negotiation rather than discovering them in a redline.
The certifications behind it are independently examined rather than asserted, including a privacy information management certification and an attestation carried out with privacy trust principles included rather than security alone, which is a narrower and more relevant scope than the security only attestations most vendors present.
The products also contain a stewardship control operating on the data rather than a statement about it, since the imaging capability performs automated detection and redaction of identifying information in study images. Held below the top grade on the question that matters most for a platform of this shape.
Nothing located states whether customer trial data is used to develop, train or refine models offered back to other customers, and this platform serves competing sponsors, so a sponsor's own trial data improving a capability sold to a rival is the first thing to establish rather than the last. Ask for that in writing, ask for a sub processor list, and ask what is retained after a study closes and locks.
Deployment scale is exceptional and long standing, with the AI reported across more than 500 clinical studies over a decade and more than 120 AI supported studies starting in 2025, against a platform footprint of 38,000 trials and 12 million patients. Named customer Eisai reports up to 80 percent faster data review with Clinical Data Studio, and the imaging capability reports 32 percent faster PHI redaction.
Held back from A because the figures are customer or vendor stated without published methodology, and the platform's scale should not be read as evidence for the newer AI capabilities specifically, which is the same distinction applied to Medable.
Strong on the instruments and unanswered on the question that matters most for the capabilities indexed here.
The instruments are real and independently examined. The company holds a privacy information management certification and a SOC 2 carried out with privacy trust principles included rather than security alone, which it describes as among the first such independent privacy certifications in the sector. It publishes a data processing exhibit to its services agreement, including standard contractual clauses for international transfer, so a counterparty can read the processing terms before contracting rather than negotiating blind. Very few vendors in this index publish the contract.
The products also contain a stewardship feature rather than only a policy: the imaging capability performs automated detection and redaction of identifying information in study images, which is a control operating on the data rather than a statement about it.
That same feature is where the open question sits. Automated redaction has a specific and asymmetric failure mode. A missed identifier does not announce itself; it travels onward into a dataset that reviewers, sponsors and regulators will handle on the assumption that redaction succeeded. The published figure for this capability is a speed improvement. No recall figure, no residual leakage rate and no human verification step is described.
The second open question applies to the whole AI layer. Nothing located states whether customer trial data is used to develop, train or refine the models offered back to other customers, which is the first thing a sponsor should ask of any vendor operating a shared platform across competing sponsors.
Ask for the redaction recall rate and the verification workflow, and for a written position on model training.
Correctly scoped, well documented for the regime that governs most of the business, and thinner on the one pathway that reaches into provider data.
For the majority of what this platform does, the health privacy rule is not the operative framework. The customer is a trial sponsor, the data is study data collected under a protocol with participant consent, and the governing instruments are good clinical practice, the sponsor's own obligations and the site's institutional review board approval. The company publishes a data processing exhibit covering personal data with standard contractual clauses for cross border transfer, and holds an independent privacy certification, which is the right assurance for that framework.
One capability crosses into different territory and should be assessed separately. The record connection feature pulls existing patient health records from provider systems and health information exchanges into trial data capture, so that sites populate study forms from the chart rather than re keying. That is a genuinely useful design, and it means the platform is receiving identifiable records that originate in a covered entity's systems rather than data generated by the study. Where that happens in the United States, a business associate relationship is the ordinary consequence.
Nothing located states the company's position on that: no availability statement, no role characterisation for this pathway, and no identification of which entity signs. A sponsor will not usually ask, because the sponsor is not the covered entity. The provider organisation supplying the records is, and it should.
Ask which agreement governs the record connection pathway specifically, what is retrieved and retained, and whether retrieval is scoped to enrolled participants or extends to screening.
The most complete security disclosure located anywhere in this index, and the grade rests on what the company publishes rather than on how much it holds.
The certification set is current and specific: ISO 27001:2022, statements of conformity to the cloud controls standard and the cloud personal information standard, and a privacy information management certificate. Alongside them sit a SOC 2 Type 2 carried out with additional privacy controls, a SOC 1 Type 2, a separate SOC 1 for the site payments offering, payment card standard validation for that same offering, and assessments against the United States federal control catalogue. The SOC evaluation runs on a six month cycle across a rolling twelve month population rather than annually, which is a materially tighter cadence than the norm.
What lifts this to A is the trust centre itself. It publishes penetration test results and vulnerability scan summaries, which almost no vendor in this index does at all, a cloud security alliance registry entry, an information security programme fact sheet, a security whitepaper in four languages, a responsible disclosure programme, and a standing statement on security incidents at study sites. That last document is unusual and worth noting: it addresses incidents occurring at the customer's own sites rather than at the vendor, which is the direction of exposure a sponsor actually worries about and which most vendors ignore entirely.
The practical test this index applies is whether a counterparty can get what it needs without a sales conversation, and here it largely can, with some documents behind a customer login.
One thing to confirm rather than assume. This disclosure describes the platform's security programme. Ask specifically whether the newer capabilities indexed on this record sit inside the audited scope, and read the scope section of the report rather than the summary.
No device pathway applies and none is claimed. What makes this an A rather than a scoping determination is that the framework which does apply is addressed more completely here than anywhere else in this category.
The products support the conduct of clinical trials. They do not diagnose or treat anyone, so device regulation does not reach them. The regime that does reach them governs electronic records and signatures, good clinical practice, and computerised systems used in regulated studies, and it is where a sponsor's inspection risk actually sits.
The company maintains a regulatory compliance library of documented position statements written for customers, setting out how the platform meets specific instruments by name. The set spans the United States electronic records and signatures rule and the agency's own question and answer guidance on electronic systems in clinical investigations, the European good manufacturing practice annex on computerised systems, both the current and the revised international good clinical practice guidelines, and the requirements of the Japanese and Chinese regulators. Publishing a position per instrument, per jurisdiction, is what a sponsor needs when its own inspection readiness depends on the vendor's, and it is rare.
For calibration within this index, another vendor assessed in the same pass is correctly outside device regulation on identical reasoning and publishes nothing at all about the electronic records framework its outputs feed. The two sit at opposite ends of the same axis for exactly that reason.
What a buyer should still confirm. Position statements are the vendor's own characterisation, not a regulator's finding. Ask for the validation documentation package, for how it is maintained across releases, and specifically for how the newer capabilities on this record are validated, since automated study construction and automated redaction change what a validated state means.
The disclosure posture that makes this vendor exceptional elsewhere does not extend to the AI capabilities themselves, and the contrast is the finding.
This company publishes certificates, penetration test results, vulnerability summaries, a data processing exhibit and per regulation compliance positions. It publishes nothing located on how the AI capabilities behave: no model documentation, no evaluation methodology, no error analysis, no statement of where a capability should not be relied upon, and no governance framework covering the AI layer.
The domain relevant questions are specific and each is answerable. The data quality workspace identifies data issues and safety signals across sources, so the material question is whether flagging is uniform: a detector can be well calibrated in aggregate while over querying particular sites, regions, therapeutic areas or source systems, and which discrepancies get investigated before database lock shapes the dataset a regulator eventually reviews. The automated study construction capability raises a different one, since a model trained on prior protocols will reproduce the conventions of the trials it learned from, including their eligibility patterns, and eligibility patterns are where trial populations narrow. The redaction capability raises a third, which is whether detection performance varies with document type, language or image modality.
None of the three is addressed. The published figures for these capabilities are speed improvements, reported by the vendor or a named customer, without methodology.
This is a gap rather than a deficiency in posture, and it should be readable as such: a company that publishes this much elsewhere can publish here, and the questions above are the ones to put.
Two passes located no accuracy figure, no evaluation methodology, no model or method named for any capability, and no warranty, indemnity or remediation commitment. The reported figures measure throughput rather than correctness: a named customer reports up to eighty per cent faster data review, which uses the ceiling construction this index treats as an absence of a number, and the company reports a percentage improvement in redaction speed with no denominator, baseline or method.
A percentage improvement in speed says nothing about whether the output is right. One capability makes that gap consequential rather than merely unsatisfying, and its failure mode deserves naming because it is asymmetric and silent.
Automated redaction of identifiers in study images either catches an identifier or it does not, and a missed one does not announce itself: it travels onward into a dataset that reviewers, sponsors and regulators will handle on the assumption that redaction succeeded, and the discovery event, if it comes, is a disclosure incident rather than a quality finding. The published figure for this capability is a speed improvement.
No recall figure, no residual leakage rate and no human verification step is described, and recall is the only number that matters for a control of this kind. Ask for redaction recall against a labelled test set, the residual leakage rate, the verification workflow, and what happens when an identifier is found downstream.
Health Record Connect is the strongest element here and solves a real problem: it uses FHIR and health information exchanges to pull existing patient health records directly into trial data capture, so sites complete Rave EDC forms from the chart rather than re keying. Clinical Data Studio separately integrates non Medidata sources including third party labs and other electronic data capture systems, which matters because sponsors rarely run a single vendor stack.
Cloud delivered, with the transfer question answered contractually and the residency question not answered publicly.
What is documented. The company holds statements of conformity to both the cloud controls standard and the cloud personal information standard, which are the certifications specific to operating as a cloud service provider rather than to information security generally, and it has been assessed against the United States federal control catalogue. It publishes a data processing exhibit incorporating standard contractual clauses, which is the mechanism that makes international transfer lawful under European rules and is the practical answer for a global sponsor. Security documentation is published in four languages, which is a reasonable proxy for where the customer base sits.
What is not documented in the material located: the regions in which data is held, whether a sponsor can elect a region, whether tenancy is dedicated or shared between sponsors, and the subprocessor list. The tenancy question deserves particular weight here. This platform serves a large share of the industry, so competing sponsors running competing programmes in the same therapeutic area are customers of the same environment simultaneously, and the isolation model is a competitive question as much as a security one.
Ask for the region list and whether residency can be elected, for the tenancy and isolation model between sponsors, and for the subprocessor register. All three are ordinary requests for a vendor with this disclosure posture and should be readily answered.
No public pricing. Contact the vendor. Enterprise agreements with sponsors, typically scoped by study count and module. As with other platform vendors indexed under the scope rule, buyers should establish whether AI capabilities carry incremental cost or are included in platform licensing, since that determines the comparison against a separately priced specialist.
Clearly bounded to sponsor side clinical development for biopharmaceutical and medtech customers, spanning study build, data capture, data quality management, imaging, and patient facing capture. No claims outside clinical research.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Contact the vendor
|
Enterprise sponsor agreements scoped by study count and module | — | — | Vendor Published |
No rate card published. Enterprise agreements with pharmaceutical and medtech sponsors, typically scoped by study count and module selection. The recurring question for platform vendors indexed under this index's product scoping rule applies here: establish whether the AI capabilities carry incremental cost or are bundled into platform licensing, since a bundled capability compares very differently against a specialist priced separately.