xCures
AI platform that retrieves, aggregates, and structures fragmented medical records into longitudinal patient histories, originally built for oncology and since extended to all therapeutic areas. The Clinical Clarity Engine applies natural language processing and machine learning to unstructured records; the company reports processing more than 300 million medical records from over 550,000 healthcare locations, with every structured output linked back to its source document.
The xDECIDE provider portal presents structured data, autogenerated summaries, and AI assisted treatment option reports with scientific rationale, ranked by the xCORE engine and reviewed through a virtual tumor board; xINFORM is the patient facing portal. Data products supply longitudinal real world oncology datasets to research and regulatory programs. Holds HITRUST e1 certification. Raised a $46 million Series B in June 2026.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
NLP and machine learning applied to unstructured medical records is the product. Without the extraction and normalization models there is no structured history, only a pile of retrieved documents.
Oversight is structural rather than asserted: treatment option outputs from the xCORE ranking engine pass through a virtual tumor board that provides expert review, training, and validation of the algorithm's outputs. Every structured data element is also linked back to its immutable source document, which makes any individual output auditable rather than opaque.
This is among the most complete model disclosures in the index. xCures names two distinct extraction approaches and describes how each works: schema based extraction applying named entity recognition and relation extraction to unstructured documents and normalising output to FHIR R4 and OHDSI vocabularies, and checklist based assertion using retrieval augmented generation across the longitudinal record with an explicit evidence hierarchy for resolving conflicting documentation, such as prioritising a pathology report over a clinic note for cancer staging.
It publishes a performance table. Four extractors are reported with accuracy, precision, recall and F1 against clinically trained reviewers, measured on a random ten per cent audit with third reviewer arbitration, ranging from 95.7 to 98.2 per cent accuracy. It publishes a deployment threshold, striving for accuracy and precision at or above 95 per cent before an extractor enters production. It names the source as a preprint with full methods, supplemental tables and raw counts available on request, and it states that the figures are retrospective and do not guarantee future performance.
Two disclosures go further than anything comparable. The scoring rule counts only explicitly stated, verifiable extractions as correct, so a correct inference that is not present verbatim in the source document is recorded as an error. A vendor choosing to score its own correct answers as failures is applying a stricter standard than the market requires. And the published limitations include that both approaches rely on semantic search over the top ranked documents only, so relevant information in lower ranked documents can be missed. For a product whose purpose is assembling a complete patient history, that is the most consequential failure mode of its own architecture, published voluntarily.
Remaining ask: the performance table covers four extractors, and a buyer should ask for the equivalent figures for whichever extractors and checklists their own use case depends on.
The controls are specific, the infrastructure is named and the commercial boundary is drawn on both sides, which together put this well above the lane. Encryption in transit and at rest, role based access applying least privilege and the minimum necessary principle by name, multi factor authentication and single sign on, immutable audit logs through named services, and full audit logging across every data access and modification event.
Source documents are preserved in the form received alongside the structured output, which is a provenance feature and a safeguard at once, because a disputed extraction can be checked against what actually arrived. The boundary statement is what distinguishes this from the unbounded permissions recorded elsewhere in this index: the notice says the business does not involve the sale of individually identifiable information or its use for marketing, and then says plainly that the company does sell de identified real world data and evidence where it has a legal basis.
Naming both halves is better practice than silence and far better than an unlimited grant. A further piece of candour deserves recording: the notice states that most collection happens at the direction and with the consent of patients, but that sometimes health information is collected without patient consent, and explains the legal basis. Very few vendors write that sentence down.
Held at B because no retention period is published and deletion is described only as occurring in accordance with contractual requirements, which makes it a negotiable term rather than a commitment. Ask for the retention schedule and which categories of de identified data are sold and to whom.
Peer reviewed publication describing the platform architecture and its clinical decision support algorithm, including in the North Carolina Medical Journal and in AI in Precision Oncology. Publishing the method rather than only outcome claims is uncommon in this index and is what the axis rewards. Operational scale is separately reported at more than 300 million records from over 550,000 locations. Graded on the existence and venue of the evidence; this index does not re verify the underlying studies.
The controls are specific and the boundaries are stated. Encryption in transit and at rest, role based access control applying both least privilege and the HIPAA minimum necessary principle by name, multi factor authentication and single sign on, immutable audit logs through AWS and Datadog with monitoring, and full create, read, update and delete audit logging across all data access and modification events. Source documents are preserved in the form they were received alongside the structured output, which is a real safeguard as well as a provenance feature.
The commercial boundary is drawn explicitly, which is uncommon and creditable. The privacy notice states that the business does not involve the sale of individually identifiable information or its use for marketing, and then states plainly that the company does sell de identified real world data and evidence where it has a legal basis. Naming both halves is better practice than the unrestricted permissions found elsewhere in this category, and considerably better than silence.
A further piece of candour deserves recording. The notice states that most collection happens at the direction and with the consent of patients, but that sometimes health information is collected without patient consent, and it explains the legal basis for that. Very few vendors write that sentence down.
Held at B on two counts. No retention period is published; deletion is described only as occurring in accordance with contractual requirements, which makes it a negotiable term rather than a stated commitment. And the sale of de identified derivatives, while bounded by a legal basis test, is a secondary use of patient derived data that a covered entity should understand before contracting. Buyers should ask for the retention schedule in writing and ask what categories of de identified data are sold and to whom.
xCures states without qualification that it operates as its customer's business associate, and its privacy notice explains what that role means in practice, including that where it collects health information without patient consent its access, use and sharing are determined and limited by its contractual and legal obligations as a business associate or as a clinical research organisation. That is the affirmative self identification the index has found missing across most of the vendors it has assessed.
The distinguishing element is ongoing validation rather than a one time claim. The company states it conducts an annual HIPAA evaluation to validate compliance, and further assessments as part of its HITRUST certification programme. A recurring formal evaluation is materially more useful to a buyer than a static assertion, because it speaks to whether the position is maintained rather than whether it was once true.
This reaches the top band by a different route from the one other vendor at this level in the category, which publishes its business associate agreement in full. Here the agreement itself is not published, and that is the gap: a buyer can see the role and the evaluation cadence but not the allocation of obligations. Buyers should request the agreement, ask for the date and scope of the most recent annual evaluation, and ask who performs it.
xCures maintains a public trust page carrying a control level table and a certification statement. Named controls: encryption in transit over TLS and at rest, role based access control applying least privilege and the HIPAA minimum necessary principle, multi factor authentication and single sign on, immutable audit logs through AWS and Datadog with security monitoring, data deletion in accordance with contractual requirements, and full create, read, update and delete audit logging across all data access and modification events.
The certification position requires care to read, and the company deserves credit for describing it accurately rather than leaving the badges to imply something stronger. The engine holds HITRUST e1 certification, achieved in 2025, with an upgrade to HITRUST i1 stated to be in progress. e1 is the entry tier of the HITRUST framework, below both i1 and the risk based r2 assessment. The HITRUST r2 certification, ISO 27001:2022 certification and SOC 2 Type 2 attestation also displayed belong to Amazon Web Services, the underlying infrastructure provider, and the company states plainly that its own certification inherits and leverages selected controls assessed in the AWS r2 certification, and that it periodically reviews the AWS reports to validate ongoing adherence.
That distinction matters more than it appears. A cloud provider's certification covers the platform beneath the application, not the application, its data handling or its staff. Inheriting controls is legitimate and common, and disclosing it is better practice than most, but a buyer relying on the r2, ISO or SOC 2 marks on this page would be relying on assurance about infrastructure rather than about the service being purchased.
Held at B on the entry tier certification, the absence of any attestation held in the company's own name beyond HITRUST e1, no penetration testing statement, no subprocessor list, and no audit dates published for the stated annual HIPAA evaluation. Buyers should confirm whether the i1 upgrade has completed, request the certification letter and its scope, and ask which controls were assessed directly versus inherited.
No clearance, authorisation or published regulatory position was located. The grade reflects the absence of a stated position rather than a judgement that clearance is required.
On the clinical decision support exclusion under the 21st Century Cures Act, this product has an unusually strong argument available to it and does not make it publicly. The exclusion turns substantially on whether the clinician can independently review the basis for the output. Here every extracted element carries provenance metadata identifying the originating document, checklist assertions cite the specific documents consulted and the hierarchy used to resolve conflicts, and the source document is preserved in its original form alongside the structured result. That is about as close to independently reviewable as a structured output gets.
Two complications are worth putting to the vendor. The provider portal presents AI assisted treatment option reports with medical rationales, which is further along the spectrum than record structuring. And the stated use cases now include risk adjustment coding capture, quality measure reporting, prior authorisation and revenue cycle work, which are reimbursement functions with their own oversight and incentive considerations rather than clinical ones. Ask the company to state its regulatory position and to say which of its output types it considers covered.
The validation governance is genuinely strong and it is published rather than asserted. Every extractor follows a stated lifecycle: independent assessment by clinically trained reviewers against source documents, classification of each field as true or false positive or negative, arbitration of discrepancies by a third reviewer with access to clinical experts, capture of errors as edge cases used to refine prompts and retrieval parameters, and version controlled models supporting rollback with A B testing across prompts, models and hyperparameters. Hallucination is handled as a defined scoring criterion rather than a general assurance.
The company also publishes its known limitations, including that extraction quality from scanned or faxed documents depends on the quality of those documents, and that semantic search retrieves only the top ranked documents so relevant material can be missed.
Held below the top band because none of this is broken down by population. There is no analysis of how accuracy varies by patient group, care setting, document source or record quality, and the company's own disclosed limitation points directly at where such an analysis would matter. Document quality is not randomly distributed: scanned and faxed records are more common in smaller, less digitised and less well resourced practices, so an accuracy gradient tied to scan quality is likely also an accuracy gradient tied to where a patient receives care. A vendor already running a rigorous audit methodology is unusually well placed to stratify those results and publish them. Doing so would move this to the top band.
This is the strongest published evidence record located in the backfill, and two choices inside it go beyond anything else here. The ordinary excellence first: a performance table reports four extractors with accuracy, precision, recall and a combined score against clinically trained reviewers, measured on a random ten per cent audit with third reviewer arbitration, ranging from roughly ninety six to ninety eight per cent accuracy.
A deployment threshold is published, with extractors expected to reach accuracy and precision at or above ninety five per cent before entering production, so a buyer knows the bar a component had to clear. Full methods and raw counts are offered, and the figures are explicitly labelled retrospective rather than predictive. Now the two that matter most.
The scoring rule counts only explicitly stated, verifiable extractions as correct, so a correct inference not present verbatim in the source is recorded as an error: a vendor choosing to score its own right answers as failures is applying a stricter standard than the market requires and reporting worse numbers as a result.
And the published limitations state that both approaches rely on semantic search over top ranked documents only, so information in lower ranked documents can be missed, which for a product assembling a complete patient history is the most consequential failure mode of its own architecture, published voluntarily. Held below the top grade only because no warranty, indemnity or remediation commitment attaches. Ask for figures on the extractors your use case depends on.
Retrieval breadth is the core competency: connectivity with local, state, and national health information networks reaching a reported 550,000 distinct healthcare sites, delivered through both web interfaces and developer APIs that embed in existing workflow platforms. This is materially broader than integration with a handful of named EHRs.
Two routes are published: an off the shelf platform with a standalone secure web portal that can be stood up quickly, or integration of the engine into an organisation's existing operational systems and data lakes. Application programming interface documentation is publicly available, which is a meaningful signal that the integration route is real rather than aspirational. Separate portals exist for patients, providers and institutional partners.
The retrieval reach is the strongest deployment property and it is specific. The platform collects records from across the United States through health information exchange networks, qualified health information networks, direct electronic health record connections and traditional record request methods, with more than 180 million patient records stated as processed to date. For an organisation whose problem is that a patient's history is scattered across institutions it will never contract with individually, that reach is the product.
Held at B because the infrastructure position is single provider, running on AWS with no alternative stated, and because no data residency commitment naming regions was located and no implementation timeline or resourcing expectation is published for either route.
No pricing information of any kind was located. There is no list price, no pricing unit, no tier structure and no indication of whether the engine is sold per record retrieved, per patient, per extractor, per seat or as an annual platform fee. For a product with several distinct delivery shapes, from a standalone portal to an embedded engine sold through channel partners, the absence of a pricing unit is the more significant omission, because the shape of the charge determines how the product behaves once deployed.
The company publishes a Series B financing of 46 million dollars and names its investors, and it publishes application programming interface documentation openly, which is a form of commercial transparency in that a prospective integrator can assess the work involved before entering a sales process. Neither substitutes for a price.
This sits against an unusually well developed published upside. The company documents extraction accuracy figures, a validation methodology, a deployment threshold and named use cases with specific operational claims. A buyer is given a great deal with which to assess whether the product works and nothing with which to assess what it costs. Buyers should establish the pricing unit before the figure, and should ask specifically how records retrieved from external institutions are charged, since retrieval volume is the variable a customer controls least.
The company began in oncology and now states that the platform is condition agnostic and supports all therapeutic areas, with published use cases spanning provider record review, diagnostics, value based care and channel partnerships where another vendor embeds the engine to solve its own customers' unstructured data problem.
The oncology depth is real and specifically evidenced. Extracted data is aligned to the HL7 Minimal Common Oncology Data Elements profile, the platform covers multiple tumour types, and named capabilities include lines of therapy and cancer staging at initial pathological diagnosis, which are oncology specific constructs that generic extraction does not handle well.
Held at B on a gap between claim and evidence that a buyer outside oncology should press. Every published validation figure is oncology shaped: cancer diagnosis, lines of therapy, surgical procedures and medications. The assertion of breadth across all therapeutic areas is not yet accompanied by published performance in any of them. A cardiology, nephrology or behavioural health buyer should ask for validation figures in their own domain rather than assuming the oncology numbers transfer, since the underlying documents, vocabularies and clinical constructs differ.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Contact the vendor
|
Platform software licenses and API access; separate data product agreements | — | — | Vendor Published |
Two commercial surfaces: software licenses and API access to the platform for provider and research organizations, and separately curated longitudinal dataset agreements with life sciences partners. No rate card published for either.