Healthcare Administrative Automation
I

Integral

Automates HIPAA Expert Determination, the statistical de-identification standard that governs whether a health dataset can lawfully be used for analytics, research, or model training. The platform detects and remediates PHI entities in structured and unstructured data, assesses re-identification risk, and produces a signed determination opinion from qualified statisticians, an approach the company describes as expert in the middle: automated analysis with human certifier sign off. Unlike Safe Harbor, which strips 18 fixed identifiers, Expert Determination is scoped per dataset and preserves more analytic utility.

For AI training corpora the platform replaces real PHI with synthetic entities rather than redacting, so text remains natural for model training. Deploys inside the customer's VPC, on premises, or air gapped environment as container images. Originated in healthcare and is extending the same methodology to financial, consumer, and AI training data, which buyers should weigh when assessing healthcare specific depth.

AI Health Index verifiedJuly 19, 2026
Compare Integral with other vendors
Founded
Headquarters
San Francisco, California
Categories
healthcare-admin-automation
Indexed Products
Privacy Workbench, Expert Determination
Buyer Segments
Pharma / Life Sciences, Payer, Large IDN
Assessment

Capability Axes

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read

AI Capability
BB on AI CentralityThe model is the engine of a core module. The platform carries other value, but this capability does not exist without it.
Vendor Published

Automated detection and remediation of PHI entities across structured and unstructured data is what makes the product fast, and the synthetic entity substitution approach for training corpora is a model driven capability rather than rule based redaction. Held back from A because the deliverable a customer buys is a signed Expert Determination opinion from a qualified human statistician; the automation accelerates the assessment but does not replace the certifier.

BB on Autonomy and Oversight ModelThe oversight structure is described and one part is missing, commonly the threshold at which the system stops or what happens after it is wrong.
Vendor Published

The oversight model is stated explicitly and is stronger than most in this index, because the human in it carries personal accountability rather than performing review.

The vendor describes an expert in the middle approach: software assesses re identification risk, streamlines the workflow and proposes remediations, and qualified human certifiers validate that work and provide expert oversight. Every engagement produces a signed determination, a formal opinion attributing the risk conclusion to a named expert.

That signature is what makes the arrangement unusual. Across this index, human review typically means a person checking output before it reaches someone else, with no consequence attaching to them personally. Here the expert's own professional standing is the artefact. The privacy rule does not define what qualifies an expert, and in an audit the regulator would examine that individual's experience, academic training and methods. So the reviewer is exposed in a way a queue reviewer is not, which aligns their incentives with the buyer's.

The questions concern the boundary. Establish what the software decides alone, whether a certifier can be presented with an automated recommendation they are unlikely to overturn, how often recommendations are rejected, and how many datasets one certifier signs.

BB on Model and Technology TransparencyThe approach or the suppliers are named without the version and update discipline behind them.
Vendor Published

Technical disclosure is better than most records in this index, and the reason is that the methods are the product rather than a proprietary edge to be protected.

The vendor names its techniques rather than describing capabilities in the abstract. It supports retention of linkage tokens so records can be joined across datasets without identifiers, preservation of geographic and temporal detail where risk permits, and named privacy models including differential privacy and k map analysis. Those are recognised approaches from the disclosure control literature, not invented terminology, and naming them lets a knowledgeable buyer evaluate the approach before engaging.

It also describes the workflow concretely: continuous privacy assessment as dataset configurations change, generation of a compliant copy alongside compliance documentation, and monitoring as data evolves.

What is not published is calibration. A risk model produces a number, and the useful question is how that number was validated. Establish what assumptions are made about an adversary's resources and auxiliary data, whether risk estimates have been tested against actual re identification attempts, and how the model treats external datasets that did not exist when a determination was issued.

Ask for the risk model's assumptions and any validation of its estimates.

AA on Model Supply Chain DisclosureEvery party is enumerated by name including the model layer. A public subprocessor list naming the model provider, with the retention and training terms that govern data once it arrives, is the canonical artefact.
Vendor Published

This is the product rather than a control wrapped around it, and the disclosure is correspondingly specific. The platform is built around the statutory expert determination route, supports differential privacy and formal risk analysis rather than only the enumerated identifier list, and reports measured rather than assumed utility retention, so a buyer learns what the de identification cost them as well as what it protected. The artefact that earns the top grade is the signed opinion.

It documents the privacy model applied, the remediation strategy and the risk justification, and it travels with the dataset, which makes it the thing this axis has been asking for on record after record: a portable, attributable, written statement of what was done to the data and on what basis, readable by whoever holds the dataset next rather than only by the party that commissioned it.

Almost every other vendor in this index offers an assurance that de identification happened; this one produces a document that says how, signed by someone who can be held to it. The company also draws an unusually precise distinction between securing the container and assessing the contents, which is exactly the confusion this index has had to correct against certification claims throughout. The residual is calibration and it is discussed on the other axis. Ask what assumptions the risk model makes about an adversary and how estimates have been validated.

CC on Clinical and Operational EvidenceNamed customers, or vendor reported percentages with no method, denominator or reference standard. Scale of use is recorded here and is not treated as evidence of benefit.
Vendor Published

No clinical evidence exists and none is required. The product does not touch care delivery, so trial and outcome literature has no application.

The evidence that would matter is of a different kind and only partly published. Operationally, the vendor makes specific and checkable claims: certification timelines reduced from months to days or hours, and customer references describing timelines halved and previously blocking compliance work becoming routine. Named customers attached to those statements make them stronger than most claims in this index, and a prospective buyer can reasonably ask to speak to them.

What is absent is evidence about the thing that actually matters, which is whether the determinations hold. The output of this product is an opinion that re identification risk is very small. The evidence that would support it is empirical: attempted re identification against remediated datasets, ideally by an independent party, reporting how many records could be linked back under realistic conditions.

That is a known and publishable form of testing, and no such work was located. It is the natural complement to the vendor's own argument that data must be defensible rather than merely stored securely.

Ask whether any adversarial re identification testing has been performed, by whom, and what it found.

AA on AI Safety and PHI StewardshipRetention windows, training use and de identification are stated specifically enough to be contradicted, alongside the safety engineering: guardrails, hallucination mitigation, and how a safety event is handled.
Vendor Published

This is the product, and the disclosure is correspondingly specific. The platform is built around HIPAA Expert Determination under 45 CFR 164.514(b)(1), supports differential privacy and k-map analysis rather than only Safe Harbor's 18 identifier list, produces a signed opinion documenting privacy model, remediation strategy, and risk justification that travels with the dataset, and reports measured rather than assumed utility retention. The company draws an unusually precise distinction between securing the container and assessing the contents.

Regulatory and Compliance
BB on HIPAA and BAA PostureBusiness associate status is stated and supported by a substantive privacy document, with the agreement or its scope not fully published. For a vendor outside the United States, an equivalent regime documented to this depth grades here.
Vendor Published

HIPAA expertise is the core competency and the company operates against the Expert Determination standard directly, with determinations also aligned to Washington MHMDA, California CPRA, and applicable state law. Held back from A because the vendor's own BAA terms and execution process, as distinct from its customers' compliance posture, were not published.

CC on Security Certifications and Trust CenterControls are described with an outside check behind them, such as independent penetration testing on a stated cadence, but no attestation against a recognised framework.
Vendor Published

No attestation, certification or trust centre was located, and this vendor makes an argument that bears directly on the gap.

Its own material distinguishes between securing infrastructure and making data defensible, observing that a secure data centre full of improperly de identified records is a compliant building with a liability problem inside it. That is correct and well put. An infrastructure attestation examines the container; a determination examines the contents, and neither substitutes for the other.

The argument runs both ways, which is the finding. Defensible data held inside an unexamined environment is the mirror image of the problem the vendor describes. This company ingests identified patient records in order to remediate them, so before any determination exists it is holding fully identified health information at scale on behalf of pharmaceutical, insurance and analytics customers. That period is exactly when an infrastructure attestation matters most.

So a buyer should want both, and the vendor is unusually well placed to explain why. Establish what protects the pre remediation environment specifically, how identified source data is segregated from remediated output, and how long the original is retained after a determination is issued.

Ask which report is held or scheduled and what its boundary covers.

BB on FDA and Regulatory StatusThe pathway is stated and in progress, or a clearance is named without the vintage and scope a buyer needs to match it to the product on offer.
Vendor Published

No clearance exists and none should. This is privacy engineering applied to datasets. It makes no clinical claim and no device framework attaches, so this axis should not read as an absence.

What governs is the de identification standard itself, and the vendor operates inside it rather than adjacent to it. The privacy rule recognises two methods: a fixed list of identifiers to remove, and a determination by a qualified expert applying statistical principles that re identification risk is very small. This company's product is the second method, delivered as software with expert sign off.

Two features of that standard are worth a buyer understanding. The rule sets no expiry on a determination, but the regulator has acknowledged that technology, social conditions and the availability of external information change over time, which is the argument for the continuous monitoring this vendor offers rather than a one off certificate. And the fixed identifier list is now decades old and predates most of the external data that makes re identification feasible, which is the argument for the method this vendor sells.

Beyond the federal rule, state privacy statutes increasingly reach health data including some de identified categories, and they do not all adopt the same definitions.

Ask how determinations are refreshed and how state variation is handled.

CC on AI Governance and Bias DisclosureResponsible artificial intelligence is committed to in policy language with no evaluation behind it. Most of the index sits here.
Vendor Published

No fairness statement or subgroup disclosure was located, and the axis needs restating for this product because the relevant question is not model bias in the usual sense.

Re identification risk is not distributed evenly across a population, and that is the governance issue here. A determination concludes that the risk of identifying an individual is very small, but that conclusion is reached across a dataset. Within it, some people are far more identifiable than others: a person with a rare diagnosis, an unusual combination of attributes, an outlier age, or residence in a sparsely populated area stands out precisely because they are unusual. Aggregate risk can be very small while individual risk for those people is materially higher.

That matters because the people who are most identifiable are often those whose records are most sensitive, and they have no visibility into the determination made about them.

Standard practice addresses this through the techniques the vendor already names, including thresholds that constrain small cells. So the questions are answerable. Ask what minimum group size is enforced, how outlier records are handled, whether suppression or generalisation is applied and how utility loss is traded against it, and whether risk is reported at the individual maximum as well as the population average.

BB on AI Liability and RecourseA published falsifiable commitment, or a real correction route for the affected person. A published error rate with its method and denominator grades here, and so does a jurisdiction whose law gives the patient an enforceable right to correct an inaccurate record.
Vendor Published

The methods are the product rather than a proprietary edge to be protected, and that shows in how they are described. The vendor names its techniques rather than describing capabilities in the abstract, covering retention of linkage tokens so records can be joined across datasets without identifiers, preservation of geographic and temporal detail where risk permits, and named privacy models drawn from the published disclosure control literature rather than from invented terminology.

That distinction matters, because a named method from a public literature can be evaluated by a knowledgeable buyer before they engage, and an invented one can only be evaluated by trusting the vendor. The workflow is also described concretely, with continuous privacy assessment as dataset configurations change and generation of a compliant copy alongside its documentation. Held below the top grade because calibration is unpublished.

A risk model produces a number, and the useful question is how that number was validated: what assumptions are made about an adversary's resources and auxiliary data, whether risk estimates have been tested against actual re identification attempts, and how the model treats external datasets that did not exist when a determination was issued.

That last is the sharpest, because a determination is a statement about the world at a moment and the world's stock of linkable data only grows, so a valid opinion can decay without anyone acting. Ask how determinations are revisited.

Integration and Deployment
CC on EHR and Interoperability DepthIntegration is claimed through standards or a middleware layer with no system named and nothing to verify.
Vendor Published

This axis maps imperfectly and the scoping should be stated rather than the row read as a deficiency. The product does not connect to record systems, is not deployed in clinical workflow, and has no reason to interoperate with an electronic health record. Its inputs are datasets a customer already holds.

The equivalent question is how data reaches it and how remediated data leaves, and that is where a buyer should press. The platform is described as covering the full lifecycle from ingestion to delivery, merging datasets including clinical with consumer data, and producing a compliant copy alongside documentation. So there is an ingestion path, a transformation stage and an output path, and none is described in technical terms publicly.

Establish what formats and interfaces are supported for ingestion, whether the vendor pulls from a customer's warehouse or receives an extract, how remediated datasets are delivered and in what form, and whether linkage tokens issued in one engagement remain usable across later ones.

That last question matters more than it appears. Tokens that persist across datasets and over time are what make longitudinal analysis possible, and they are also the mechanism by which separately remediated datasets can be joined.

Ask for the ingestion and delivery interfaces, and the token lifecycle.

AA on Deployment Model and Data ResidencyDeployment options, residency and tenant isolation are all documented, including where data rests and which processing crosses a border.
Vendor Published

Strongest deployment disclosure in the index. Deploys inside the customer's VPC, on premises data center, or air gapped environment as container images running the same core engine, so regulated data never leaves the customer network. For a product whose entire function is handling PHI, that architecture is the material fact.

Commercial
CC on Commercial TransparencyNo price is published and the posture is discoverable: a buyer can establish how the product is sold and what drives the cost before contacting the vendor. Most of the index sits here.
Vendor Published

No public pricing. Contact the vendor. Engagements are scoped per dataset and determination; no rate card published.

BB on Setting and Specialty CoverageCoverage is named with validation behind part of it.
Vendor Published

Coverage should be read as data types and customer sectors rather than clinical settings, since this product is not deployed in care delivery.

On that reading it is broad and the vendor is clear about it. Customers span pharmaceutical, insurance, health technology and analytics organisations, with named references including a data company and an agency group's health data arm. The platform handles dataset merging as well as remediation, including combining clinical data with consumer data, which the vendor identifies as creating new data that did not previously exist.

That last point defines the real scope of this product better than any sector list. Its work begins where datasets meet, and the combinations it enables are the ones that raise the hardest privacy questions, because attributes that are innocuous separately can identify a person together.

One aspect of the customer base deserves stating plainly rather than obscuring. Health data monetisation is an explicit and marketed use case. That is lawful where data is properly de identified and it is the commercial engine of much real world evidence research. It also means the people whose records are being remediated are not the customer and did not choose the arrangement, which is the ordinary position for secondary use and worth naming.

Ask which data types and combinations a determination covers.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Contact the vendor
Scoped per dataset and per Expert Determination engagement Vendor Published

Engagements are scoped per dataset and per determination rather than sold as a flat license, which matters because the unit of cost is the assessment, not the seat. No rate card published.