Helix
Enterprise genomics platform for health systems, payers, and life sciences. Health systems run population genomics programs on Helix infrastructure, covering recruitment, clinically actionable disease screening, return of results, and research, built on the proprietary Exome+ assay and a CLIA certified, CAP accredited sequencing lab. The Helix Research Network spans roughly 20 health system members and a reported 400,000+ sequenced participants, feeding GenoSphere, a linked clinico-genomic dataset that passed 500,000 records in June 2026, each combining Exome+ sequencing with an average of 13 years of EHR history and 8 years of claims.
AI tooling is a recent layer rather than the foundation: an AI Cohort Builder launched June 2026 and a Model Context Protocol connector for querying the dataset launched July 2026. Health system programs include Cone Health, Sanford Health, WellSpan, St. Luke's University Health Network, HealthPartners, Nebraska Medicine, and Memorial Hermann.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The core of the product is sequencing chemistry (the Exome+ assay), a CLIA certified lab, and the data linkage that produces GenoSphere. AI tooling is real but recent and additive: a Cohort Builder for research cohort assembly launched June 2026 and an MCP connector for querying the dataset launched July 2026. The company markets AI powered genomic intelligence, but the asset a buyer is acquiring is sequencing capability and a linked dataset, not a model. Included in this index because the AI capability is genuine; graded here because it is a layer rather than the mechanism.
Two quite different activities sit on this record and the oversight design is published for neither.
The clinical activity is population screening with return of actionable results. The research protocol commits to a process for sharing individual results with participants, which is the right commitment, but nothing describes how a result is confirmed before it reaches a person, who delivers it, what counselling accompanies it, or how a variant reclassified later reaches someone who was told years earlier that they were clear. That last is the distinctive oversight problem in population screening: the result is provisional in a way the recipient does not expect, and the obligation to revisit it has no natural owner.
The analytical activity is cohort building and agent driven exploration. Here the oversight question is verification rather than escalation. A researcher asking a question in natural language receives a cohort and a statistic, and whether they can see and check the definition that produced it determines whether the tool accelerates good research or produces confident errors faster. The aggregate only design with small count suppression is a privacy control and does not address analytical correctness.
Neither is described publicly, which is why this sits at C rather than lower or higher: the structures around both activities appear sound, and their internal design is simply not visible.
Ask who confirms and delivers a clinical finding, what happens on reclassification, and what a researcher is shown to verify a generated cohort.
The assay and the data architecture are described with reasonable specificity. The analytical layer this record indexes is not.
What is public: a proprietary sequencing assay extending beyond a standard exome, an accredited laboratory performing it, a dataset harmonised to a named observational research data model and refreshed quarterly, and a stated linkage to medical, pharmacy and mortality claims. Those are real technical facts a buyer can reason about, and the data model reference in particular tells an analyst what shape the data will be in before they see it.
What is absent concerns the tools added during 2026. A cohort building tool described as AI enabled and an agent query connector are both named as products, and neither is described technically. Nothing states what the cohort builder does that a query interface does not, what model underlies it, how a natural language request is translated into a cohort definition, or what happens when that translation is wrong. That last question is the important one: a cohort definition that quietly differs from what the researcher intended produces a valid looking analysis of the wrong population, and unlike a failed query it does not announce itself.
No versioning, evaluation or accuracy statement was located for either tool.
Ask what underlies the cohort builder, how a generated cohort definition is shown to the user for verification before analysis, and how the tools are versioned and validated.
Several genuine controls sit here and one technical design answers a question this index has been putting to other vendors without getting an answer. The controls: participation runs through an ethics committee approved protocol with informed consent rather than through terms of service, a published notice describes collection, use and disclosure, participants receive individual results and annual reports on what the programme found so contribution is not one directional, and recontact for further studies runs through a stated consent mechanism rather than being assumed.
Reciprocity is worth naming on its own, since almost every research corpus in this index takes from patients and returns nothing to them. The design worth naming is the query connector, which the company states returns only de identified aggregate statistics, gives no access to individual level records, automatically suppresses small counts, and is reachable only through provisioned accounts issued to approved partners.
Small count suppression is the control that matters, because the practical attack on an aggregate interface is narrowing a cohort until it contains one person, and building that in and saying so makes this the best designed data access surface assessed here.
What holds it below the top grade is the underlying dataset, where each record links an exome to well over a decade of clinical history and several years of claims including mortality, a combination among the hardest in existence to de identify, and no method is named. Ask which method applies and who assessed re identification risk.
Deployment evidence is strong and specifically named: population genomics programs at Cone Health, Sanford Health, WellSpan, St. Luke's, HealthPartners, Nebraska Medicine, and Memorial Hermann, with a reported 400,000+ participants sequenced across roughly 20 health systems. The company also states that network scale improves variant of uncertain significance resolution, which is a credible mechanism.
Held back from A because the headline outcome metric on the vendor's own site, estimated life years saved, is explicitly a modeled projection using a published cost effectiveness framework rather than measured outcomes, and the vendor labels it as such.
Several genuine controls, including one technical design that answers a question this index has been asking of other vendors.
The controls. Participation runs through an institutional review board approved protocol with informed consent rather than through terms of service. A published privacy notice describes collection, use and disclosure. Participants are given individual results and annual reports on what the programme found, so contribution is not one directional. Recontact for further studies is handled through a stated consent mechanism rather than assumed.
The design worth naming is the agent query connector released in 2026. The company states it returns only de identified aggregate statistics, provides no access to individual level records or identifiable information, automatically suppresses small counts, and is reachable only through provisioned accounts issued to approved partners. Small count suppression is the control that matters most there, because the practical attack on an aggregate interface is narrowing a cohort until it contains one person. Building that in, and saying so, is the best designed data access surface this index has assessed.
What holds this at B is the underlying dataset. Each record links an exome to roughly thirteen years of clinical history and several years of claims including mortality. That combination is among the hardest in existence to de identify, because the sequence alone is durably identifying and the longitudinal history narrows a person further. No de identification method is named.
Ask which method applies, who assessed re identification risk, and how consent quality is assured across around twenty recruiting institutions.
Correctly framed by the company as a clinical laboratory question and a research question at once, with the contracting detail unpublished.
The company describes itself as a population genomics company and a clinical laboratory, and operates an accredited sequencing laboratory. A laboratory receiving specimens on a physician's order and returning results is a health care provider transmitting health information electronically, which makes it a covered entity in its own right for that activity. It publishes a privacy notice covering collection, use and disclosure across its site, testing services and provider portal.
The research side runs on a different footing, and this is where the structure is distinctive. Participants are enrolled by health system members under an institutional review board approved protocol with informed consent, which means the recruiting institution obtains the consent and the platform aggregates the result. Research use of identifiable information at a covered entity is governed by that approval and the consent rather than by business associate terms, so the frameworks are correctly different for the two activities.
What is not published: an explicit statement of role, notice of privacy practices, the contracting entity, business associate availability for the health system relationships, or the subprocessor list. Two vendors assessed in this same lane publish all of that and sit a grade higher for it.
Ask which entity signs with a health system, what governs data flowing between the clinical and research activities, and how consent variation across around twenty recruiting sites is reconciled.
No SOC 2, HITRUST, ISO 27001 or equivalent information security attestation and no trust centre were located.
The laboratory accreditation distinction applies as elsewhere in this lane: certification and accreditation of the sequencing laboratory examine analytical validity, proficiency and specimen handling, not information security, and the estate at issue is far larger than the laboratory. It comprises a linked dataset of over five hundred thousand records combining sequence, longitudinal clinical history and claims, a research workspace served to external life sciences customers, delivery of data into customers' own environments, and an agent reachable query interface.
One technical control is published and deserves credit rather than being lost in the absence. The agent connector is stated to return only de identified aggregate statistics, to expose no individual level records, to suppress small counts automatically, and to require a provisioned account. That is a designed control addressing the specific risk of a flexible query interface, and it is the strongest such statement this index has recorded.
What it does not do is speak to the rest. Access control for the research workspace, authentication, logging, export restriction, staff access to identified sequence, and the development environment are all unaddressed.
Ask what independent security examination exists and what it covers, how workspace access is granted and monitored, what limits export from the workspace, and who inside the company can reach identified data.
A scoping determination that closes, with the framework that does apply named.
The sequencing assay is a laboratory developed test performed in the company's own accredited laboratory rather than a distributed device, so it operates under clinical laboratory regulation rather than device clearance, and no clearance is claimed. That is the ordinary and correct position for this product class. The regulator's attempt to bring such tests under device regulation through rulemaking was vacated by a federal court in 2025, so enforcement discretion continues and the laboratory framework remains the operative one. A buyer should understand that as a live area rather than a settled one.
What governs in practice is therefore laboratory certification and accreditation, covering analytical validity, proficiency testing and personnel qualification, alongside professional standards for variant classification and reporting.
One feature of this deployment sits outside all of that and deserves attention. Population screening returns actionable findings to people who presented with no symptoms, which creates downstream obligations: genetic counselling capacity, confirmatory testing, cascade testing for relatives, and a care pathway for a person newly told they carry risk. Those obligations fall on the health system rather than the platform, and the failure mode is a result delivered with nowhere for the patient to go.
Ask what the return of results pathway requires of the institution, what counselling is provided and by whom, and how findings for relatives are handled.
The research programme rests on an externally reviewed governance structure rather than a vendor written policy, which is rare in this index and is what earns the grade.
The network operates as a multi centre research programme under a formal protocol with single institutional review board oversight, registered publicly so the protocol, eligibility criteria and stated purposes can be read by anyone rather than described by the company. Participants are enrolled with informed consent obtained in accordance with applicable regulations and the reviewing board's requirements. The protocol commits to a process for sharing individual results with participants and to annual reporting back to them on study outcomes. Independent ethical review, a public protocol and a return of results commitment together constitute governance in the sense this axis means, not a statement of values.
Two gaps keep it below the top. Nothing published covers the newer analytical layer specifically, the cohort building tool and the agent query connector added during 2026, in the way the research programme is covered. And the company describes its population as diverse without publishing what that means or how performance varies across it, which matters because variant interpretation and polygenic risk estimation both degrade for populations under represented in reference data, and a population genomics programme is precisely where that inequity becomes operational.
Ask for the ancestry composition of the dataset, how interpretation performance varies across it, and what governs the analytical tools as distinct from the research protocol.
The assay and the data architecture are described with reasonable specificity and the analytical layer this axis indexes is not. What is public is genuinely useful to an analyst: a proprietary sequencing assay extending beyond a standard exome, an accredited laboratory performing it, a dataset harmonised to a named observational research data model and refreshed quarterly, and a stated linkage to medical, pharmacy and mortality claims.
Naming the data model in particular tells an analyst what shape the data will be in before they see it, which is a practical disclosure most vendors never make. What is absent concerns the tools added more recently. A cohort building tool described as artificial intelligence enabled and a query connector are both named as products and neither is described technically: nothing states what model underlies the cohort builder, how a natural language request is translated into a cohort definition, or what happens when that translation is wrong.
That last question is the important one, because a cohort definition that quietly differs from what the researcher intended produces a valid looking analysis of the wrong population, and unlike a failed query it does not announce itself. The output is a publishable result nobody has reason to doubt. No versioning, evaluation or accuracy statement was located for either tool, and no warranty, indemnity or remediation commitment. Ask how a generated cohort definition is shown to the user for verification before analysis runs.
A named standard, a stated refresh cadence and two distinct delivery models, which is more concrete than most records in this lane.
The clinical data underlying the research dataset is described as harmonised to a recognised common data model for observational research, refreshed quarterly, and enriched with medical, pharmacy and mortality claims. Naming that model is meaningful rather than decorative: it is the standard the observational research community actually uses, it means analyses are portable to other datasets built the same way, and it implies real extraction and mapping work against each contributing health system's record system rather than a bespoke export per site.
The operational side supports the same conclusion. Around twenty health systems run population genomics programmes on this infrastructure, covering recruitment, screening, return of results and research. A programme that returns clinically actionable findings into care has to reach the record where clinicians work, not merely produce a report.
Held at B rather than A because no interface documentation, named record platform integration or specification of how results are represented in the chart was located. For a screening programme that distinction is consequential: whether an actionable finding lands as a discrete coded result that decision support can act on years later, or as a document, determines whether the finding survives beyond the visit that produced it.
Ask how results are returned into each health system's record, in what form, and what happens when a variant is reclassified after the result was filed.
Two delivery models are described explicitly, which is the substantive part of this axis for a data platform, with the infrastructure detail unpublished.
The company offers managed access to the dataset in place, described as working without data transfers or infrastructure overhead, and separately offers delivery of data into a customer's own trusted research environment. Those are genuinely different propositions and naming both lets a buyer choose according to their own governance posture: the first keeps the data under the platform's controls and logging, the second puts it inside the customer's perimeter and makes them responsible for it. Most vendors in this position describe one and leave the other to negotiation.
The agent connector adds a third access path, deliberately constrained to aggregate statistics with small count suppression and provisioned accounts.
What is not published: hosting location, region, tenancy separation between the roughly twenty contributing health systems, retention schedules, and the subprocessor register. Tenancy deserves attention here because contributing institutions are also, in effect, contributors to a shared asset, and what one member can see of another's population is a governance question as much as a technical one.
Ask where the platform runs, how member institutions are separated, what a customer receives when data is delivered into their own environment and what obligations travel with it, and the retention position for sequence and linked records.
No public pricing. Contact the vendor. Two commercial surfaces: enterprise genomics programs contracted with health systems, and separate multi year data access agreements with life sciences organizations. No rate card published for either.
Clearly bounded and consistently described: population genomics programs for health systems, plus data access for payers and life sciences. The company does not claim clinical decision support or diagnostic interpretation beyond genomic screening and return of results.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Head to head
Vendors the index assesses as direct competitors to Helix for the same buyer.
Adjacent comparisons
Products a buyer researches alongside Helix that do a different job: a different category, a different layer of the stack, or a specialist scope. These pages exist to settle whether the comparison is real before it settles which one to pick.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Contact the vendor
|
Enterprise genomics program contracts; separate data access agreements | — | — | Vendor Published |
Two commercial surfaces: enterprise population genomics programs contracted with health systems, and separate multi year data access agreements with life sciences organizations such as those announced with Alnylam and AstraZeneca in 2026. No rate card published for either.