Ellipsis Health
Voice AI care management for health plans and at risk providers, originating in vocal biomarker research for detecting depression and anxiety and now productized as Sage, an AI care manager. The distinguishing technical claim is the Empathy Engine, patented vocal biomarker technology trained on millions of live clinical patient calls, which reads emotional state from voice in real time rather than only transcribing words, and adapts tone and approach accordingly.
Sage conducts inbound and outbound member calls covering health risk assessments, post discharge follow up, risk profiling, care gap closure, program enrollment, medication reminders, and triage, and is positioned explicitly as added staffing capacity rather than a chatbot, reserving clinicians for complex high touch interventions. Integrations include Salesforce Health Cloud and, through an April 2026 partnership with HealthEdge, the Care Solutions suite drawing on GuidingCare member data and returning outcomes to the platform.
Company reported figures, which it acknowledges are not independently audited, include roughly 60 percent reduction in administrative tasks, 6x faster program enrollment, and 4x return on investment. Named customers per the company website include UnitedHealth Group, Aetna, Duke Health, and Highmark. Founded 2017 by CEO Mainul Mondal; approximately $75 million raised including a $45 million round in June 2025 co-led by Salesforce, Khosla Ventures, and CVS Health Ventures.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
Vocal biomarker analysis is the product and the company's origin: the technology began as research into detecting depression and anxiety from voice, and Sage applies it during live calls to read emotional state and adapt. That is inference on the acoustic signal rather than on the transcript, which is a materially different capability from the other voice agents in this index.
Positioning is consistent and correct in shape: added staffing capacity rather than a chatbot, with the explicit purpose of reserving skilled clinicians for complex, high touch interventions while the agent handles high volume outreach.
Held back from A because escalation criteria are not published, and the gap matters given what the system detects: if vocal analysis surfaces acute distress or crisis signal during a routine care gap call, the handoff protocol is the single most important control and it is undisclosed.
The training basis is stated specifically: a patented Empathy Engine trained on millions of live clinical patient calls, with vocal biomarker work grounded in published research on detecting mental health conditions from voice. Held back from A because no accuracy, sensitivity, or specificity figures for the emotional or mental health inference were retrieved, and this is precisely the axis where a vocal biomarker claim needs numbers. Detecting depression from voice is a strong claim and should be evidenced as one.
The corpus is described and no party is named. The vendor states its engine was trained on millions of live clinical patient calls, with the vocal biomarker work grounded in published research on detecting mental health conditions from voice, so a buyer knows what kind of material the model learned from and can read the underlying literature.
That is more than most vendors disclose about any input, and it makes provenance the question rather than an accusation: millions of live clinical calls is a corpus of recorded patients, and whose calls those were, on what consent basis, and whether a current customer's calls join it are all unanswered.
Voice raises this above the ordinary training question, because a recording carries speaker identity independently of what is said, so a voice corpus is identifying in a way a text corpus is not, and de identification of transcripts does nothing about it.
On enumeration there is nothing: no model or model family, no foundation model provider, no hosting arrangement and no sub processor list was located, and no statement covers retention of the recordings themselves or of the inferred signals derived from them. Ask whose calls formed the corpus and on what basis, whether customer calls join it, and where the inferred mental health signals are stored and who receives them.
Distribution and investor signals are strong: a strategic HealthEdge partnership embedding Sage into Care Solutions with GuidingCare member data, Salesforce Health Cloud integration, and CVS Health Ventures participating in the funding round. Named customers per the company website include UnitedHealth Group, Aetna, Duke Health, and Highmark.
Held back because the operational figures, roughly 60 percent administrative task reduction, 6x faster enrollment, and 4x return on investment, are acknowledged by third party review as company reported and not independently audited, and no clinical validation of the biomarker capability in deployment was retrieved.
The most sensitive inference surface among the voice agents in this index, and disclosure does not match it. The system infers emotional and mental health state from vocal characteristics during routine care management calls, which means it may derive behavioural health signal from members who called about something else and never consented to a mental health assessment.
No disclosure was retrieved covering member notification, consent for biomarker analysis, retention of voice recordings, or how inferred mental health signals are stored and shared with the plan. When those signals return to a health plan alongside member data, the stakes are higher than for transcription.
Compliance with the health privacy rule is stated on the company's own materials, and customers are health plans and at risk providers, so business associate agreements exist as a matter of course. No availability statement, scope description, contracting entity or subprocessor list was located.
A second framework applies here that most vendors in this index never encounter, and it is stricter than the first. The product analyses the voice itself rather than only the words. A voiceprint, meaning a measurement of vocal characteristics used to identify or characterise a person, is a biometric identifier under several state biometric privacy statutes. Those statutes generally require informed written consent before collection, disclosure of the retention schedule and destruction policy, and prohibit sale. One of them carries a private right of action with statutory damages assessed per violation, which has produced substantial class litigation against companies that assumed their existing consent covered it.
The critical point is that the two frameworks do different work. The health privacy rule permits use for treatment and health care operations without specific consent, so a health plan may lawfully call a member without one. That permission does not extend to collecting a biometric identifier, where the state statute asks its own question and expects its own consent. A programme that is entirely compliant on the health privacy side can still be exposed on the biometric side.
Ask what consent is obtained before voice analysis, how it is captured on an inbound or outbound call, what the retention and destruction schedule for voice data is, and how the position varies by state.
A real and independently examined posture, described with more care than most.
What is stated: a SOC 2 Type II, which the company explains correctly as demonstrating that controls were operationally effective over a period rather than present at a point in time; compliance with the health privacy rule; alignment with European data protection law; and regular third party security audits. It also publishes a plain language explanation of what each framework is and does, which is unusual and genuinely useful for a buyer without a security function of their own.
One status needs checking rather than accepting. Across materials published roughly a year apart, certification under the healthcare control framework is described as in process. That may simply reflect undated pages, and the certification may since have completed. It may also mean the process has been open for some time. Either way, in process is not certified, and a buyer should ask for the current position rather than reading the phrase as an accomplishment.
What is missing is a trust centre or documented request path, a scope statement covering which systems and entities fall inside the attestation, and penetration testing disclosure.
One scope question is specific to this product. The company holds recordings and derived acoustic features from millions of clinical calls. Whether that corpus and the model training environment sit inside the audited boundary, or only the production service does, is worth establishing directly.
No clearance was located, no regulatory rationale is published, and the company's own descriptions of the technology span both sides of the boundary. That gap is the finding.
On one side, the operating product is positioned as care management: outreach, assessments, enrolment, follow up and coordination, delivered as staffing capacity for health plans, with clinicians reserved for complex intervention. Framed that way it is administrative and no device pathway attaches.
On the other, the company's own research materials describe vocal biomarkers as accelerating the diagnosis, treatment and monitoring of mental health conditions, and trade coverage describes the technology as identifying disorders such as anxiety and depression. Software that identifies a mental health condition, or measures its severity, is doing what a diagnostic aid does. Independent commentary describes the outputs as advisory and positions the technology as a supplement to standard screening instruments, which is the framing that keeps it outside device regulation, but that framing appears in third party analysis rather than in a published position from the company.
This is not an allegation of non compliance. The likely reality is that the marketing language predates the current product and has not been reconciled with it. But a buyer cannot tell from the outside which claim governs, and the distinction determines whether what they are deploying is a screening instrument requiring clinical governance or an outreach tool requiring operational governance.
Ask for the regulatory position in writing, the formal intended use statement, whether any component has been assessed against device criteria, and what the product must not be used to conclude.
Two real governance controls are published, and the disclosure stops short of the results a buyer in this domain most needs.
What exists. The company states it operates under a written artificial intelligence ethics policy, and that every update to the product is reviewed by its own clinical team of physicians, nurses and care managers before deployment. A named clinical review gate on every release is a specific, checkable control rather than a principle, and few vendors in this index publish one. The underlying vocal biomarker technology is backed by more than ten peer reviewed publications, which is a stronger scientific base than almost anything else in this lane.
What is absent is subgroup performance, and the reason it matters here is acoustic rather than abstract. Models that infer state from voice are sensitive to accent, dialect, language, age, vocal pathology and the audio channel itself, since telephony compresses and degrades exactly the signal the model reads. Independent commentary on this technology names those as the primary risks. The company acknowledges the direction of travel, describing work to enrich datasets and make algorithms more robust and inclusive, which is candid and is not a measurement.
The consequence is asymmetric. A member whose voice is well characterised gets accurate prioritisation; one whose voice is not may be scored as lower need and simply receive less outreach, and nothing in the workflow reveals it.
Ask for performance by accent, primary language, age and audio channel, the calibration approach for a new population, and what the ethics policy actually commits to.
Two passes located no accuracy, sensitivity or specificity figure for the emotional or mental health inference, no published limitations and no warranty, indemnity or remediation commitment, and this is the record in the backfill where that absence is hardest to justify. The system infers mental health state from vocal characteristics during routine care management calls.
Detecting depression from voice is a strong clinical claim and needs to be evidenced as one, because the failure modes are not symmetrical and both land on a person who did not ask for the assessment. A false positive attaches a mental health signal to a member who called about a prescription refill. A false negative in a product marketed as surfacing risk means a plan believes it has screened someone it has not. The consent position compounds it rather than mitigating it.
The member called about something else, may not know that vocal biomarker analysis is running, and has not consented to a mental health assessment, so an inference about their psychological state is generated without their knowledge and then returns to their health plan alongside their other data.
That is the harmed party structure this axis tracks in one of its sharpest forms: the person assessed has no relationship with the vendor, no notice, and no route to see or contest what was inferred. Ask for sensitivity and specificity by population, what members are told, and whether an inferred signal can be challenged or removed.
The right integrations for a payer facing product, named specifically and in one case bidirectional.
This is not an electronic health record story and should not be graded as one. The buyer is a health plan or an at risk provider organisation, and the systems that matter are care management platforms rather than clinical record systems. The company names two: a major customer relationship platform's health cloud, where it maintains a listing in that vendor's marketplace, and through a 2026 partnership, a care management suite where it draws on member data and returns outcomes to the platform.
That second arrangement is the more substantive one, because the return path is what makes an outreach tool useful rather than merely busy. An assessment completed by voice is only valuable if its result lands in the system the care manager works in, attached to the right member, in a form the next workflow can act on. Naming a partnership that includes both directions is a real capability claim.
Held at B rather than A because no interface documentation, supported data specification or list of production integrations beyond those two was located, and because marketplace listing is a lower bar than deployed integration. There is also no evidence of connectivity to clinical record systems, which matters if a plan wants outreach outcomes visible to treating clinicians rather than only to its own care team.
Ask which integrations are live in production, what is written back and where it lands, and whether anything reaches the member's treating providers.
Delivered as a cloud service, with no hosting location, region, tenancy model, retention schedule or subprocessor list located.
Three questions are specific to this product rather than generic.
Retention of voice is the first and most important. The company describes technology trained on millions of live clinical patient calls, which means recordings and derived acoustic features are held at scale. How long raw audio is kept, whether derived features persist after audio is deleted, and whether a member can have their voice data destroyed are questions with legal weight as well as operational weight, because state biometric statutes impose retention and destruction requirements independently of health privacy law.
The second is the telephony layer. Inbound and outbound calling necessarily involves carriers and voice infrastructure providers, so audio passes through parties the company does not own. That layer belongs in the residency answer rather than being treated as plumbing.
The third is training separation. Where a company improves models from the calls it conducts, the boundary between one health plan's member conversations and the models served to another plan is a commercial question as much as a privacy one, and several of this company's named customers compete directly with one another.
Ask where audio and features are stored and for how long, what the destruction process is, which carriers are in the path, and whether one customer's calls contribute to models serving another.
No public pricing. Contact the vendor. Sold to health plans, at risk providers, and care management organizations, increasingly through platform partners including HealthEdge and Salesforce, which means some buyers will encounter it as a component of a larger contract rather than a standalone purchase. Worth establishing whether the biomarker capability is separately priced or bundled.
Clearly bounded by customer and workflow, with an important asymmetry between where the evidence sits and where the product now operates.
The customer boundary is explicit: health plans and at risk providers, not consumers and not clinical settings. The workflow boundary is equally clear, covering health risk assessments, post discharge follow up, risk profiling, care gap closure, programme enrolment, medication reminders and triage. The company is also explicit about what it is not, positioning the product as added staffing capacity rather than a chatbot and reserving clinicians for complex intervention. That kind of stated ceiling is creditable.
The asymmetry concerns evidence. The scientific foundation, more than ten peer reviewed publications, concerns vocal biomarkers for depression and anxiety. The product now spans behavioural, physical and social needs across a wide range of care management tasks. Evidence that a voice signal carries information about depression does not establish that it carries information relevant to medication adherence, chronic disease management or social needs, and the broader the application, the further it travels from what was validated.
That is not a criticism of the expansion, which is a reasonable commercial path. It means a buyer should establish which capabilities rest on the published research and which rest on general conversational competence, because those warrant different levels of confidence and different oversight.
Ask which use cases the biomarker informs and which it does not, and what evidence supports the non behavioural applications.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Contact the vendor
|
Health plan and care management agreements; also sold via platform partners | — | — | Third Party Estimated |
No rate card published. Sold to health plans, at risk providers, and care management organizations. Increasingly distributed through platform partnerships including HealthEdge Care Solutions and Salesforce Health Cloud, so some buyers will meet it embedded in a larger platform contract rather than as a direct purchase.
Reported returns of roughly 4x and 60 percent administrative task reduction are company stated and, per third party review, not independently audited, so they should be treated as a starting point for diligence rather than a basis for a business case.