Canary Speech
Canary Speech sells the model rather than the application. Its vocal biomarker engine extracts roughly 2,590 acoustic and linguistic features from the human voice every 10 milliseconds and returns behavioral and cognitive indicators from as little as 20 to 45 seconds of natural speech, delivered through an API that other products embed. Canary Ambient is the API first real time offering, running on patient and clinician conversations, with a continuous monitoring variant added subsequently. Canary Cognitive addresses the cognitive side. Because the analysis reads how someone speaks rather than what they say, the company describes the approach as language agnostic and device agnostic, workable on anything with a microphone.
The published evidence is unusual in this lane for what it admits. A 2026 paper in the Proceedings of Artificial Intelligence in Medicine reports unweighted average recall of 0.70 for depression and 0.68 for anxiety, with a combined behavioral health assessment reaching sensitivity of 0.76 and specificity of 0.65 on a remote test set, and sensitivity of 0.68 with specificity of 0.80 on an independent in clinic dataset collected by tablet. Those are moderate numbers and the company published them anyway, at an operating point, which almost no competitor here does. The same work introduces an uncertain classification that flags low confidence predictions near the decision boundary rather than forcing a call, which is a deliberate clinical safety choice rather than a performance figure.
Markets span health systems, payers, pharmaceutical and clinical trial work, employers, telehealth and clinical call centres. In February 2026 the company partnered with Intermountain Ventures on an institutional review board approved study led by a neurologist at Intermountain Health, testing whether vocal features can identify multiple sclerosis. That same month it entered consumer health through JubileeTV, embedding passive speech analysis into video calls between older adults and their families, its first deployment outside clinical and research settings, producing indicators the company is careful to describe as nondiagnostic.
Founded in 2017 in Provo, Utah by Henry O'Connell, who began his career at the National Institutes of Health in a neurological disease group before a long medical device career, and Jeff Adams, a speech recognition specialist. Roughly 22.35 million dollars raised in total, including a 13 million dollar Series A in June 2024 led by Cortes Capital and a further round in January 2025. Investors include Sorenson Communications, SMK Corporation, Plug and Play Japan and Hackensack Meridian Health, the last of which is a health system and therefore a party whose commercial and investment interests a buyer should hold separately when weighing any reference from it.
One disclosure detail deserves attention before a security review. The company's security page lists an extensive set of assurance programs covering the SOC family, FedRAMP, FISMA and four ISO standards, and attributes them accurately to the cloud partners it builds on rather than claiming them. Canary's own stated position is HIPAA compliance, control audit against a recognised benchmark, vulnerability scanning and cleared API penetration testing. No attestation held by Canary itself was located. The wording is honest; the visual effect of the list is not, and a buyer skimming it may credit the company with certifications it does not hold.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The purest case in this lane. Canary Speech sells the model itself through an API, and there is no application wrapped around it to carry independent value. No portal, no scheduling, no documentation suite, no patient experience layer. A customer buys inference and builds the product around it.
The analytical work is substantial rather than nominal. The engine extracts roughly 2,590 acoustic and linguistic features every 10 milliseconds and returns behavioral and cognitive indicators from 20 to 45 seconds of natural speech. Reading how something is said rather than what is said is the entire basis of the offering, which is why the company can claim device independence and language independence at all.
The API first architecture makes centrality unusually easy to verify here. Where a platform vendor's centrality must be argued by separating intelligence from workflow, this vendor has already performed that separation commercially: the intelligence is the unit of sale. The corresponding exposure is that every risk in this record concentrates in the models, since nothing else exists to absorb it.
The best oversight posture recorded in this lane, resting on three things that reinforce one another rather than on a statement of intent.
First, the operating point is published. Sensitivity and specificity are given at the deployed decision threshold, so a clinician or an integrator knows the approximate miss rate before deployment rather than after an incident. Every other multimodal assessment record here publishes either nothing or a threshold independent statistic, and this index has written into two of them that only a published operating point would move their liability grade. This vendor meets that condition.
Second, and rarer, the system can decline to answer. The published work introduces an uncertain classification that flags low confidence predictions near the decision boundary instead of forcing a binary call. Building abstention into a screening model is a deliberate safety choice with a commercial cost, since it reduces the proportion of inputs the product appears to resolve. It routes genuinely ambiguous cases to human judgement rather than manufacturing a confident answer, which is the behaviour this axis is meant to identify.
Third, scope language is disciplined. The offering is described as clinical decision support, and the consumer product's outputs are called nondiagnostic indicators rather than assessments. Correctly narrowing a claim in the market where the incentive to inflate it is greatest is a good sign.
What is missing keeps this from being complete. No escalation protocol is published for a high risk indication, no notification path, and nothing states what an integrator is required to have in place before embedding a behavioral risk signal in a consumer product. Because the model ships inside someone else's application, the vendor also cannot fully control how its output is presented.
Ask what integrators are contractually required to do with a high risk result, and what proportion of inputs return uncertain in production.
The most useful model disclosure in this lane, because it publishes the number a deployment decision actually turns on.
Most vendors here publish either nothing or a threshold independent statistic. This one publishes the operating point. Sensitivity of 0.76 and specificity of 0.65 on a remote test set, and sensitivity of 0.68 with specificity of 0.80 on an independent in clinic dataset, with unweighted average recall of 0.70 for depression and 0.68 for anxiety. Those figures tell a buyer roughly how many cases are missed and how many people are flagged unnecessarily, which is what determines whether a screening programme is viable at a given volume.
They are also unflattering, and that is the point. A specificity of 0.65 is a number a marketing department would prefer to omit, and reporting performance on a second independently collected dataset invites exactly the comparison a vendor is tempted to avoid. Disclosure that costs something is the disclosure worth grading on.
The method is described concretely too. Roughly 2,590 features extracted every 10 milliseconds, ensemble models, and an explicitly described uncertain class for predictions near the decision boundary. A patent covering neural network based speech analysis was granted in 2024, and a peer reviewed paper documents the approach.
Three gaps remain. No model card exists. Architecture beyond the ensemble framing is undescribed, and whether any third party speech recognition component sits in the pipeline is unstated. And the published figures cover depression and anxiety only, so the cognitive and neurological claims carry no numbers at all.
Ask for a model card, calibration data, and published performance for the cognitive claims.
The infrastructure tier is characterised, the model layer is proprietary and plausibly self contained, and no component is actually named.
What can be established. Cloud partners are described by their assurance programs rather than by name, which narrows the provider without identifying it. The models themselves appear to be built in house rather than assembled: a patent covering neural network based speech analysis was granted in 2024, feature extraction is described at a level of detail consistent with proprietary work, and one co founder is a speech recognition specialist by background, which makes an internally built speech pipeline plausible in a way it would not be for most vendors in this lane.
Plausible is not disclosed. No cloud provider, no base or foundation model, no speech recognition component, no sub processor register and no third party library position was located. Whether any external service ever receives customer audio is unstated, and for a product whose input is recorded human voice that is the first question a privacy reviewer asks.
Training provenance is described only as clinically labelled datasets. Where those recordings came from, under what consent, and whether audio captured through customer or partner deployments feeds further development are all unaddressed. The consumer partnership makes this more pressing, since conversations between older adults and their families are an unusually rich source of natural speech.
Graded C rather than lower because de identification in the analysis environment and anonymous configurations are described, which is a partial answer about data flow, and because the assurance list characterises the infrastructure tier.
Ask for the sub processor register, whether any third party performs speech recognition, and the consent basis of the training corpus.
Honest published performance, vendor authored, with commercial outcome evidence still missing.
The 2026 paper in the Proceedings of Artificial Intelligence in Medicine reports unweighted average recall of 0.70 for depression and 0.68 for anxiety, and a combined assessment reaching sensitivity of 0.76 with specificity of 0.65 on a remote test set. Performance was then examined on an independent in clinic dataset collected by tablet, where sensitivity was 0.68 and specificity 0.80. Testing on a second dataset gathered by a different method and in a different setting is a meaningful robustness check rather than a restatement of training performance.
Those are moderate numbers. A specificity of 0.65 means roughly a third of unaffected people are flagged, which matters for a screening tool deployed at volume. Publishing them anyway, rather than asserting equivalence without figures, is the behaviour this index exists to reward, and it is the reason this record is graded above competitors making larger claims on thinner disclosure.
Authorship is the qualification. All five authors are company staff, including the chief executive. Peer review is real and vendor authored research is normal in this field, but no independent replication exists.
Two further gaps. Deployment evidence is absent: no customer count, no facility count and no outcome from any named site was located, so nothing shows that using the technology changes care. And the validated scope is depression and anxiety only. Dementia, Alzheimer's disease and multiple sclerosis are all marketed or under study, with the multiple sclerosis work an institutional review board approved study with Intermountain Health launched in February 2026 rather than a completed result.
One relationship to hold separately: Hackensack Meridian Health is a health system and an investor, so any reference originating there carries a commercial interest.
Ask for a deployment outcome at a named customer, and for published figures on the cognitive and neurological claims.
Real controls described at the right level of specificity, with one category of risk left entirely unaddressed.
The substance is there. Anonymous configurations are offered, API research data is de identified inside the analysis environment, the development lifecycle and continuous infrastructure monitoring are described, and three concrete assurance activities are named: control audit against a recognised benchmark, vulnerability scanning, and cleared API penetration testing. Naming penetration testing of the API specifically is apt, since the API is the entire attack surface of this product.
The unaddressed category is biometric law. A voiceprint is biometric data. Several jurisdictions regulate the collection and retention of biometric identifiers under statutes distinct from health privacy law, with their own consent, notice and retention requirements and, in some cases, a private right of action. A vendor whose product extracts thousands of features from a person's voice sits squarely inside that regime, and nothing published engages it in any form. This is the most consequential legal gap in the record and it is specific to this vendor rather than generic.
The consumer product sharpens it. Passive analysis of an older adult during a family video call raises the question of whose consent was obtained and whether the person being analysed knows, and no published material describes the consent flow.
The training position is also unstated. Models are described as trained on clinically labelled datasets, with nothing saying whether audio captured through customer deployments contributes to further development.
Ask about biometric statute compliance and retention, the consent design in the consumer product, and whether customer audio trains models.
A compliance claim stated plainly, with architecture that genuinely reduces the exposure and no contracting position behind either.
The architectural point is worth crediting because it is a design decision rather than a policy sentence. The company states that fully anonymous solutions are offered and that research data collected through the API is de identified in the analysis environment. For an API embedded in someone else's product, that arrangement can mean the vendor never holds identity at all, which lowers the stakes of the compliance question rather than answering it more loudly.
What is asserted is HIPAA compliance, described as full. HIPAA has no certifying body, so this is the company's own assessment of itself, and no external verification of the health privacy position specifically was located.
Business associate agreement handling is not addressed. No template, no execution requirement, no statement of which entity contracts, and no description of how the agreement chain works when the API sits inside a partner's product and the patient relationship belongs to that partner rather than to Canary.
That last point is the distinctive gap here. In an embedded model the flow of responsibility runs vendor to integrator to provider, and nothing published describes who signs what at each hop, or whether Canary contracts directly with covered entities at all.
Ask for the agreement template, whether Canary signs as a business associate or a subcontractor, and which deployments run in the anonymous configuration.
A list of certifications appears prominently, correctly attributed, and none of them belongs to this company.
The security page enumerates assurance programs spanning SOC 1 and ISAE 3402, SOC 2, SOC 3, FISMA, DIACAP, FedRAMP, and ISO 9001, 27001, 27017 and 27018. The sentence introducing that list states that these are programs the company's cloud partners adhere to. That attribution is accurate and the company deserves credit for wording it correctly rather than claiming the credentials outright, which several vendors in this index do.
The effect nonetheless requires flagging, because a buyer skimming a security page sees ten certifications and forms an impression. Inherited infrastructure assurance is not the same assurance as an attestation covering the vendor's own controls, its own staff, its own development practice and its own handling of customer data. A hyperscaler holding FedRAMP says nothing about how the application built on it behaves.
What Canary states about itself is narrower and more specific than a badge: HIPAA compliance, control audit against a recognised security benchmark, vulnerability scanning, and cleared penetration testing of the API. Naming penetration testing of the API is well judged given that the API is the whole product surface. These are real activities and they lift this above the vendors whose entire posture is a logo.
What is missing is independent attestation of any kind held by the company. No trust center, no report available under agreement, no assessor named, no audit period, no scope statement and no vulnerability disclosure policy was located. Given that leadership has publicly written about the distinction between the two SOC report types and about ISO 42001, the absence appears to be a stage of maturity rather than an oversight.
Ask whether an attestation is in progress, what documentation can be shared under agreement, and who performs the penetration testing.
Careful scope language throughout and no stated determination behind it.
The care is evident and worth recording, because it contrasts with the lane. Outputs in the consumer product are described as nondiagnostic indicators. The clinical offering is described as clinical decision support providing actionable insight to a clinician. Screening rather than diagnosis is the consistent framing. No registration is presented as a credential, no designation is described as a clearance, and no approval language appears anywhere. Three separate records in this lane carry exactly those overstatements, and this vendor carries none of them.
What is absent is the determination itself. No device classification, no rationale, no clearance, and no published statement that the products sit outside device regulation was located.
The question is live at the edges of the marketed range rather than at its centre. Screening for depression and anxiety with a clinician deciding is comfortably within decision support territory. Identifying dementia or Alzheimer's disease from voice moves closer to detecting disease, and the multiple sclerosis study, being institutional review board approved and aimed at identifying individuals with a neurological condition from vocal features, is investigating a capability that would be difficult to characterise as anything but diagnostic if it succeeded and shipped.
The consumer deployment raises a second question. Wellness positioning is a recognised route to remaining outside device regulation, and the boundary is tested when the same underlying models that carry clinical claims are used to generate indicators for families.
Ask for the written device determination, whether it covers the cognitive and neurological claims, and what regulatory path the multiple sclerosis work is on.
Strong performance disclosure with no fairness disclosure attached to it, against a modality where the bias mechanism is well established.
What exists is aggregate. Performance is published across two datasets, and the abstention category is a safety mechanism rather than a fairness one. No subgroup breakdown by sex, age, race, accent or primary language was located, no calibration analysis, no fairness testing, no external audit and no responsible artificial intelligence certification of the kind the company's own leadership has publicly discussed as an emerging standard.
The risk is specific to voice and it is documented. Speech models carry differential error by sex, by age, by accent and by dialect, and automatic speech processing has a long record of degraded performance for speakers outside the dominant training distribution. This lane also holds a direct precedent: a competitor that shut down in early 2025 faced published concern that its models had been trained largely on the speech of a majority white population and might not perform equivalently for young people of colour.
The language agnostic claim is where the gap bites hardest, because it is a fairness claim presented as a technical property. Asserting that a model works across languages, without publishing per language performance, tells a buyer in a non English market to assume equivalence that has not been shown.
The consumer product adds an age dimension. Older adults' speech is altered by hearing loss, dentition, medication and comorbidity, all of which are plausible confounds for the features being read.
Graded C rather than lower because published aggregate performance at an operating point is itself a governance artefact and outranks unmeasured principle statements. Ask for sensitivity and specificity broken down by sex, age band, race and primary language.
No contractual allocation exists, and this is the one vendor in the lane that supplies the number such an allocation would have to be built on.
The missing half is conventional. A dedicated pass located no service level agreement, no accuracy warranty, no performance guarantee, no indemnity and no remediation commitment. Terms governing the API were not reachable without contact.
The half that is present is why this sits above the records graded lower. This index wrote into two competitor records that only a contractual term or a published operating point with a stated false negative rate could move their liability grade. This vendor publishes the operating point. Sensitivity of 0.76 on the remote test set means roughly one in four cases is missed at the deployed threshold, and the company put that figure in a peer reviewed paper rather than leaving a buyer to discover it. A miss rate that is known can be planned around, insured against and written into a contract. A miss rate that is unpublished cannot.
The abstention class reinforces the same point by limiting confidently wrong outputs near the decision boundary.
What remains unallocated is consequence, and the embedded model complicates it beyond the usual. When the analysis ships inside a partner's product, responsibility for a missed indication is divided between a model with a published error rate, an integrator who chose how to present it, and a clinician or family member who acted or did not. Nothing published describes how that division is handled contractually, and the consumer deployment puts a behavioral signal in front of family members with no professional standing at all.
Ask what the integration agreement says about responsibility for a missed indication, and whether consumer embedding carries different terms from clinical use.
Integration capability is the architecture rather than a feature list, and not one clinical record system is named.
What is genuine is that this product is built to be embedded. An API first design with documented integration capability is a real interoperability posture, and one named embedding exists in production: the consumer caregiving partnership announced in February 2026, which places the analysis inside another company's video calling product. Telehealth organisations, health systems and clinical call centres are described as an existing partnership network. A vendor whose commercial model depends on other people integrating it has stronger incentives here than one selling a portal.
What is absent is everything specific to clinical systems. No electronic health record is named, no interface standard such as FHIR or HL7 version 2 is described, no marketplace or partner listing was located, and no public API documentation was reachable without contact.
The consequential question for a screening product is where the result lands. A behavioral or cognitive indicator that reaches a clinician inside the chart as a discrete, trendable result is a different clinical object from one displayed in a partner application, and nothing describes which occurs. For a product whose value proposition rests on detecting change earlier than scheduled screening, longitudinal storage in the record is close to essential.
Graded C rather than lower because an API first architecture with a live production embedding is demonstrated integration capability, and rather than higher because no clinical system is named and the standard is unstated.
Ask which record systems are integrated in production today, through what standard, and whether results write to the chart as discrete values.
The hosting tier can be characterised from what is published and the hosting itself cannot.
What the security material conveys indirectly is useful. Cloud partners are described as adhering to assurance programs spanning the SOC family, FedRAMP, FISMA and several ISO standards, and environments are described as routinely audited with certifications across geographies. That set narrows the provider to a major hyperscaler and tells a buyer the underlying infrastructure sits at the tier where government and regulated workloads run.
Everything specific is missing. The provider is not named. No region is stated, no residency commitment is made, no tenancy model is described, no segregation approach is given for an API serving many integrators, and no customer controlled or private deployment option is mentioned.
Residency matters more here than the omission alone suggests. The company positions on language independence, holds an investor based in Japan, and describes certifications spanning geographies, all of which point at deployment outside the United States. European and other regimes impose location and transfer requirements on health and biometric data that a global voice analysis service cannot avoid, and nothing published addresses where audio is processed or where derived features are stored.
A related question follows from the architecture. Whether audio leaves the integrator's environment at all, or whether feature extraction can occur locally with only derived values transmitted, materially changes the residency analysis, and the answer is not published.
Ask which provider, which regions, whether audio or only derived features cross the boundary, and what is offered for deployments outside the United States.
Cost is absent from every published surface. A dedicated pass located no pricing page, no unit of charge, no range, no tiering, no implementation fee position, no minimum commitment and no trial terms. Every route terminates in a demo request.
The one commercial adjective published is that the API is cost effective, offered without a figure, a comparison or a basis. That is a claim a buyer cannot act on.
An API first product makes the omission more surprising rather than less. Usage priced infrastructure is the category most able to publish a rate, because the unit is obvious and inspectable: a call, a minute of audio, a completed assessment. Competing developer facing services routinely publish exactly that. Nothing indicates whether Canary charges per analysis, per minute of speech, per monitored individual, per seat or by annual licence, and an integration partner has to model cost of goods before committing engineering effort to embed a third party model.
Segment spread compounds it. A health system, a payer, a pharmaceutical sponsor running a trial, an employer, a clinical call centre and a consumer caregiving product cannot plausibly share one rate, and nothing describes how they differ.
Ask for the unit of charge, the volume tiers, whether pharmaceutical and consumer embedding price differently from clinical use, and the minimum commitment.
Genuinely broad reach for a company of this size, with a validated core considerably narrower than the marketed range.
The breadth is real and structurally enabled. Because the product is an API rather than an application, it reaches settings an application could not: health systems, payers, pharmaceutical and clinical trial work, employers, telehealth services and clinical call centres. February 2026 added consumer caregiving through JubileeTV, embedding passive analysis into video calls between older adults and their families, which the company describes as its first deployment outside clinical and research environments. Covering the clinic, the trial, the call centre and the living room from one model is unusual.
Condition coverage is where the claim outruns the evidence. Depression and anxiety carry published performance. Dementia and Alzheimer's disease are longstanding marketing claims with no figures located. Multiple sclerosis is an active study rather than a product. A buyer reading the company description would reasonably conclude all four are equally supported, and they are not.
The language agnostic claim needs the same treatment. Analysing acoustic form rather than lexical content is a sound reason to expect cross language transfer, and expectation is not evidence. No per language performance was located, and the investor base and partnerships point at international deployment.
Age coverage is also unaddressed at both ends. The consumer product targets older adults, where speech is affected by hearing loss, medication and comorbidity, and nothing published separates those effects from the signal being measured.
Ask which conditions have published performance, in which languages, and how age related speech change is handled.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published
|
Not disclosed. No unit of charge is described anywhere. Whether pricing runs per analysis, per minute of speech, per monitored individual, per seat, by annual licence or by usage tier is unstated, and the interface first architecture means the natural unit is a call or a completed assessment rather than a user, so the basis cannot be inferred from the product shape as it sometimes can. Nothing indicates how the continuous monitoring offering is charged relative to episodic assessment, nor whether the cognitive and behavioral models are licensed separately. | Not disclosed as a template or posture, and complicated by the embedded model. HIPAA compliance is asserted as full, with anonymous configurations offered and application programming interface research data described as de identified in the analysis environment, which can mean the vendor holds no identity at all in some deployments. No business associate agreement template, negotiation stance or execution requirement was located. The distinctive question is the contracting chain: when the analysis runs inside a partner's product and the patient relationship belongs to that partner, whether Canary signs as a business associate directly with covered entities or as a subcontractor to the integrator is unstated, and the answer determines who carries breach notification duties. Note separately that a voiceprint is biometric data regulated in some jurisdictions under statutes distinct from health privacy law, with their own consent and retention requirements, and nothing published engages that regime. Ask which deployments run anonymous, where Canary sits in the agreement chain, and how biometric statutes are handled. | Not disclosed. No implementation, integration or onboarding fee position was located and no deployment timeline is published. The likely effort can be characterised from the architecture even though it is never costed: an interface first product shifts integration work onto the customer or partner, who builds the surrounding application, handles audio capture and decides how results are presented. Whether Canary charges for integration support, certification of a partner implementation, or model tuning against a customer population is unstated, as is whether the institutional review board approved study work with a health system partner represents a funded engagement or a research collaboration. | Vendor Published |
Cost is absent from every published surface. A dedicated pass located no pricing page, no unit of charge, no range, no volume tiering, no implementation fee position, no minimum commitment and no trial terms. Every commercial route terminates in a demo request.
The single commercial adjective published is that the offering is cost effective, given without a figure, a comparison or a basis, which is a claim a buyer cannot act on or verify.
An interface first product makes this omission more surprising rather than less. Usage priced infrastructure is the category best placed to publish a rate, because the unit of consumption is obvious and inspectable: a call, a minute of audio, a completed assessment. Developer facing services in adjacent categories publish exactly that, and a published rate is close to a precondition for the audience this company sells to. An integration partner has to model cost of goods sold before committing engineering effort to embed a third party model in its own product, and that calculation cannot begin without a unit price.
Segment spread compounds it. A health system, a payer, a pharmaceutical sponsor running a trial, an employer, a clinical call centre and a consumer caregiving application cannot plausibly share a single rate. Continuous monitoring in particular implies a fundamentally different consumption profile from episodic assessment, and the continuous variant is marketed as a distinct offering without any indication of how it is charged.
No return proxy is supplied either. The commercial argument rests on earlier detection of behavioral and cognitive change, and no cost per assessment, cost per case identified, or avoided cost figure is published to support it. The market sizing claims that appear instead, including a stated multi trillion dollar savings opportunity from early detection of cognitive illness, describe a category rather than a customer's economics.
Ask for the unit of charge, volume tiers, how continuous monitoring is priced against episodic assessment, and whether consumer and pharmaceutical embedding carry separate rates.