Behavioral Health AI
C

Callyope

Callyope is a Paris company building an audio and language foundation model trained specifically on psychiatric data rather than adapted from a general purpose speech model. Patients use a smartphone application for voice journals, short structured speech tasks, symptom self reports, or passive analysis of calls with designated caregivers, and the model combines that speech with smartphone sensor signals covering sleep and activity and with clinical information to produce continuous symptom scores between appointments.

The clinician side adds case summarisation from uploaded records, consultation notes, referral and hospitalisation documents, and identification of gaps in medical coding. The founding team is a research team: Rachid Riad completed a doctorate at the Ecole Normale Superieure on automatic assessment of cognitive, linguistic and emotional disorders affecting speech in Huntington disease, and the company publishes at speech science venues, including two papers at Odyssey 2026, one on learning health related speech representations and one on approximate signal processing as a route toward homomorphic encryption for audio. A registered trial of voice based biomarkers for predicting schizophrenia relapse is running with an estimated 200 participants and completes in October 2026.

AI Health Index verifiedAugust 3, 2026
Compare Callyope with other vendors
Founded
2023
Headquarters
Paris, France
Categories
behavioral-health, remote-monitoring, ambient-scribes
Assessment

Capability Axes

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read

AI Capability
AA on AI CentralityThe artificial intelligence is the product. Remove the model and there is nothing left to sell.
Vendor Published

The model is the product and it was built rather than borrowed. The company trained an audio and language foundation model on psychiatric data specifically, rather than fine tuning a general purpose speech system, which is the distinction that separates this from most speech based health products. Every function on the clinician side, from symptom scoring to case summarisation to document generation, runs on that same model.

The team is a research team publishing on the model at speech science venues, and the technical founder's doctorate was on automatic assessment of speech in a neurological disease, so the mechanism claim is backed by primary work rather than asserted in marketing.

BB on Autonomy and Oversight ModelThe oversight structure is described and one part is missing, commonly the threshold at which the system stops or what happens after it is wrong.
Vendor Published

Conventional and clearly stated, with candour credited. The company states plainly that the system is designed to support clinicians rather than replace them, and that the clinician remains in full control of diagnosis, treatment planning and care. Output is a symptom score and a set of drafted documents that a clinician reviews.

That is the same restraint credited on the MD-Staff and Genomind records, and it matters more than usual here, because a continuous score presented longitudinally carries an implied authority that a single reading does not. What is not addressed is the escalation question: if the model detects deterioration between appointments and nobody opens the dashboard, no published material describes what happens next or who is accountable for noticing.

BB on Model and Technology TransparencyThe approach or the suppliers are named without the version and update discipline behind them. Naming a supplier is the entry to this band both here and on Model Supply Chain Disclosure, which ask different questions of the same disclosure: who receives the data, and what produces the output.
Vendor Published

Well above the norm for this index and the transparency is in the scientific literature rather than in a marketing page. The company presented two papers at Odyssey 2026, one on learning speaker and health related representations from natural language supervision and one on quantized approximate signal processing as a path toward homomorphic encryption for audio.

Separate published work examines which pretext tasks in speech foundation models transfer to mental health detection and how different model layers encode the relevant features, including segment length and pooling strategy. Crucially that work reports results on the Androids dataset, a public third party benchmark, alongside the company's own data, which makes part of the performance externally checkable.

Held at B because the production model is a different artefact from the research papers: weights are not released, the proprietary psychiatric training data is not characterised, and the headline claim of assessing ten or more symptoms from thirty seconds of speech at over ninety percent accuracy is not tied to any published evaluation. Publishing the evaluation behind that specific number would move this to an A.

BB on Model Supply Chain DisclosureSubstantial partial disclosure, or a chain that is structurally short: an in house build, a cleared model that cannot be quietly swapped, or a deployment where the transfer does not occur at all. Naming the model provider exits the band below into this one; the axis rises from here on the completeness of the party list and on the terms that govern data once it arrives.
Vendor Published

The boundary is stated in the strongest terms found in this index and the chain is implied rather than enumerated. The company publishes the commitment this index has asked repeatedly of others and rarely received: patient data is never used to train its models without explicit consent, is never shared with third parties, and is encrypted in transit and at rest.

A never shared statement bounds the chain to the company itself, which is a structurally short chain rather than a disclosed one, and the consent gate on training places the decision with the customer rather than leaving it to a use limitation the vendor interprets. Rarer still, the company conducts primary research on privacy preserving computation for its own data type, presenting work on approximate signal processing as a route toward computing on encrypted audio.

Voice is among the most identifying data any health product holds, because it carries speaker identity independently of what is said, and a company researching how to compute on it without decrypting it is addressing the risk at its root rather than wrapping it in policy.

Held below the top grade because nothing is named: no model or provider, no hosting arrangement, no sub processor list, and the proprietary psychiatric training corpus is uncharacterised, which for recorded speech from psychiatric populations is a provenance question worth asking. Ask whose speech that corpus contains and on what consent basis.

CC on Clinical and Operational EvidenceNamed customers, or vendor reported percentages with no method, denominator or reference standard. Scale of use is recorded here and is not treated as evidence of benefit.
Vendor Published

The right studies are running and the results are not in yet, which is an unusual and honest position rather than a weak one. A registered trial of voice based biomarkers for monitoring and predicting schizophrenia relapse enrols an estimated 200 participants across six months of repeated voice interviews and completes in October 2026, so it is live at the time of this assessment.

The company states seven clinical studies across more than 1,000 patients and a research partnership with one of Europe's largest psychiatric institutions. Against that, no completed clinical outcome study was located. The published work is detection accuracy in machine learning venues, which establishes that speech signal carries symptom information but not that acting on it changes what happens to patients, and those are different claims.

One item is owed: the chief executive said in late 2023 that a scientific paper on general population results would follow the next year, and that paper was not located in two passes. This is the inverse of the evidence paradox recorded in the medication safety lane, where the vendors running the most rigorous designs got the least favourable results. Here a company is running a proper prospective design and simply has not finished.

AA on AI Safety and PHI StewardshipRetention windows, training use and de identification are all stated specifically enough to be contradicted. Safety engineering, meaning guardrails, hallucination mitigation and how a safety event is handled, does not gate this grade: no vendor in this index publishes it, so every note on this axis states that gap rather than leaving it implied.
Vendor Published

The strongest stewardship posture encountered in this index, on two independent grounds. First, the company publishes the commitment this index has spent months asking other vendors for and almost never receiving: patient data is never used to train its models without explicit consent, is never shared with third parties, and is encrypted in transit and at rest.

Secondary use of clinical data for model improvement is the single most requested and least published position in this whole market, and it is stated here in plain terms. Second, and rarer still, the company is doing primary research on privacy preserving computation for its own data type, presenting work at Odyssey 2026 on approximate signal processing as a route toward homomorphic encryption for audio.

Voice is among the most identifying data any health product can hold, since it carries speaker identity independently of what is said, and a company researching how to compute on it without decrypting it is addressing the risk at its root rather than wrapping it in policy. Very few vendors anywhere in this index conduct original privacy research.

Regulatory and Compliance
BB on HIPAA and BAA PostureBusiness associate status is stated and supported by a substantive privacy document, with the agreement or its scope not fully published. For a vendor outside the United States, an equivalent regime documented to this depth grades here.
Vendor Published

Graded against the framework that actually governs the company rather than the American one, following the precedent set on the LGPD scoping call. This is a French company operating under the General Data Protection Regulation, and it states full compliance with it and the use of hosting certified under the French Hebergeur de Donnees de Sante regime for all health data.

That regime is an audited health data specific certification rather than a self declaration, which puts it well ahead of the vendors in this index whose entire published position is the phrase HIPAA compliant. It also answers, for this vendor, the certification question left open against Synapse Medicine. Held at B for a distinction worth drawing carefully: the company states that it uses certified hosting, which places the certification with the hosting provider.

Using a certified host and holding certification yourself are different assurances, and the second covers the vendor's own handling of the data once it arrives. No data processing agreement terms, data protection officer contact or processing register were located.

BB on Security Certifications and Trust CenterA recognised certification is named in the vendor own material without the artefact, or with a scope or renewal question the buyer has to raise. A certification has a scope and a clock, and both are part of this grade.
Vendor Published

Among the better positions in the recent run of builds, and notable because it is an audited certification rather than an assertion. The company states that it is certified to ISO 27001 and that data is encrypted in transit and at rest. Very few vendors assessed in this index in the last several lanes publish any external certification at all.

Held at B rather than A because the certification is stated without a certificate number, an issuing body or a declared scope, and scope is where ISO 27001 certifications differ most: a certificate covering a corporate function is not the same as one covering the production platform that processes patient speech. No SOC 2 report, penetration testing statement, trust centre or vulnerability disclosure policy was located.

CC on FDA and Regulatory StatusNo device claim is made and the product is scoped accordingly. Most administrative and operational products sit here and are not penalised for it, because this axis grades the appropriateness of the positioning rather than possession of a clearance.
Vendor Published

No FDA involvement and no United States market entry located. On the European side, no medical device certification was located either. The relevant statement of intent is on the record from late 2023, when the chief executive described the plan as freezing a version of the proprietary model and submitting it to health authorities for certification, which is the correct sequence and is stated more clearly than most vendors manage.

So the product currently appears to operate outside device certification while positioned to seek it. The boundary is close: software intended to monitor psychiatric symptoms and inform clinical decisions sits within the medical device definition under the European regulation on medical devices, and the company's own description of providing objective data to inform clinical decisions is close to that line. Recorded as a status to re check rather than as a criticism, since a company saying it will certify later is being straightforward about where it stands.

CC on AI Governance and Bias DisclosureResponsible artificial intelligence is committed to in policy language with no evaluation behind it. Most of the index sits here.
Vendor Published

The published research shows awareness of generalisation, examining transfer across languages and speech tasks explicitly, which is more than most vendors attempt. But no bias testing methodology, subgroup performance, model card or drift monitoring concept was located, and one specific exposure is visible in the company's own trial design. The registered schizophrenia study recruits French speakers only and excludes anyone with a condition that impairs French.

Speech models carry documented accuracy variation across accent, dialect and first language, and this index has recorded that the affected populations overlap with those already underserved. A psychiatric speech model validated in one language and marketed in another needs published cross language performance before a clinician can know whether a falling score reflects deteriorating mental state or an accent the model handles less well. That question is sharper here than for a scribe, because the output is a clinical symptom score rather than a transcript a human immediately checks.

BB on AI Liability and RecourseA published falsifiable commitment, or a real correction route the affected person can exercise against the vendor. A published error rate with its method and denominator grades here. A statutory right that runs to the covered entity rather than to the vendor does not reach this band on its own: every other route here asks something of the vendor, and being located in a particular jurisdiction is not conduct.
Peer Reviewed Publication

The route into this band is published research evaluated against a public benchmark, which is the strongest falsifiability signal available short of a contractual commitment. The company has presented work at a named venue on learning speaker and health related representations from natural language supervision, and separate published work examines which pretext tasks in speech foundation models transfer to mental health detection and how model layers encode the relevant features.

Crucially that work reports results on a public third party dataset alongside the company's own, so part of the performance record can be checked by someone outside the company against a standard the company did not set. Very little in this index reaches that bar. One caution belongs here rather than as a footnote, because it is the specific way rigour can mislead a buyer.

The production model is a different artefact from the research papers: weights are not released, the proprietary psychiatric training data is not characterised, and the headline commercial claim, assessing ten or more symptoms from thirty seconds of speech at better than ninety percent accuracy, is not tied to any published evaluation. So the credible work and the marketed number are not connected, and a reader should not treat the first as validating the second. Publishing the evaluation behind that specific claim would move this to the top grade. No warranty, indemnity or remediation commitment was located.

Integration and Deployment
CC on EHR and Interoperability DepthIntegration is claimed through standards or a middleware layer with no system named and nothing to verify.
Vendor Published

The product does move clinical data: it ingests uploaded medical records for case summarisation, takes smartphone sensor signals covering sleep and activity, and emits structured documents including consultation notes, referral letters and hospitalisation reports. What was not located in two passes is any published integration: no named record system, no interoperability standard, and no partnership with a hospital information system supplier.

Delivered as a patient smartphone application plus a clinician web application, which means documents are produced for a clinician to place into the record rather than written back into it, and that manual step is where adoption of this kind of product usually fails. The integration burden in European hospital systems differs from the American one, so an absence here is less telling than it would be for a United States vendor, but it is still an absence.

BB on Deployment Model and Data ResidencyOptions and residency are stated with isolation or the processing path left open.
Vendor Published

A clear and well governed answer to where the data goes, and a different answer from the one Carenostics gives. Health data is processed vendor side, hosted under the French health data hosting certification regime, which fixes it inside a European jurisdiction and inside an audited hosting framework rather than leaving residency to a contract term.

That is the strongest posture available to a vendor that must centralise data, and centralising is unavoidable here, since the model consumes speech from a patient's own phone. Held at B rather than A precisely because the data does move: patients transmit voice recordings from personal devices to vendor infrastructure, which is a materially different exposure from software that runs inside a hospital and never sends anything out. No hosting provider, region detail or architecture description was published.

Commercial
DD on Commercial TransparencyNothing a buyer can establish before a sales conversation. A published pricing claim contradicted by evidence also grades here.
Vendor Published

Nothing published. No pricing, no pricing mechanism, no contracting model and no stated basis of charge, and no reimbursement pathway identified in any market, which matters for a product whose value proposition is continuous monitoring between appointments, a service most health systems have no existing code to pay for.

The visible commercial facts are funding: a seed round of roughly 2.2 million euros co led by 360 Capital and Bpifrance Digital Venture with No Label Ventures participating, plus selection for Google's artificial intelligence for health programme and an Amazon Web Services cohort. That is a small round against a product carrying both a research programme and a clinical trial, which is worth a buyer's attention on supplier continuity grounds.

CC on Setting and Specialty CoverageCoverage is claimed broadly without specifics, or stated clearly with nothing validating it yet.
Vendor Published

Deliberately narrow on all three of specialty, language and geography. Specialty is psychiatry, covering depression, bipolar disorder, schizophrenia, anxiety, insomnia and psychosis, with the company's own framing extending to psychiatric, cognitive and motor symptoms of brain disorders more broadly.

Setting is outpatient psychiatric care between appointments, plus hospital use through the institutional partnership, which is a real gap to target: the company's stated problem is that a patient with schizophrenia may see a psychiatrist only every four to six weeks and deterioration in that interval goes unobserved.

Coverage is constrained by language before it is constrained by anything else, since a speech model is only validated in the languages it was tested in, and the registered trial is French only. Expansion is therefore a revalidation problem per language rather than a translation problem.

Comparisons

Compared With

Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

No pricing data has been verified for this vendor. Pricing information will be published here once confirmed through vendor disclosure or third-party estimation.