Sporo Health
Sporo Health, founded 2024, builds a multi agent documentation system on fine tuned medical language models rather than a single general purpose model, with agents that adapt to each clinician's specialty and learn from their edits. Alongside the scribe it runs a patient chart review agent that assembles a history from hundreds of documents and surfaces allergies, chronic conditions and social determinants, plus a cardiology workflow covering echo interpretation and risk stratification. Its positioning is explicitly against pure transcription tools, which it calls faster horses for speeding up an inefficient process without rethinking it.
Two things separate it from the rest of the long tail. It published a comparative evaluation of its own scribe on arXiv, measuring clinical content recall, precision and F1 against clinician written notes as ground truth and adding clinician satisfaction ratings on the validated PDQI-9 instrument. And it runs private model instances per client, so one customer's content does not commingle with another's model. Capture spans English, German, Spanish, Hindi, Greek and Arabic among others.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
A multi agent architecture built on fine tuned medical language models is the product, and the company frames its differentiation explicitly as not using general purpose models. The chart review and cardiology workflow agents are further model work rather than software wrapped around one.
Conventional draft and review, with an explicit feedback loop: the clinician edits the generated note and the system refines against those edits over time. Notes are produced in under two minutes for review rather than filed automatically. Held at B because no acceptance rate, edit burden figure or confidence threshold is published, which is a notable omission given that the vendor has already demonstrated it can measure recall and precision.
The most methodologically explicit vendor in the long tail of this category. Sporo published a comparative evaluation on arXiv describing its architecture as a multi agent system of fine tuned medical language models, and evaluating output against clinician written notes as ground truth using clinical content RECALL, PRECISION and F1, supplemented by satisfaction ratings on the modified PDQI-9, a validated documentation quality instrument.
Naming your metrics and your instrument is a different order of disclosure from claiming a percentage. Held at B because the paper is a preprint rather than peer reviewed, no model card or ongoing accuracy figure for the shipping product is published, and the specialised medical training data is described only in general terms.
One architectural commitment does substantial work on this axis: private model instances for each client, so one customer's content does not commingle with another's model. That is the strongest structural answer available short of self hosting, and it resolves what would otherwise be a tension in the design, because the system learns from clinician edits and those edits improve the client's own instance rather than a shared model.
Several vendors in this lane describe learning from customer content without ever defining what crosses between customers; this one forecloses it by architecture rather than by promise. The architecture is also described in real terms as a multi agent system of fine tuned medical language models, which tells a buyer the shape of what is running even without naming it. Held below the top grade because nothing is enumerated and the central claim is asserted rather than evidenced.
No base model, provider, version or hosting arrangement is named, no sub processor list was located, the specialised training data is characterised only in general terms, and the per client isolation is a statement rather than something covered by an attestation a buyer could read. Ask what the fine tuned models were fine tuned from, for a sub processor list, and whether the isolation claim is covered by any independent assessment.
Real evidence with a weak comparator, and the limitation matters as much as the result. The published study reports Sporo outperforming its comparator on recall, precision and F1, with output rated more favourably on accuracy, comprehensiveness and relevance and with fewer hallucinations.
Read what that actually establishes. The comparator is OpenAI's GPT-4o Mini, a small general purpose model rather than a competing medical scribe, so outperforming it is close to the expected result for any fine tuned medical system and says nothing about how Sporo compares with Abridge, Nabla or Heidi. The study is also vendor authored, published as a preprint, and the satisfaction ratings come from two raters, a medical student and a physician.
Graded B because a documented method with a validated instrument genuinely outranks the testimonials that carry most of this tail, not because the finding is strong.
One architectural commitment does real work here: private model instances for each client, which means one customer's content does not commingle with another's model. That also resolves what would otherwise be a tension, since the system learns from clinician edits. Those edits improve the client's own instance rather than a shared model, which is a coherent design rather than a training pipeline. HIPAA and GDPR compliant architecture with end to end encryption is claimed alongside.
Held at B because no retention schedule for audio or transcripts and no de identification practice was located, and the per client isolation claim is stated rather than evidenced by an attestation.
Claims compliance across both HIPAA and GDPR, consistent with a product marketed into multiple jurisdictions and supporting European languages. Business associate agreement terms are not published for inspection.
A second pass located no named or dated attestation, no SOC 2 report of either type, no HITRUST certification, no ISO 27001 and no trust centre.
What the vendor does publish is architectural rather than examined. Company authored material describes a HIPAA and GDPR compliant architecture, end to end encryption, and private model instances for each client. The last of those is the most substantive claim on this record and the one worth pursuing. A dedicated model instance per customer is a real isolation posture if implemented as described, and it is a stronger statement than most vendors at this size make. It remains a description of a design rather than an examination of controls, and there is no certification for the health privacy rule in any case, so the compliant architecture phrasing carries less than it appears to.
Measured against its own segment the absence is conspicuous rather than merely unstated, since competitors publish an independently examined posture and several surface it without being asked.
Fairness is due to the stage of the company. It was founded recently and a full attestation programme is a material cost at that size, so this grade reflects what a counterparty can verify before contracting rather than a judgement that controls are absent. Worth noting alongside it: the company has submitted its documentation quality to comparative evaluation in a published preprint using a modified documentation quality instrument, which is more external scrutiny than most vendors in this tier accept. That is evidence about product quality rather than about security posture and the two should not be read across, but it does indicate a company willing to be measured.
No clearance claimed and none located for the documentation product. No United States device pathway attaches to a note the clinician reviews and signs.
The flag raised in the earlier assessment stands and is now confirmed from a second source. The cardiology workflow is described as covering echo interpretation, note generation and risk stratification. Echo interpretation is diagnostic and risk stratification is prognostic. Neither is documentation, and both sit in territory where software is ordinarily regulated as a device rather than treated as a productivity tool. Nothing published resolves it: no clearance, no statement scoping the cardiology functions as clinician performed with the software assisting, and no description of what the system actually produces for an echo study. Establish whether the product interprets the study or organises a clinician's interpretation of it, because those are different regulatory objects wearing similar marketing language.
Markets are the second open question, and they matter more here than for a pure scribe precisely because of that module. The company is United States based and recently founded, multilingual support is claimed, and third party listings enumerate European languages including German and Greek. Ambient documentation is treated as software as a medical device in the United Kingdom and European Union, and a cardiology interpretation function would classify higher there than a scribe would. Establish which markets are actually served before relying on the United States position.
A sourcing caution belongs on this record. At least one third party directory attributes this product to a different company entirely, and a similarly named scribe from an unrelated vendor is separately indexed here. Confirm the corporate identity on any material before relying on it.
Partial rather than absent. To its credit, the published study measures hallucinations as an explicit outcome, which is a vendor putting a failure mode on record rather than only reporting successes. The chart review agent also surfaces social determinants of health, which is an equity relevant capability, though surfacing them is not the same as demonstrating fair performance across the populations they describe.
What is missing is the fairness work itself: no subgroup analysis, no accent or dialect performance disclosure, and no breakdown across the six or more languages the product claims to support, including Hindi, Greek and Arabic, which are precisely the cases where a medical speech model is most likely to degrade.
The most methodologically explicit record in the long tail of this category, and its existence is a useful corrective: rigour here is not correlated with size or funding. The company published a comparative evaluation describing its architecture as a multi agent system of fine tuned medical language models and assessing output against clinician written notes as ground truth, using clinical content recall, precision and an aggregate score, supplemented by satisfaction ratings on a validated documentation quality instrument.
Naming your metrics, your ground truth and your instrument is a different order of disclosure from publishing a percentage, because it tells a reader what was measured and how, and therefore what the result would mean if it were reproduced or contradicted. It also gives a customer something specific to hold the vendor to and something specific to re run. Held below the top grade for three reasons that should be read together rather than as fatal.
The work is a preprint rather than peer reviewed, so it has not been through external methodological scrutiny. No ongoing accuracy figure for the shipping product is published, so the evaluation describes a version at a point in time rather than what a customer receives today. And no warranty, indemnity or remediation commitment attaches to any of it, so the disclosure is transparency rather than accountability. Ask whether the evaluation is re run against current releases, and whether the published metrics can be written into an agreement as a service level.
Described as an EHR agnostic platform integrating with various record systems, but no named integration, architecture or write back mechanism was located, so agnostic here means compatible rather than connected. The access surface is genuinely broad, spanning web, iOS and Android applications, browser extensions and an API, and the API in particular means a group with development capacity can build the integration the vendor has not. Confirm what your own system actually gets.
Private model instances per client is a real deployment architecture rather than a policy, and it is the substantive claim here: isolation is structural, so a buyer is not relying on a promise that their data will be kept separate inside a shared system. Held at B because no hosting region, residency option or sub processor detail was located, and GDPR compliance is claimed without a stated European hosting arrangement.
No published rate card, tier structure or pricing model was located, though self serve account creation exists. For a vendor otherwise willing to publish evaluation methodology, the absence of any published price is an odd asymmetry.
Specialty adaptation is the design rather than a template library: agents are described as adapting to each clinician's specialty, with primary care, rheumatology and emergency medicine named as target settings chosen for high turnover and constrained conditions, and cardiology addressed through a separate workflow.
Language coverage is unusually spread for a small vendor, naming English, German, Spanish, Hindi, Greek and Arabic among others, which reaches beyond the European and Latin American languages most competitors stop at. Held at B because no specialty count is published and the named settings are aspirations as much as documented deployments.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published. Self serve account creation available.
|
Not disclosed. Sold to individual clinicians and to clinics and hospitals through group access. | HIPAA and GDPR compliance claimed. BAA terms not published. | None published. Access through web, iOS and Android apps, browser extensions and an API, so a group with development capacity can build its own integration. | Vendor Published |
No price published despite self serve signup being available, which is an odd asymmetry for a vendor willing to publish its evaluation methodology. Two questions carry more weight than the rate. First, ask what the private model instance actually means operationally, since per client isolation is the strongest claim in this record and the difference between a genuinely separate instance and a logically partitioned tenant is substantial.
Second, if the cardiology workflow is in scope, price and scope it separately: echo interpretation and risk stratification are diagnostic functions rather than documentation, and they belong in a different review than a scribe purchase.