Sporo Health
Sporo Health, founded 2024, builds a multi agent documentation system on fine tuned medical language models rather than a single general purpose model, with agents that adapt to each clinician's specialty and learn from their edits. Alongside the scribe it runs a patient chart review agent that assembles a history from hundreds of documents and surfaces allergies, chronic conditions and social determinants, plus a cardiology workflow covering echo interpretation and risk stratification. Its positioning is explicitly against pure transcription tools, which it calls faster horses for speeding up an inefficient process without rethinking it. Two things separate it from the rest of the long tail. It published a comparative evaluation of its own scribe on arXiv, measuring clinical content recall, precision and F1 against clinician written notes as ground truth and adding clinician satisfaction ratings on the validated PDQI-9 instrument. And it runs private model instances per client, so one customer's content does not commingle with another's model. Capture spans English, German, Spanish, Hindi, Greek and Arabic among others.
Capability Axes
A multi agent architecture built on fine tuned medical language models is the product, and the company frames its differentiation explicitly as not using general purpose models. The chart review and cardiology workflow agents are further model work rather than software wrapped around one.
Conventional draft and review, with an explicit feedback loop: the clinician edits the generated note and the system refines against those edits over time. Notes are produced in under two minutes for review rather than filed automatically. Held at B because no acceptance rate, edit burden figure or confidence threshold is published, which is a notable omission given that the vendor has already demonstrated it can measure recall and precision.
The most methodologically explicit vendor in the long tail of this category. Sporo published a comparative evaluation on arXiv describing its architecture as a multi agent system of fine tuned medical language models, and evaluating output against clinician written notes as ground truth using clinical content RECALL, PRECISION and F1, supplemented by satisfaction ratings on the modified PDQI-9, a validated documentation quality instrument. Naming your metrics and your instrument is a different order of disclosure from claiming a percentage. Held at B because the paper is a preprint rather than peer reviewed, no model card or ongoing accuracy figure for the shipping product is published, and the specialised medical training data is described only in general terms.
Real evidence with a weak comparator, and the limitation matters as much as the result. The published study reports Sporo outperforming its comparator on recall, precision and F1, with output rated more favourably on accuracy, comprehensiveness and relevance and with fewer hallucinations. Read what that actually establishes: the comparator is OpenAI's GPT-4o Mini, a small GENERAL PURPOSE model, not a competing medical scribe, so outperforming it is close to the expected result for any fine tuned medical system and says nothing about how Sporo compares with Abridge, Nabla or Heidi. The study is also vendor authored, published as a preprint, and the satisfaction ratings come from two raters, a medical student and a physician. Graded B because a documented method with a validated instrument genuinely outranks the testimonials that carry most of this tail, not because the finding is strong.
One architectural commitment does real work here: PRIVATE MODEL INSTANCES FOR EACH CLIENT, which means one customer's content does not commingle with another's model. That also resolves what would otherwise be a tension, since the system learns from clinician edits: those edits improve the client's own instance rather than a shared model, which is a coherent design rather than a training pipeline. HIPAA and GDPR compliant architecture with end to end encryption is claimed alongside. Held at B because no retention schedule for audio or transcripts and no de identification practice was located, and the per client isolation claim is stated rather than evidenced by an attestation.
Claims compliance across both HIPAA and GDPR, consistent with a product marketed into multiple jurisdictions and supporting European languages. Business associate agreement terms are not published for inspection.
Not assessed. No named or dated attestation and no trust centre located in this pass, which is the obvious gap next to an otherwise specific architectural privacy claim.
No clearance claimed and none located for the documentation product. Flag rather than grade: the cardiology workflow is described as covering echo interpretation and risk stratification, which are diagnostic and prognostic functions rather than documentation, and sit far closer to software as a medical device territory than note generation does. Ask how that capability has been scoped.
Partial rather than absent. To its credit, the published study measures HALLUCINATIONS as an explicit outcome, which is a vendor putting a failure mode on record rather than only reporting successes. The chart review agent also surfaces SOCIAL DETERMINANTS OF HEALTH, which is an equity relevant capability, though surfacing them is not the same as demonstrating fair performance across the populations they describe. What is missing is the fairness work itself: no subgroup analysis, no accent or dialect performance disclosure, and no breakdown across the six or more languages the product claims to support, including Hindi, Greek and Arabic, which are precisely the cases where a medical speech model is most likely to degrade.
Described as an EHR agnostic platform integrating with various record systems, but no named integration, architecture or write back mechanism was located, so agnostic here means compatible rather than connected. The access surface is genuinely broad, spanning web, iOS and Android applications, browser extensions and an API, and the API in particular means a group with development capacity can build the integration the vendor has not. Confirm what your own system actually gets.
Private model instances per client is a real deployment architecture rather than a policy, and it is the substantive claim here: isolation is structural, so a buyer is not relying on a promise that their data will be kept separate inside a shared system. Held at B because no hosting region, residency option or sub processor detail was located, and GDPR compliance is claimed without a stated European hosting arrangement.
No published rate card, tier structure or pricing model was located, though self serve account creation exists. For a vendor otherwise willing to publish evaluation methodology, the absence of any published price is an odd asymmetry.
Specialty adaptation is the design rather than a template library: agents are described as adapting to each clinician's specialty, with primary care, rheumatology and emergency medicine named as target settings chosen for high turnover and constrained conditions, and cardiology addressed through a separate workflow. Language coverage is unusually spread for a small vendor, naming English, German, Spanish, Hindi, Greek and Arabic among others, which reaches beyond the European and Latin American languages most competitors stop at. Held at B because no specialty count is published and the named settings are aspirations as much as documented deployments.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published. Self serve account creation available.
|
Not disclosed. Sold to individual clinicians and to clinics and hospitals through group access. | HIPAA and GDPR compliance claimed. BAA terms not published. | None published. Access through web, iOS and Android apps, browser extensions and an API, so a group with development capacity can build its own integration. | Vendor Published |
No price published despite self serve signup being available, which is an odd asymmetry for a vendor willing to publish its evaluation methodology. Two questions carry more weight than the rate. First, ask what the private model instance actually means operationally, since per client isolation is the strongest claim in this record and the difference between a genuinely separate instance and a logically partitioned tenant is substantial. Second, if the cardiology workflow is in scope, price and scope it separately: echo interpretation and risk stratification are diagnostic functions rather than documentation, and they belong in a different review than a scribe purchase.