Ambient Scribes
S

Sully.ai

Sully.ai sells what it calls AI Medical Employees, a set of role based agents covering reception, scribing, coding, nursing intake and triage, interpretation and clinical consultation, integrated with Epic, Oracle Cerner and athenahealth. Ambient documentation is one agent among several rather than the whole product, and the scribe extends past notes into diagnosis identification and care plan suggestion. The company is explicit about its trajectory: it frames its market as the roughly 800 billion dollars a year the United States spends on medical staff salaries, and describes its destination as a fully autonomous healthcare system.

Set against that ambition, it also publishes genuine technical research, including a benchmark across more than a dozen models and medical specialties reporting the best performer at 45 percent overall accuracy, and an architecture in which agents reach agreement through structured proposal and critique cycles with weighted scoring. A vendor publishing a 45 percent figure while marketing autonomous clinical agents is holding two positions at once, and both belong in a buyer's assessment.

AI Health Index verifiedJuly 23, 2026
Compare Sully.ai with other vendors
Founded
Headquarters
San Francisco, California, United States
Website
www.sully.ai/
Categories
ambient-scribes
Assessment

Capability Axes

An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read

AI Capability
AA on AI CentralityThe artificial intelligence is the product. Remove the model and there is nothing left to sell.
Vendor Published

The agents are the product and the company has no non AI business. Worth noting one dependency for completeness: speech infrastructure is partly supplied through a partnership with Speechmatics announced in January 2026, so the transcription layer is not wholly first party even though the agent architecture above it is.

CC on Autonomy and Oversight ModelAutonomy is claimed and oversight is asserted without a mechanism. Human in the loop appears as a phrase rather than a described control.
Vendor Published

The widest autonomous surface in this category and the least bounded in public. Agents independently answer inbound calls, schedule, perform intake and TRIAGE, process claims, and the scribe extends past documentation into identifying a diagnosis and suggesting a plan. No confidence threshold, escalation path, abstention behaviour or human review requirement was located for any agent.

The consensus mechanism deserves credit as a real verification design, agents critiquing one another and scoring proposals before acting, but it is AI checking AI: it reduces single model error without introducing a clinician, and it should not be read as an oversight model in the sense this axis measures. The stated destination of a fully autonomous healthcare system makes the disclosure gap more consequential rather than less.

BB on Model and Technology TransparencyThe approach or the suppliers are named without the version and update discipline behind them.
Vendor Published

Publishes real technical work rather than adjectives, and some of it is unflattering to its own category. Sully documents a SuperAgent architecture of isolated composable agent packages and a consensus mechanism in which agents reach verifiable agreement through structured proposal and critique cycles with weighted scoring and reputation tracking.

More notably it published a benchmark across more than 12 models and multiple medical specialties reporting the leading model at 45 percent overall accuracy and recommending ensemble use, which is a vendor telling the market that frontier models are far from reliable on medical tasks. That is candour in a direction that costs it something. Held at B because no accuracy figure, error rate or evaluation result was located for its own shipping agents, and the Doctor-LM personalisation layer is named but not explained.

CC on Model Supply Chain DisclosureThe architecture is described and no provider is named.
Vendor Published

The shape of the chain is described and its members are not named. The company publishes a benchmark comparing more than a dozen models across medical specialties and recommends ensemble use, and it ships an architecture of isolated composable agent packages with a consensus mechanism among agents, so a reader can infer that several third party models are invoked and that no single one is authoritative. Inference is not enumeration.

Nothing states which models are in the production ensemble, which providers supply them, what hosting arrangement runs them, or which parties appear on a sub processor list, and a personalisation layer is named without being explained.

The training boundary is the better half of the record: protected health information is not used for model training unless explicitly permitted, which is a permission gate rather than a use limitation, and a gate at least places the decision with the customer instead of leaving it to vendor interpretation.

What is missing alongside it is any retention schedule, which matters more here than for a single purpose scribe because the platform spans inbound calls, intake, triage, documentation and claims, and recorded triage or intake calls are patient speech captured outside a clinical encounter entirely. Ask which models the ensemble actually invokes, and for a retention schedule per content type.

CC on Clinical and Operational EvidenceNamed customers, or vendor reported percentages with no method, denominator or reference standard. Scale of use is recorded here and is not treated as evidence of benefit.
Vendor Published

Adoption claims without denominators and no clinical evidence. Published figures include partnership with more than 100 healthcare organisations, adoption across platforms serving over 100,000 providers and seven figure annual recurring revenue within ten months of launch, none of which measures clinical or operational benefit. Pilot feedback is described in terms of enthusiastic reactions. No peer reviewed study, controlled evaluation, third party performance rating or named health system outcome data was located.

BB on AI Safety and PHI StewardshipCategorical commitments are published, such as no training on customer data, without the retention schedule or the safety engineering behind them.
Vendor Published

The earlier assessment found nothing on training use, retention or de identification. The training question is now answered on the vendor's own material, which is the most important half of this axis.

The vendor states that protected health information is not used for model training unless explicitly permitted. That is a permission gate rather than a carve out, and it is the stronger of the two formulations available: a flat commitment never to train is cleaner still, but a permission gate at least places the decision with the customer rather than leaving it to a use limitation the vendor interprets. One peer in this lane uses the same construction and it is a good one.

Audit logging is described as carrying retention policies, and access to data is governed by role based controls with provisioning and deprovisioning processes.

What is still missing is the retention period itself. No stated schedule for encounter audio or transcripts was located, no deletion commitment with a time bound, and no de identification practice. A peer in this category publishes a default period and names the standard used to destroy data at the end of it, and that remains the bar.

The gap is wider than for a single product because of what this platform holds. Its agents touch inbound calls, patient intake, triage, clinical documentation and claims. Each produces a different class of content, and recorded triage or intake calls are patient speech captured outside a clinical encounter, which peers elsewhere in this lane also leave unaddressed. Retention is unlikely to be identical across them.

Ask for the schedule per content type, and for the training commitment in contract language rather than the help centre.

Regulatory and Compliance
BB on HIPAA and BAA PostureBusiness associate status is stated and supported by a substantive privacy document, with the agreement or its scope not fully published. For a vendor outside the United States, an equivalent regime documented to this depth grades here.
Vendor Published

The earlier assessment located no product specific posture. That is overturned. The vendor states it will sign a business associate agreement with covered entities or business associates before processing protected health information, and describes a data processing agreement whose security measures are aligned to the security rule.

The commitment is stated in the right shape. Before processing is the correct sequencing, since using a tool on protected health information without an executed agreement is itself a violation regardless of how secure the product is, and a vendor that says so has understood the obligation rather than treating the agreement as paperwork to follow adoption.

The surrounding commitments are the ones a counterparty actually needs: encryption in transit and at rest, role based access with provisioning and deprovisioning, audit logging with retention policies, breach notification commitments, and stated visibility into subprocessors. Subprocessor visibility is the item most often missing in this category and it is the one that determines whether the agreement means anything downstream.

One question holds this short of the top grade and it follows from the product's breadth. This is not a scribe but an agent suite whose components touch inbound calls, patient intake, triage, documentation and claims. Each of those creates or transmits protected health information in a different form and, in the case of claims, sends it onward to third parties. A single agreement may well cover all of it, but nothing published establishes the scope, and a buyer adopting one agent should not assume the terms they signed anticipated the others.

Ask for the agreement, its scope across the agent suite, and the subprocessor list rather than the promise of visibility.

BB on Security Certifications and Trust CenterA recognised certification is named in the vendor own material without the artefact, or with a scope or renewal question the buyer has to raise. A certification has a scope and a clock, and both are part of this grade.
Vendor Published

The earlier assessment found no attestation. That is overturned on the vendor's own material. Its help centre states that the security programme holds SOC 2 Type II and ISO 27001 attestations covering the platform and hosting environment, and that these are accessible through a trust portal.

Two attestations rather than one matters, because they examine different things. SOC 2 Type II tests whether controls operated effectively across a period. ISO 27001 certifies a management system for information security, meaning the processes by which risk is identified and controls are maintained over time. A vendor holding both has been examined on operation and on governance, and that combination puts this record near the top of this lane.

The supporting detail is also more specific than the category norm: encryption in transit and at rest with stated key management, role based access with provisioning and deprovisioning processes, session controls, audit logging with retention policies, breach notification commitments and stated visibility into subprocessors.

One further certification is reported by a third party comparison published by a direct competitor. Under this index's treatment that is a better class of evidence than an aggregator, since conceding a rival's credentials runs against the publisher's interest, but it is not the vendor evidencing its own posture and is not graded on.

What would move this to the top grade is specificity rather than more logos: the report dates, the audit periods, and a scope statement naming which of the platform's agents are inside the assessed boundary. That last point matters because this is an agent suite rather than a single product.

CC on FDA and Regulatory StatusNo device claim is made and the product is scoped accordingly. Most administrative and operational products sit here and are not penalised for it, because this axis grades the appropriateness of the positioning rather than possession of a clearance.
Vendor Published

No clearance claimed and none located. The flag from the earlier assessment stands and the second pass adds an architectural claim that bears directly on it.

Several described capabilities sit closer to regulated clinical decision support than to documentation: identifying a diagnosis, suggesting a care plan, and operating a triage agent. Where a product moves from recording what a clinician decided to proposing what the decision should be, the software as a medical device question becomes live, and documentation status does not carry across. A buyer should ask how the vendor has assessed each capability rather than assume the scribe's position covers the suite.

The new element is the vendor's stated approach to reliability. It describes a consensus mechanism in which multiple expert models review a case and agree on the most reliable clinical output, rather than relying on a single model's response. That is a considered design and it is more than most vendors offer by way of explaining how they reach an answer.

It also sharpens the regulatory question rather than settling it. The reasoning that keeps decision support outside device regulation depends on a professional being able to review the basis of a recommendation independently. A consensus across several models is harder to review than a single chain, because what the clinician sees is an agreed output rather than the disagreement that preceded it. Ask what is surfaced when the models do not agree, whether a confidence or dissent signal reaches the clinician, and what the interface shows for a low agreement case.

The triage agent remains the sharpest item, since urgency determinations are made in exchanges where no clinician is present.

CC on AI Governance and Bias DisclosureResponsible artificial intelligence is committed to in policy language with no evaluation behind it. Most of the index sits here.
Vendor Published

No fairness statement, subgroup analysis or accent and dialect disclosure was located, and no governance framework covering agents that triage patients or propose diagnoses. The commercial framing compounds the gap: the market is described as the roughly 800 billion dollars spent annually on United States medical staff salaries, and the destination as a fully autonomous healthcare system, which is a labour substitution thesis rather than an augmentation one and carries different obligations.

To its credit the published 45 percent benchmark result cuts directly against its own marketing trajectory, and a vendor willing to publish that number is better placed than most to publish the rest. The absence here is of governance disclosure, not of technical seriousness.

CC on AI Liability and RecourseMechanisms exist that let someone challenge an output, such as audit trails, source traceability or review before commit, with nothing standing behind the output and no route for the harmed party.
Vendor Published

Real published technical work sits on this record, and the way it is distributed produces an odd result a buyer should notice. The company published a benchmark across more than a dozen models and several medical specialties, reporting the leading model at 45 percent overall accuracy and recommending ensemble use.

That is a vendor telling the market that frontier models are far from reliable on medical tasks, which is candour in a direction that costs it something and is exactly the sort of falsifiable published work this axis rewards. The oddity is what it leaves the buyer holding. The company has published a rigorous and alarming number about the models underneath this class of product, and no accuracy figure, error rate or evaluation result for its own shipping agents.

So the only measurement available is the discouraging one, and the reassuring one does not exist. A buyer cannot tell whether the ensemble and consensus design closes the gap the benchmark describes, which is the single question the benchmark raises.

The architecture is documented in genuine detail, with isolated composable agent packages and a consensus mechanism in which agents reach agreement through structured proposal and critique with weighted scoring, and a permission gate governs training so the decision on whether protected health information is used sits with the customer. Those are mechanisms a customer can use. No warranty, indemnity or remediation commitment was located. Ask for the same benchmark run against the shipping agents.

Integration and Deployment
BB on EHR and Interoperability DepthNamed systems with read access or one directional writing, or standards support with named deployments behind it.
Vendor Published

Named integrations with Epic, Oracle Cerner and athenahealth, with outputs described as landing directly where clinicians work and an EHR integration agent offered in the higher tier. Held at B because integration depth, write back mechanism and certification status were not documented or verified in this pass, and connects instantly is a claim rather than an architecture.

CC on Deployment Model and Data ResidencyA single hosted option with location implied rather than committed.
Vendor Published

No hosting region and no residency commitment were located on the vendor's own material, and no subprocessor list was published, though the vendor states subprocessor visibility is available through its trust portal and that its attestations cover the hosting environment.

Subprocessor visibility as a stated commitment is better than silence and is not the same as disclosure. A buyer cannot compare vendors on a portal they have to request access to, and the chain is the thing that determines whether the agreement and the training commitment survive downstream. The speech partnership already on this record makes that concrete rather than theoretical: at least one external party processes audio, so the question is what else does.

One reported capability would change this answer entirely if confirmed and is the thing to chase first. A third party comparison published by a direct competitor describes a self hosted deployment option for organisations with data residency requirements. If that exists it resolves residency outright, since the customer holds the environment, and it would place this vendor among the very few in this category offering an on premise path. It was not confirmed on the vendor's own material in this pass and is not graded on. Ask for it directly, and ask which components can run self hosted, because an agent suite that reaches inbound calls and claims is unlikely to be fully deployable inside a customer boundary.

Also unresolved: which model providers process the encounter, and whether the multi model consensus approach the vendor describes means more than one external service sees the same clinical content.

Ask for the region, the subprocessor list itself, and the self hosted scope.

Commercial
BB on Commercial TransparencyA price or a pricing basis is published without full tiers, so a buyer can size the cost before making contact.
Third Party Estimated

Tiered pricing is reported at 79 US dollars per provider per month for a Pro tier covering the scribe, chat, custom templates and support, and 99 US dollars for a Premium tier adding the EHR integration agent, decision support and research, writer and interpreter agents.

Held at B rather than A on sourcing: those figures were located through third party comparison pages including one published by a competitor, not confirmed on the vendor's own pricing page in this pass, and the go to market is described as demo led with sales negotiation for larger organisations. Confirm the tiers directly before relying on them.

BB on Setting and Specialty CoverageCoverage is named with validation behind part of it.
Vendor Published

Broad across the encounter rather than deep in one setting: coverage spans before, during and after the visit, from reception and intake through documentation to coding and follow up, across a claimed 50 or more medical and surgical specialties with multilingual support through a dedicated interpreter agent. Targeted at hospitals, health systems and larger clinics. Held at B because the specialty count is asserted without enumeration and no setting specific validation was located.

Tracked Since Listing

What Changed

Material product, regulatory, evidence and commercial changes at Sully.ai, each verified against a live source and tagged to the capability axis it bears on. Funding rounds and awards are not product changes and are not logged.

Aug 6, 2026Product / capability

Sully.ai released an AI scribe specifically designed for nursing shifts, expanding its capabilities beyond standard physician visit documentation. The new tool is tailored to handle nursing-specific workflows, including timestamps, flowsheets, and shift handoffs.

Bears on: Setting and Specialty CoverageSource
Our read on this change →Tracked since Aug 2026
Comparisons

Compared With

Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Reported at $79 per provider per month Pro, $99 Premium. Not confirmed on vendor pricing page.
$79 baseline
Per provider per month tiers by agent bundle, with demo led sales and negotiated agreements for hospitals and health systems. Not retrieved in this verification pass Not published. Third Party Estimated

Tier figures here are third party reported, including from a competitor's comparison page, rather than confirmed on the vendor's own pricing page in this pass, so treat them as indicative. The more important commercial point is what is being priced.

Sully sells agents as staff substitutes and frames its market as the United States medical salary bill, so a per provider per month comparison against other scribes understates the intended scope: the buying decision is presented as replacing roles rather than adding a documentation tool. Buyers should price the agents they actually intend to run, and should establish oversight requirements for the triage and diagnosis suggestion capabilities separately from cost.