Ambient Scribes
S

Sully.ai

Sully.ai sells what it calls AI Medical Employees, a set of role based agents covering reception, scribing, coding, nursing intake and triage, interpretation and clinical consultation, integrated with Epic, Oracle Cerner and athenahealth. Ambient documentation is one agent among several rather than the whole product, and the scribe extends past notes into diagnosis identification and care plan suggestion. The company is explicit about its trajectory: it frames its market as the roughly 800 billion dollars a year the United States spends on medical staff salaries, and describes its destination as a fully autonomous healthcare system. Set against that ambition, it also publishes genuine technical research, including a benchmark across more than a dozen models and medical specialties reporting the best performer at 45 percent overall accuracy, and an architecture in which agents reach agreement through structured proposal and critique cycles with weighted scoring. A vendor publishing a 45 percent figure while marketing autonomous clinical agents is holding two positions at once, and both belong in a buyer's assessment.

Last VerifiedJuly 23, 2026
Compare Sully.ai with other vendors
Founded
Headquarters
San Francisco, California, United States
Website
www.sully.ai/
Categories
ambient-scribes
Assessment

Capability Axes

AI Capability
AI Centrality
A
Vendor Published

The agents are the product and the company has no non AI business. Worth noting one dependency for completeness: speech infrastructure is partly supplied through a partnership with Speechmatics announced in January 2026, so the transcription layer is not wholly first party even though the agent architecture above it is.

Autonomy and Oversight Model
C
Vendor Published

The widest autonomous surface in this category and the least bounded in public. Agents independently answer inbound calls, schedule, perform intake and TRIAGE, process claims, and the scribe extends past documentation into identifying a diagnosis and suggesting a plan. No confidence threshold, escalation path, abstention behaviour or human review requirement was located for any agent. The consensus mechanism deserves credit as a real verification design, agents critiquing one another and scoring proposals before acting, but it is AI checking AI: it reduces single model error without introducing a clinician, and it should not be read as an oversight model in the sense this axis measures. The stated destination of a fully autonomous healthcare system makes the disclosure gap more consequential rather than less.

Model and Technology Transparency
B
Vendor Published

Publishes real technical work rather than adjectives, and some of it is unflattering to its own category. Sully documents a SuperAgent architecture of isolated composable agent packages and a consensus mechanism in which agents reach verifiable agreement through structured proposal and critique cycles with weighted scoring and reputation tracking. More notably it published a benchmark across more than 12 models and multiple medical specialties reporting the leading model at 45 percent overall accuracy and recommending ensemble use, which is a vendor telling the market that frontier models are far from reliable on medical tasks. That is candour in a direction that costs it something. Held at B because no accuracy figure, error rate or evaluation result was located for its own shipping agents, and the Doctor-LM personalisation layer is named but not explained.

Clinical and Operational Evidence
C
Vendor Published

Adoption claims without denominators and no clinical evidence. Published figures include partnership with more than 100 healthcare organisations, adoption across platforms serving over 100,000 providers and seven figure annual recurring revenue within ten months of launch, none of which measures clinical or operational benefit. Pilot feedback is described in terms of enthusiastic reactions. No peer reviewed study, controlled evaluation, third party performance rating or named health system outcome data was located.

AI Safety and PHI Stewardship
Not rated

Not assessed. No statement on audio or transcript retention, de identification or training use was located in this pass, which is a significant gap for a platform whose agents touch calls, intake, triage, notes and claims.

Regulatory and Compliance
HIPAA and BAA Posture
Not rated

Not assessed. No product specific business associate agreement posture was located in this pass.

Security Certifications and Trust Center
Not rated

Not assessed. No named or dated attestation and no trust centre located in this pass.

FDA and Regulatory Status
Not rated

No clearance claimed and none located. Flag rather than grade: several described capabilities sit closer to regulated clinical decision support than to documentation, specifically identifying a diagnosis, suggesting a care plan and operating a triage nurse agent. Where a product moves from recording what a clinician decided to proposing what the decision should be, the software as a medical device question becomes live, and a buyer should ask how the vendor has assessed it rather than assume documentation status carries across.

AI Governance and Bias Disclosure
C
Vendor Published

No fairness statement, subgroup analysis or accent and dialect disclosure was located, and no governance framework covering agents that triage patients or propose diagnoses. The commercial framing compounds the gap: the market is described as the roughly 800 billion dollars spent annually on United States medical staff salaries, and the destination as a fully autonomous healthcare system, which is a labour substitution thesis rather than an augmentation one and carries different obligations. To its credit the published 45 percent benchmark result cuts directly against its own marketing trajectory, and a vendor willing to publish that number is better placed than most to publish the rest. The absence here is of governance disclosure, not of technical seriousness.

Integration and Deployment
EHR and Interoperability Depth
B
Vendor Published

Named integrations with Epic, Oracle Cerner and athenahealth, with outputs described as landing directly where clinicians work and an EHR integration agent offered in the higher tier. Held at B because integration depth, write back mechanism and certification status were not documented or verified in this pass, and connects instantly is a claim rather than an architecture.

Deployment Model and Data Residency
Not rated

Not assessed. Multimodal access across voice, web, phone and SMS is described, but no hosting region, residency option or sub processor detail was located, and the Speechmatics speech partnership makes the sub processor question concrete rather than theoretical.

Commercial
Commercial Transparency
B
Third Party Estimated

Tiered pricing is reported at 79 US dollars per provider per month for a Pro tier covering the scribe, chat, custom templates and support, and 99 US dollars for a Premium tier adding the EHR integration agent, decision support and research, writer and interpreter agents. Held at B rather than A on sourcing: those figures were located through third party comparison pages including one published by a competitor, not confirmed on the vendor's own pricing page in this pass, and the go to market is described as demo led with sales negotiation for larger organisations. Confirm the tiers directly before relying on them.

Setting and Specialty Coverage
B
Vendor Published

Broad across the encounter rather than deep in one setting: coverage spans before, during and after the visit, from reception and intake through documentation to coding and follow up, across a claimed 50 or more medical and surgical specialties with multilingual support through a dedicated interpreter agent. Targeted at hospitals, health systems and larger clinics. Held at B because the specialty count is asserted without enumeration and no setting specific validation was located.

Commercial

Pricing

Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.

Entry Price Pricing Basis BAA Tier Implementation Source
Reported at $79 per provider per month Pro, $99 Premium. Not confirmed on vendor pricing page.
$79 baseline
Per provider per month tiers by agent bundle, with demo led sales and negotiated agreements for hospitals and health systems. Not retrieved in this verification pass Not published. Third Party Estimated

Tier figures here are third party reported, including from a competitor's comparison page, rather than confirmed on the vendor's own pricing page in this pass, so treat them as indicative. The more important commercial point is what is being priced. Sully sells agents as staff substitutes and frames its market as the United States medical salary bill, so a per provider per month comparison against other scribes understates the intended scope: the buying decision is presented as replacing roles rather than adding a documentation tool. Buyers should price the agents they actually intend to run, and should establish oversight requirements for the triage and diagnosis suggestion capabilities separately from cost.

AI Health Index

An independent reference for evaluating AI vendors in healthcare. No vendor pays for inclusion, placement, or rating.

Index Status
Last index update
July 23, 2026
The AI Health Index is an editorial reference, not a regulatory body. Vendor data is verified against published sources and public regulatory filings. Figures labeled “Estimated” have not been confirmed by the vendor. See the Methodology page for evaluation standards and limitations.
© 2026 AI Health Index
3801 N Capital of Texas Hwy, Ste E240 · Austin, TX 78746