Microsoft Dragon Copilot vs TORTUS
The two vendors in this category that take error seriously, and they treat it differently. Microsoft names its failure modes in a transparency document and asks the clinician to review the draft. TORTUS filters before the clinician sees anything: the Shell verifies every generated statement against the consultation and removes what it cannot support, then publishes the rate, 92.7 percent of detected hallucinations removed, with a 75 percent reduction in major hallucinations published in npj Digital Medicine. It also co ran a nine site NHS evaluation across more than 17,000 patients with public funding and independent assessment. Naming a risk is good practice; measuring and filtering it is better. The constraint is geography. TORTUS is deep in the NHS and unevidenced outside it, so for a United States health system Dragon Copilot is the practical answer and TORTUS is the benchmark to hold it to.
- Availability and reach. Native Epic embedding, SMART on FHIR, a partner developer kit for other record systems, and clinician availability across ten countries, against a product whose evidence and integrations are concentrated in the NHS estate.
- Role based products for physicians, nurses and radiologists, with flowsheet capture at the bedside, plus multilingual multi party capture and generated referral and patient letters.
- Enterprise contracting your organisation already has, with a central trust portal, published privacy and security documentation and a support relationship that does not require a new vendor onboarding.
- It filters hallucinations before the clinician rather than relying on the clinician to catch them, and publishes how well that works: 92.7 percent of detected hallucinations removed as a live platform metric, and a 75 percent reduction in major hallucinations in a peer reviewed journal.
- The evidence base is prospective, multi site and publicly funded rather than vendor commissioned, spanning nine clinical sites and more than 17,000 patients with NHS England backing and independent evaluation, plus a paediatric study at Great Ormond Street.
- Accent performance was tested by the deploying hospital, not the vendor, and the hospital states publicly that transcription held accurate across a wide variety of accents. That is the failure mode every speech product carries, measured by the party with the regulatory exposure.
Side by Side
| Axis | M Microsoft Dragon Copilot |
T TORTUS |
|---|---|---|
| AI Centrality | ||
| Autonomy and Oversight Model | ||
| Model and Technology Transparency | ||
| Clinical and Operational Evidence | ||
| AI Safety and PHI Stewardship | ||
| HIPAA and BAA Posture | ||
| Security Certifications and Trust Center | ||
| FDA and Regulatory Status | ||
| AI Governance and Bias Disclosure | ||
| EHR and Interoperability Depth | ||
| Deployment Model and Data Residency | ||
| Commercial Transparency | ||
| Setting and Specialty Coverage |
Related comparisons
Other published head to head assessments involving these vendors or their closest peers. The full set for this category is on the Ambient Scribes page.
TORTUS holds a UKCA Class I device registration, which is self certified by the manufacturer rather than reviewed by a regulator, and its selection for the United Kingdom regulator's AI Airlock sandbox is engagement with a pathway rather than approval under one. Its HIPAA alignment is asserted across third party listings with no business associate agreement template or scope statement located, so a United States buyer cannot verify it before contracting.
Neither vendor publishes a rate card. Dragon Copilot's own randomized result, 1.7 percent reduction in time in note, was not statistically significant, and TORTUS reports an average saving of four minutes per consultation from early adopters rather than from a controlled comparison.