HealthOrbit AI
HealthOrbit AI offers transcription, clinical note generation, automated ICD-10 and CPT coding and what it calls a doctor co pilot providing diagnostic recommendations during consultations. Its described workflow is upload driven, with audio files submitted to the platform and transcripts returned, alongside claims of multilingual support and EHR interoperability. This record is deliberately thin and the thinness is the finding.
Almost nothing about this vendor could be verified: no named customer, no funding, no deployment, no third party assessment and no independent coverage was located, its accuracy claim appears as both above 90 percent and up to 95 percent in its own materials, its pricing is described as free in one place and available only by demo elsewhere, and its profile on a major software review site has been dormant for over a year. Buyers encountering the name should treat the record as an inventory of unverified claims rather than an assessment, and should establish basic commercial and security facts before any pilot.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
Transcription and note generation are the entire product with no services or platform business underneath. This is the one axis that can be graded confidently, because it follows from what the product is rather than from anything the vendor has evidenced.
No review gate, sign off step, confidence threshold or human oversight mechanism was located anywhere in the material, and the second pass makes that omission more serious rather than less.
The product is described as generating diagnosis and procedure codes automatically, providing diagnostic and treatment recommendations during consultations, and aggregating information from records, laboratory results and imaging to surface potential risks. Three different kinds of output, each with different consequences, and no published account of what a clinician confirms for any of them.
The reasoning that keeps clinical decision support outside device regulation depends almost entirely on this axis. It assumes a professional can review the basis of a recommendation and reach an independent judgement before acting. A vendor making diagnostic recommendations and publishing nothing about how the clinician engages with them has left unaddressed the single mechanism its regulatory position rests on.
The questions are specific and answerable. What does the clinician see alongside a diagnostic suggestion, and can they see what it was derived from. Is a suggestion presented before or after the clinician forms their own impression, since a recommendation shown first anchors judgement in a way one shown second does not. Are generated codes committed automatically or held for confirmation. What happens when the system has low confidence.
The last of those matters most for a product claiming to improve diagnostic accuracy, because a system that is right more often than a clinician on average will still be wrong sometimes, and the value of the oversight step is entirely in catching those cases.
Ask what the clinician confirms, for each output type separately.
Accuracy is claimed twice at different values in the vendor's own materials, above 90 percent in one place and up to 95 percent in another, with no methodology, reference standard, denominator or date attached to either. The up to construction is unfalsifiable in any case. No model card, named models or evaluation protocol located.
Nothing identifies any party in the chain: no model or model family, no foundation model provider, no hosting arrangement and no sub processor list was located in two passes. What distinguishes this record is how much content the unnamed chain holds.
The platform ingests from record systems, laboratory results and imaging alongside transcriptions and builds consolidated historical records on top, so the store contains clinical data about patients whose encounters this product never recorded, and derived analysis built from it. A patient can therefore be represented inside this system without ever having been in a room where it was running. Two questions follow and neither is addressed.
Whether ingested laboratory and imaging data persists or is fetched on demand determines whether that store is a copy or a view, and the answer changes what a breach or a departure would expose. And whether any of it, or the derived records, contributes to model development is unstated in either direction, which matters more for a product making diagnostic claims because the system was trained on something. Ask for retention per data type, the deletion path on termination, and a sub processor list.
No product specific evidence of any kind was located: no study, no controlled evaluation, no accuracy benchmark, no third party rating, no named customer and no documented deployment. The figures the vendor publishes are market statistics rather than results, such as clinicians losing 65,000 US dollars a year to wasted time and two hours saved daily, which describe the category's premise rather than this product's performance. Its profile on a major software review site has been inactive for over a year, which is not evidence of anything on its own but is consistent with the absence of everything else.
No retention schedule, de identification practice, encryption detail or training use statement was located.
The earlier assessment noted that the workflow involves uploading recorded conversations, which makes the absence more consequential rather than less, and that stands. Uploaded audio is a file that exists somewhere until something deletes it, which is a different exposure from streamed audio discarded on completion, and several peers in this category distinguish the two explicitly.
The second pass widens what is held. The platform is described as aggregating from record systems, laboratory results and imaging alongside transcriptions, and as generating consolidated historical records that surface potential risks. So the store contains ingested clinical data about patients whose encounters this product never recorded, and derived analysis built on top of it.
That should be asked about separately from encounter audio. Establish what is retained after a note is produced, whether ingested laboratory and imaging data persists or is fetched on demand, and what happens to the generated historical records when a practice leaves.
The training question is unanswered in either direction and matters more for a product making diagnostic claims. A system asserted to improve diagnostic accuracy was trained on something, and a buyer should establish whether their own patients' data joins that corpus. Peers now state a position plainly, some committing never to train on clinical content and at least two operating an explicit permission gate.
Ask for the retention schedule per data type, the deletion path on termination, and the training position in contract language rather than marketing.
HIPAA compliance is asserted, but the assertion located appears in the vendor's own comparison blog referring to itself in the third person as one of a group of compliant vendors, rather than as a direct product commitment on a product or trust page. No business associate agreement terms were located. Graded C rather than B because the form of the claim is weaker than a plain vendor statement.
No named or dated attestation, no report of either type and no trust centre were located.
The vendor publishes an article on its own site asking whether medical records are safe with AI scribes, setting out the privacy and security risks such tools introduce and the measures needed to protect patient information. That is accurate and it is the fifth instance this index has recorded of a vendor authoring the argument it does not answer for itself. The pattern is now consistent enough to state as a general observation: in this category, publishing security guidance correlates poorly with holding security evidence.
The scope of what an examination would need to cover is wider than a scribe's. The platform is described as ingesting not only encounter audio but record content, laboratory results and imaging, and as generating consolidated historical records. That is a substantially larger data estate than a documentation tool holds, and the questions an assessment answers, how access is scoped, how ingestion is authorised, what an administrator can restrict, become correspondingly more important.
The published workflow adds a specific item. Recorded doctor patient conversations are described as being uploaded to the platform, so audio files exist and move, rather than being streamed and discarded. Where and how those files are stored is exactly what an attestation examines.
Nothing here suggests controls are absent. This grade records what a counterparty can verify before contracting, which is nothing.
Ask what external testing has been performed, and specifically how access to ingested record, laboratory and imaging data is scoped and logged.
No clearance claimed and none located. The flag raised in the earlier assessment is confirmed and the second pass finds a claim that goes further than diagnostic recommendation, into quantified clinical efficacy.
The product's co pilot is described as a diagnostic assistant providing treatment recommendations during consultations, and as raising the diagnostic accuracy rate by twenty five per cent.
That is not a documentation claim. It is an assertion that using the software makes clinicians diagnose more correctly, by a stated amount. Intended purpose is the test that determines whether software falls inside the device framework, and a product whose stated purpose includes improving diagnostic accuracy has described itself in the terms the framework uses. No regulatory assessment, intended use statement or disclaimer addressing that was located.
The evidentiary problem is separate and equally serious. A quantified improvement in diagnostic accuracy is a clinical efficacy claim, and efficacy claims require a study: a comparator, a reference standard for correctness, a defined population and a method. None accompanies it. A figure of that kind either rests on evidence the vendor has not published, or it does not rest on evidence at all, and a buyer cannot tell which.
Where the claim appears matters too. It sits in blog articles on the vendor's own domain framed as comparisons of scribe products, rather than on a product page alongside supporting material. Claims of clinical performance published as search optimised comparison content, without evidence, is the weakest form this can take.
Ask for the study behind the figure, and for a written intended use statement.
No fairness statement, subgroup analysis or accent and dialect performance disclosure was located, despite multilingual transcription being marketed.
The second pass adds a pattern that bears on how every unevidenced claim from this vendor should be read, and it belongs on this axis because this axis is about what a vendor can demonstrate rather than assert.
The vendor publishes quantified performance figures across several claims: a twenty five per cent improvement in diagnostic accuracy, a forty per cent reduction in coding mistakes. Neither is accompanied by a study, a methodology, a comparator or a population. They appear in blog articles on the vendor's own domain framed as industry comparisons rather than on product pages with supporting material.
A vendor willing to publish precise figures without evidence for the things it wants to claim is not a vendor whose silence on subgroup performance can be read as considered restraint. The absence here is not the modest silence of a company that has not yet run the evaluation; it sits alongside a habit of quantifying what suits.
The multilingual claim is the specific version. Marketing multilingual transcription is a claim about performance across speaker populations, which is what this axis measures, and no language list, per language accuracy or evaluation is published.
One setting reference makes it sharper. Use in emergency departments is referenced in the vendor's content, and that is the environment with the widest range of language, distress and acuity.
Ask for the methodology behind each published figure, and for accuracy by language and accent.
Two passes located no warranty, indemnity or remediation commitment, and the accuracy position is worse than an absence because the vendor's own materials disagree with each other. Accuracy is claimed above 90 percent in one place and up to 95 percent in another, with no methodology, reference standard, denominator or date attached to either, and the second figure uses the up to construction that is satisfied by any result at all.
A buyer cannot rely on a number when the vendor publishes two, and cannot rely on either when neither is defined. The stakes are raised by what the product claims to do. This is not documentation alone: the platform aggregates record system data, laboratory results and imaging alongside transcription, generates consolidated historical records, and is positioned as surfacing potential risks and improving diagnostic accuracy.
A system that asserts an effect on diagnosis is making a clinical claim, and a clinical claim with no measurement behind it and no commitment attached is the weakest position a product in this index can occupy. It also widens who can be harmed, because the consolidated record covers patients whose encounters this product never captured. Ask which of the two accuracy figures the vendor stands behind, what it measures, and what the vendor commits to when a surfaced risk is wrong or a real one is missed.
Integration and interoperability are claimed in general terms, described as enabling generated notes to reach the record and as aggregating information from record systems, laboratory results and imaging. No named system, integration architecture, standard, certification or write back mechanism was located, and the earlier assessment's conclusion holds: a claim of integration without a single named system cannot be graded.
The second pass makes the gap wider rather than narrower, because the claimed data flow now runs both directions and reaches beyond documentation. Writing a note into a record is one capability. Reading laboratory results and imaging out of connected systems, aggregating them into a consolidated history and generating risk flags is a substantially deeper integration requiring broader access, and no mechanism for it is described at all.
That asymmetry between what is claimed and what is evidenced is the finding. Ingesting results and imaging from a record system requires either a permissioned interface with defined scopes, a data feed the institution configures, or access credentials the vendor holds. Those have very different security and audit properties, and a buyer cannot determine which applies.
So the questions are the basic ones and none has a published answer. Which record systems are supported. By what mechanism. Is it bidirectional. What scopes does it request. Whose identity do its reads and writes carry in the audit log. Who configures and maintains it.
A buyer should treat integration as unevidenced until demonstrated in their own environment, and should specifically confirm what the product can read as distinct from what it can write.
No hosting region, residency option or subprocessor detail was located, and nothing establishes which model service processes encounters or what it retains. Delivery is cloud, with tablet based capture described for in clinic and telehealth use.
The upload based workflow gives this axis a specific shape. Where recorded conversations are uploaded rather than streamed, audio files exist as objects in storage, and the questions become concrete: which region holds them, for how long, under whose encryption keys, and whether an administrator can see what has been uploaded.
The ingestion claims widen it further. A platform described as drawing from record systems, laboratory results and imaging is holding a broader clinical estate than encounter audio, and the residency question covers all of it rather than only recordings.
The model provider question is unanswered and matters more than usual here. A product asserting diagnostic recommendations is either running its own clinical model or building on an external foundation model with clinical prompting. Those are materially different propositions for a buyer assessing where clinical suggestions come from, what they were trained on, and who else sees the content. Nothing published distinguishes them.
Telehealth use adds a further consideration. Where a consultation is remote, the patient may be in a different jurisdiction from the clinician, and the recording captures a person who is not in the clinic. Establish how the product handles cross border telehealth encounters.
Ask for the hosting region, the retention of uploaded audio files, the subprocessor list, and which model provider generates the clinical output.
No published rate, and the available statements conflict. The product is described as a free AI medical scribe in one place while every route to it is a demo booking, and a major software review site records that pricing details are not available.
A free claim without terms, alongside no published rate, gives a buyer nothing to plan against and is arguably less useful than saying nothing. Establish what free covers, for how long, for how many users, and what happens at the boundary.
No specialty count, specialty tuning, note format list or supported language list was located.
The earlier assessment made the distinction that matters and it holds after a second pass. Multilingual support and use in high pressure settings such as emergency departments are referenced in the vendor's general blog content rather than stated as product capability. That difference is not pedantic. A capability statement is something a buyer can hold a vendor to and a supplier can be asked to evidence. A passing reference in an article about the industry is neither, and this vendor's content mixes the two throughout.
So a buyer cannot currently establish the basic scoping facts: which specialties the product has been used in, which note formats it produces, which languages it handles, or whether templates exist for their own documentation requirements.
That is a more consequential gap for this product than for a narrow scribe, because of what it claims to do. A system offering diagnostic and treatment recommendations is making specialty dependent judgements, and the appropriate suggestion in primary care differs from the appropriate suggestion in oncology or psychiatry. Establishing which specialties the clinical layer has been developed and tested for is therefore a safety question rather than a fit question.
The emergency department reference deserves specific confirmation for the same reason, since that setting combines the hardest acoustic conditions with the highest consequence of a missed diagnosis.
Ask which specialties are supported and evidenced, which languages, and whether the diagnostic layer behaves differently by specialty or applies one model everywhere.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Conflicting. Described as free in vendor material, demo request only in practice, no published rate.
|
Not disclosed. Access is by demo booking. | Not retrieved. HIPAA compliance asserted only indirectly in the vendor's own blog. | Not published. | Vendor Published |
Conflicting and unusable as published. The product is described as a free AI medical scribe in vendor material while every access route is a demo request, and an independent software directory records that pricing is not available.
Before any pilot, establish the basics that could not be verified here: whether the service is genuinely free and on what terms, which EHRs are actually supported and by what mechanism, what happens to uploaded conversation audio, and on what basis the product offers diagnostic recommendations.