MEDILIT
AI scribe and clinical document generator built by a team the company describes as doctors, engineers and support staff, led by cofounder and chief executive Naeim Abedi. The product generates clinical notes in real time during the consultation and extends into letters, referral documents, patient instructions, reports, plans and custom document types, with the stated design goal of handling long and complex consultations consistently rather than short simple ones. Two architectural decisions make this record worth carrying, and both are uncommon in the small vendor tail. The first is a privacy design rather than a privacy policy.
MEDILIT states that transcripts are redacted before any analysis, with names, dates of birth, numbers and email addresses replaced by random values before the content is sent for processing, so that only clinically relevant information reaches the generation layer and no personally identifiable information is linked to the data being analysed. That is de identification as an engineering step rather than a promise about handling, and it belongs in the same class as the small number of vendors in this index that removed a risk by design instead of governing it. The company also states that data is encrypted in transit, at rest, and after processing.
The second is that MEDILIT names its architecture: a multi agent system with a self auditing layer intended to catch omissions and maintain consistency across long consultations. Very little in this lane will describe an internal checking mechanism at all. A published AMD case study documents the product running on AMD Ryzen Embedded 8000 Series processors, which is a real named technology partnership and raises a question worth asking directly, since embedded silicon points toward local or edge inference rather than pure cloud processing.
Pricing is published as a two tier Standard and Premium ladder with the gating disclosed and a genuinely open trial that grants full premium features with no payment details taken. Prices are quoted excluding GST, indicating the company sells into a GST jurisdiction rather than the United States, which matters because the HIPAA framing this index applies to US vendors may not be the relevant regime here.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
Capture, redaction, note generation, the self auditing pass, and letter and document generation are all model output, delivered through a multi agent architecture the company describes as its core technology. No human scribe tier, no services layer and no host platform that the AI is a feature of.
The only checking mechanism described is the system auditing ITSELF. No human review gate, sign off step, confidence threshold, abstention behaviour or escalation path was located anywhere in the retrieved materials. This index has ruled on this pattern before, holding that a machine confidence or quality primitive is not an oversight mechanism: SimboAlphus earned transparency credit for exposing a per note quality score and was still graded down on autonomy because a score says nothing about default behaviour or whether a human is required.
The same applies more strongly here, because a self audit run by the same system that produced the output is an internal consistency check rather than independent verification. The likelihood is that clinicians do review notes before use, since that is how every product in this category works in practice, but the vendor does not say so and this index grades what is published.
Graded B for describing how the system is built, which most of this tail will not do. MEDILIT names a multi agent architecture with a self auditing layer intended to catch omissions and hold consistency across long consultations, and separately names its compute platform through a published AMD case study documenting the product running on AMD Ryzen Embedded 8000 Series processors. Naming both the agent design and the silicon is more disclosure than almost anything else at this scale offers.
Held at B because no model card, named language model, accuracy figure, error rate or evaluation methodology exists for any of it, and because the product page headline calls it the most accurate AI medical scribe without publishing a single number to support the superlative. A self auditing claim with no measured catch rate is an architecture description, not a result.
Two disclosures put this above most of the tail, and one of them bounds the chain structurally rather than describing it. The structural one first: transcripts are redacted before any analysis, with names, dates of birth, numbers and email addresses replaced by random values before content is sent for processing, so identifiers never reach the generation layer.
Whatever parties sit downstream therefore receive content with the direct identifiers already removed, which limits what the chain can expose rather than promising that it will not. Encryption is stated in transit, at rest and after processing, which is a more complete formulation than the usual pairing.
The second is unusual for its specificity: the compute platform is named through a published case study identifying the processor family the product runs on, and the agent architecture is described as multi stage with a self auditing layer. Naming the silicon is more disclosure than almost anything else at this scale offers. Held below the top grade because the model layer itself is unnamed, with no language model, provider or version identified, and no sub processor list was located.
One limit of the redaction claim deserves a question: identifiers that appear conversationally rather than as structured fields are the hard case for any de identification pipeline, and nothing published states how the redaction performs against them. Ask that, and ask whose language model sits behind the agents.
No study, controlled evaluation, accuracy benchmark, third party clinical assessment, named customer or deployment count located. The AMD case study is a genuine externally published document and is more than most vendors at this scale have, but it is a technology partner marketing artifact rather than clinical evidence: it describes the solution as making note taking more efficient and allowing doctors to save time, with no measurement, cohort or comparator. Treat it as confirmation that the product exists and runs on named hardware, not as evidence of performance.
The strongest privacy architecture found in the long tail of this wave, and it is a design decision rather than a policy statement.
Transcripts are redacted before any analysis. Names, dates of birth, numbers and email addresses are replaced with random values before content is sent for processing, so identifiers never reach the generation layer and no personally identifiable information is linked to the data being analysed. That removes a risk by construction, in the same way this index credited voize for on device inference and Veradigm for claiming an audio artifact that never exists. Encryption is stated in transit, at rest, and after processing, which is a more complete statement than the usual in transit and at rest pairing.
Held at B rather than A because three questions central to this axis are unanswered: whether audio is retained and for how long, whether customer content is used to train models, and how robust the redaction is against identifiers that appear conversationally rather than as structured fields, which is the hard case for any de identification pipeline. Scribeberry holds the reference position on this axis because it answers all of those explicitly.
The earlier assessment inferred from pricing quoted excluding goods and services tax that this vendor sells outside the United States. That inference is now confirmed directly. The company is Australian, states that all processing happens on servers located within Australia, and describes compliance with Australian privacy law and healthcare standards.
So this axis does not read as an absence. The health privacy rule is not the governing regime for this vendor, and its absence is expected rather than a gap. What governs instead is the Australian privacy framework and its health specific principles, which impose their own obligations on collection, use, disclosure, storage and cross border transfer, and the vendor has answered the hardest of those by committing to onshore processing rather than leaving location unstated.
That commitment is worth more than a general compliance claim, because residency is the specific thing Australian buyers are told to check and most vendors in this category do not state it plainly.
One clear instruction for a United States buyer, which is the reason this axis stays open rather than being closed entirely. Do not assume a business associate agreement is available. A vendor operating solely under a foreign privacy regime, with onshore processing as its stated architecture, may have no mechanism to offer one, and using the tool on protected health information without it is itself a violation regardless of how secure the product is. Establish whether the vendor will sign, and where a United States customer's data would be processed, before any trial.
Nothing located suggests a United States offering exists at present.
A second pass located no named or dated certification, no audit report, no penetration testing statement and no trust centre.
What is published is a set of described controls, and they are more specific than the tier norm. Voice recordings are stated never to be stored, with speech converted directly to text in real time rather than captured to a file. Transcriptions are encrypted and stated to be permanently removed once the note is generated. Data is stated to reside on Australian cloud infrastructure. Access is described as restricted to the treating clinician and authorised clinic staff, logged and audited. Audio processing is described as automated, with no human listening.
Those are the right commitments and several are checkable in principle. They remain descriptions of design rather than examinations of it, and the earlier assessment made the point that matters most: a pipeline that claims never to store audio and to delete transcriptions on note generation is precisely the kind of claim an independent assessment exists to test. The stronger and more specific the architectural claim, the more an attestation would add, because there is more to verify.
The segment context matters here. Australian buyer guidance in this market routinely tells practices to check for ISO 27001 or SOC 2 before adopting a scribe, and at least one competitor in the same market holds one. So the certification is a published purchasing criterion in this segment rather than an enterprise nicety, and its absence is visible to buyers being told to look for it.
Fairness is due to a company offering a free tier at this scale. Ask for whatever external testing has been performed, and for the report if one exists.
No device pathway attaches and none is claimed. This axis does not read as an absence, for two reasons that reinforce each other.
The first is scope, and the earlier assessment established it: no diagnostic suggestion, no coding automation and no order drafting were found. The decision adjacent surface that puts most of this lane on a software as a medical device watchlist is simply not present. In the Australian market that restraint is more pointed than it looks, because peers there do offer billing item suggestion against the national benefits schedule, so this vendor has declined a feature its competitors ship.
The second is what the vendor does publish, and it is rare enough to carry the grade. Alongside material for clinicians there is an explanation written for patients: what the tool is, that it captures the conversation and writes the notes, that only the treating clinician and authorised staff can see the output, that access is logged, and that audio is processed automatically without anyone listening. It states plainly that use is entirely voluntary, that declining will not affect the quality of care received, and that a patient may change their mind at any point.
In Australia the obligations on an ambient scribe come from privacy law, the therapeutic goods regulator's guidance on digital scribes and medical defence organisation positions, and all of them converge on informed patient consent. Peers in this category supply consent wording to clinicians. This vendor addresses the patient directly, in language a patient can read, including the right to refuse without consequence. That is the harder half and almost nobody does it.
The gap to track is market expansion, since none of this transfers to a jurisdiction treating ambient documentation as a regulated device.
One real governance positive and one complete gap. The positive is structural: choosing to strip identifiers before the analysis layer is a data governance decision taken at design time rather than a control bolted on afterwards, and it limits what the vendor itself can see.
The gap is that no fairness statement, subgroup analysis, accent or dialect performance disclosure, or model limitations document was located, while the product markets itself on accuracy and on handling complex and lengthy consultations. Notably absent, and worth recording as a positive contrast, is any revenue, coding intensity or reimbursement optimisation language, which keeps this vendor off the coding gradient this index tracks across much of the rest of the lane.
Two passes located no accuracy or error figure, no published limitations and no warranty, indemnity or remediation commitment, while the product page headline calls this the most accurate medical scribe available without publishing a single number to support the superlative. A comparative claim with no measurement of the product or of any competitor asserts a ranking that rests on nothing a reader can check.
The architecture includes a self auditing layer intended to catch omissions and hold consistency across long consultations, and that is where the disclosure gap bites hardest rather than merely being a missing statistic. A self audit is a claim about error correction, so the only thing that would make it meaningful is a catch rate: how many omissions it finds, how many it misses, and how a clinician would know when it has failed.
Without that it is an architecture description presented as a result, and it may lead a clinician to review less carefully on the understanding that something else is checking. Nothing published states what the self auditing layer surfaces to the user, whether it flags uncertainty on the note itself, or what happens when it disagrees with the primary output. Ask for the catch rate, and for what the clinician sees when the audit layer fires.
The earlier assessment recorded that no integration mechanism, named system, standard or transfer workflow was located, and that where the generated documents end up is a first order question. The second pass finds a partial answer and it is less reassuring than a partial answer usually is.
The product is a browser based application requiring no local installation. On integrations the vendor states that it aims to partner with all Australian practice management software providers and will update users as each is finalised. That is a roadmap rather than a capability. No system is named as connected today, no interface or standard is identified, and a buyer cannot determine from published material whether their own software is supported.
What makes this more than an ordinary gap is that the vendor's retention policy depends on the integration existing. Transcriptions and outputs are described as temporarily stored to prevent data loss until successfully transferred to the record system, after which nothing remains on the vendor's servers. Retention is therefore defined by a workflow event rather than by a clock, which is an elegant design when the transfer is automated and confirmable.
Where no integration exists, that event has no clear definition. If a clinician copies text out of a browser by hand, it is not obvious what confirms transfer, when the deletion trigger fires, or what happens to output that is never marked as transferred at all. The retention commitment and the integration roadmap are load bearing for each other, and only one of them is delivered.
Ask which systems are connected today, by what mechanism, and precisely what triggers deletion for users whose system is not among them.
Partially described, with one detail worth chasing. The company refers to its own secure cloud servers, so a hosted service is the baseline. But a published AMD case study documents the solution running on AMD Ryzen Embedded 8000 Series processors, and embedded silicon points toward local or edge inference rather than pure cloud processing.
If clinical audio is processed on hardware inside the practice, that would place MEDILIT alongside voize in the small group of vendors where the data physically does not leave the site, which would be a materially stronger position than the grade here reflects. It is graded C rather than higher because the vendor itself does not describe a deployment model, and no hosting region, residency option or sub processor detail was located. Ask directly where inference runs.
A published two tier ladder with the gating stated, which is better than most of this lane. Standard and Premium both include the AI scribe plus letter and custom document generation, and Premium adds unlimited AI document generation and additional AI assistant tools, so a buyer can see exactly what the upgrade buys before talking to anyone.
The trial terms are unusually clean and stated plainly: the trial grants full access to all premium features, and no payment information is required to start, which removes the card capture pattern common at this tier. Prices are quoted excluding GST, so the buyer should confirm the tax inclusive figure for their jurisdiction. Held at B rather than A because enterprise or multi clinician terms are not addressed and no per seat versus per practice basis is stated.
General clinical use with breadth in document TYPES rather than in specialties. The output range is genuinely wide for a small vendor: consultation notes, letters, referral documents, patient instructions, reports, plans and user defined custom documents, with the design explicitly aimed at long and complex consultations rather than short visits. What is missing is specialty depth.
No specialty is named, no specialty specific instrument, form or scoring system is referenced, and no care setting beyond general consultation is addressed. On this index's test, naming the instrument is what distinguishes real domain work from a general model with a template, and nothing here does that.
Compared With
Each comparison carries a written verdict, the buyer conditions that favor each vendor, and a graded side by side. Pairs that cross a category boundary are grouped separately, and their verdicts state where the boundary sits rather than manufacturing a head to head.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Two published tiers, Standard and Premium, with prices quoted excluding GST. Free trial with full access to all premium features and no payment information required.
|
Two tier subscription. Standard and Premium both include the AI scribe plus letter and custom document generation. Premium adds unlimited AI document generation and additional AI assistant tools. Enterprise or multi clinician terms are not addressed. | No HIPAA statement or BAA reference located. GST pricing indicates a non US market where HIPAA may not be the governing regime. US buyers should confirm availability before proceeding. | None published. | Vendor Published |
Three questions, in order of importance. First, where does inference actually run: the vendor describes secure cloud servers, but a published AMD case study documents the product on AMD Ryzen Embedded 8000 Series processors, and if audio is processed on hardware inside the practice this vendor sits in a much stronger data residency position than the grade currently reflects.
Second, how robust is the pre processing redaction against identifiers spoken conversationally rather than appearing as structured fields, since a patient naming their own street or employer mid sentence is the hard case for any de identification pipeline, and the strength of this whole record rests on that step working. Third, where does the note actually go, because no EHR integration, API or export mechanism is described anywhere and that determines the real workflow cost.
Also confirm audio retention and whether customer content trains models, neither of which is addressed, and note that the product markets itself as the most accurate AI medical scribe while publishing no accuracy figure at all.