CombineHealth
Autonomous medical coding sold as one member of a named agent lineup rather than as a standalone engine. Amy is the coder; the same platform carries Mark for billing, Adam for accounts receivable, Rachel for appeals and Taylor for analytics. Founded 2022 in San Francisco by Sourabh Agrawal, chief executive, and Shikha Mohanty.
Amy reads encounter notes directly from the record system and assigns ICD-10, CPT, HCPCS Level II, evaluation and management levels, modifiers and hierarchical condition categories, with configurable coding grids and payer specific rules. Two design choices distinguish it. Decisions are explainable line by line, with rationale and evidence attached to each assignment, which is the same audit trail argument Nym Health makes. And the platform learns continuously from payer outcomes including denials, reimbursements and underpayments rather than from chart data alone, which is uncommon in this category and carries a governance question the company does not address.
Oversight is specified more fully than anywhere else in this lane. A four stage quality process runs model confidence scoring with uncertainty flags, secondary model validation, human review by certified coders, and a live compliance feedback stage tracking claims and corrections in production. The confidence threshold governing when work routes to a human is configurable by the customer rather than fixed by the vendor.
Published performance is 97.2 percent coding accuracy, an 85 percent claim automation rate and a 64 percent reduction in overall denials. The company published a parallel coding study across 1,000 emergency department charts comparing its output against expert human coders on the same charts, reporting 97 percent accuracy, turnaround roughly halved and five times more documentation gaps surfaced. That study was designed, run and reported by the vendor.
Evidence concentrates in emergency departments and anesthesia. Named customers include Medcor, Homeward, McFarland Clinic, SignatureCare ER and El Mirage ER, alongside anonymised references at a 500 bed hospital, a 400 bed emergency focused hospital and a 150 provider emergency physician group. Integration is claimed across twelve named record and practice management systems, the broadest coverage in the category, though no vendor marketplace listing was located and the company describes custom interfaces built per customer.
Data stewardship is the strongest in this lane: customers own their data, it is not sold, shared or repurposed, deletion or export can be requested at any time, and the company states it does not train on customer data without permission, using de identified data or obtained consent.
Two cautions for anyone quoting this record. The public material contradicts itself, asserting fully compliant and accurate outputs on the frequently asked questions page against the 97.2 percent figure published elsewhere on the same site, and that page still describes a scribe and a policy reviewer agent that no longer appear in the product navigation. And nothing about pricing is published anywhere.
Funding is a single institutional round of undisclosed amount, making this the earliest stage record in the category and the one most likely to need a status recheck within a year.
Capability Axes
An AI Health Index grade measures what a buyer can verify from public sources on the date shown. It is not a rating of how good the product is. A vendor can build an excellent system and grade low on an axis because it publishes nothing an outsider can check. How grades read
The models are the product and the company says so in its own framing: an artificial intelligence workforce, sold as named agents that perform work people previously did. Amy codes, Mark bills, Adam works accounts receivable, Rachel drafts appeals, Taylor analyses. Remove the models and there is no workflow layer, network or platform left that a customer would pay for.
One ambiguity belongs on the record because it bears on this grade and the company does not resolve it. The published quality process names a final review stage performed by certified coding experts, without stating whose experts they are. Elsewhere the material says the customer's coders review flagged cases, which points to the buyer's staff. If instead the vendor supplies that review layer at scale, part of what is sold is labour rather than software, and this grade would need revisiting. Ask who staffs the final review tier.
The oversight architecture is the most fully described in this lane and it would earn an A on its own. Four stages are named and distinguished: model self assessment producing confidence scores and uncertainty flags, secondary model validation, human review by certified coders, and a compliance feedback stage tracking claims and corrections in production. Critically, the confidence threshold that governs when work routes to a human is not merely disclosed but configurable by the customer, alongside workflow logic, payer rules and review requirements. No other vendor in this category lets the buyer set the autonomy boundary.
Explicit commitments accompany it: the models do not make irreversible decisions, outputs can be reviewed, approved, overridden and tracked as exceptions, and every action carries a traceable rationale.
What holds it at B is that the automation claims do not reconcile across the company's own pages. A claim automation rate of 85 percent, a reduction in manual coding effort of up to 85 percent, and a statement that human review is built into every step are three different propositions, and a buyer cannot tell from published material what share of charts reaches billing untouched. Add the absence of any contractual service level agreement on automation or accuracy, and a strong oversight design sits on an unclear autonomy figure. Ask what percentage of charts is submitted with no human touch at the default threshold.
The technique stack is enumerated rather than gestured at: natural language processing for unstructured notes, machine learning over coding patterns and payer outcomes, deep learning for pattern recognition, agentic automation for rule based steps, and explainable output as a stated design principle. Output scope covers evaluation and management levels, procedure and diagnosis codes, supply codes, modifiers and hierarchical condition categories. The company also states models can be fine tuned on a customer's historical data, coding grid and preferred guidelines.
Explainability is the substantive claim and it is described concretely: line by line rationale and evidence attached to each assignment, traceable back to the documentation that justified it.
Two problems hold it at B. No foundation model, model class or version is named behind the word proprietary. And the published material contradicts itself in ways that matter: the frequently asked questions state the continuous learning models ensure fully compliant and accurate outputs, which cannot be reconciled with the headline accuracy figure of 97.2 percent published on the same site. That page also still describes a scribe product and a policy reviewer agent that no longer appear anywhere in the product navigation, so the public account of what the platform contains is stale in at least one place.
The training data position is the best in this lane and the party enumeration is absent, which is why this sits at the top of the C band rather than higher.
On training, the company states plainly that it does not train on customer data without permission, that models are trained on de identified data or under obtained permission, that customers own their data, that it is not sold, shared or repurposed, and that deletion or export can be requested at any time. Every one of those is a question left blank by Fathom, Arintra and Maverick, and answering them is a genuine disclosure rather than a posture.
On who is in the chain there is nothing usable. No foundation model provider, model class or version is named. Two major cloud and model vendors appear as logos on the trust page with no accompanying statement, which indicates infrastructure relationships without disclosing their scope, and a logo is not an enumeration. No sub processor list was located. Given the product is described as generative and agentic, at least one external model provider almost certainly sits in the chain and is unnamed. Ask for the base model, the sub processor list, and what those two platform relationships actually cover.
The strongest evidence design in this lane, undercut by who ran it. The company published a parallel coding study across 1,000 emergency department charts, comparing its output against expert human coders on the same charts, reporting 97 percent accuracy, turnaround roughly halved, and five times more documentation gaps surfaced. A head to head design against the human baseline is what this category needs and what no competitor here has published.
Four further customer studies carry specific figures: eligibility related denials down 8 percent with verification 80 percent faster, 150 claims generated in minutes with billing errors down 32 percent and denials down 18 percent at an anesthesia billing team, and denials down 20 percent across more than 10,000 claims at a health center where 250 false denials were identified at 97.4 percent accuracy. Seven customer organisations are named by logo.
Held at B for two reasons. The parallel study was designed, run and reported by the vendor, with no independent audit, no protocol publication and no third party validation, so it demonstrates capability under conditions the vendor chose. And the testimonials are anonymised by role and facility size rather than attributed, which is weaker than the named attributed quotes XpertDox publishes. No research organisation coverage of the kind Fathom and Arintra have exists.
The strongest stewardship disclosure in the autonomous coding lane, and it answers every question the other records in this category leave open.
On data rights the company states that the customer owns all data, that it is not sold, shared or repurposed, that no third party data sharing occurs, and that deletion or export can be requested at any time. On model training it states that it never trains on customer data without permission and that models are trained on de identified data or under obtained permission. Those two paragraphs resolve the exact questions recorded as unanswered on Fathom, Arintra and Maverick.
On protection the controls are named individually rather than implied by a certification: encryption at rest and in transit using a named cipher, secure transmission protocols, single sign on, role based permissions, session management, interface security, data minimisation limiting collection to what the service requires, real time threat detection, automated incident response, a security operations function, vulnerability management, continuous monitoring, physical data center controls, and regular third party penetration testing. Audit logs cover model decisions, user actions and system access, and are exportable.
Graded A because the disclosure is specific, covers both protection and rights, and states a training position most vendors avoid entirely. What would make it unimpeachable is a formal retention schedule with periods, and independent assessment against a health specific framework rather than a general one.
The health privacy programme is enumerated as controls rather than claimed as a badge, which is the distinction this axis rewards. Encryption of patient data in transit and at rest, strict role based access controls, comprehensive audit trails logging all data access and modification, regular security audits and vulnerability assessments, mandatory staff training, and de identification of data used for model training and analysis are each stated individually.
The framing is also correct, which is not universal in this lane. The compliance badge is labelled as compliance rather than as certification, avoiding the category error XpertDox makes by presenting a framework with no certification scheme as a credential. Audit logs are stated to be exportable and available to a customer's compliance officer specifically for regulatory checks, which is a practical answer to how a covered entity would evidence its own compliance.
Held at B because two things are missing. No business associate agreement posture, template, negotiation stance or execution requirement was located, despite the company contracting with hospitals. And no health specific external certification exists of the kind two competitors in this lane hold, so the control programme is described in detail and assessed only against a general framework.
One credential, disclosed properly, plus the most substantial trust center in the lane.
The report type is stated as the second type, the one covering operating effectiveness over a period rather than design at a point in time, and the company states the report is available on request. Type plus an availability process is the full form of this disclosure and it passes the credential test cleanly, where Arintra's unqualified mention of the same framework does not.
The trust center is a standing page organised into data protection, access management, and monitoring and response, with controls enumerated inside each: a named encryption cipher, secure transmission, single sign on, role based permissions, interface security, session management, real time threat detection, automated incident response, a security operations function and vulnerability management. Regular third party penetration testing is disclosed separately, which almost nobody in this category publishes.
Graded A on the completeness of the disclosure rather than the count of badges. XpertDox reaches A by a different route, holding three certifications with a thinner trust center; this vendor holds one and documents it and the surrounding programme far more fully. Two gaps remain: no audit period, auditor or certification date is given, and there is no health specific certification of the kind two competitors hold.
No device pathway applies and none is claimed. Assigning billing codes from documentation is an administrative determination rather than a clinical one, so the absence of a clearance is correct and is not a gap in the record.
The exposure sits in claims submission, where codes are representations to a payer and error is governed by federal false claims enforcement, landing on the billing provider rather than the software vendor. The company's posture is well aimed at that risk: audit trails built to meet administrative audit standards, exportable logs for regulatory checks, and a documentation integrity function checking that notes support the codes billed.
One published statement cuts against it. The frequently asked questions assert that continuous learning ensures fully compliant and accurate outputs, which is not a claim any coding system can support and which contradicts the accuracy figure published elsewhere on the same site. In a domain where the provider carries false claims exposure, a vendor asserting perfect compliance is a statement a compliance officer should raise rather than rely on. Graded C rather than lower because the underlying regulatory position is correctly represented.
The instruments are real and well described. Every model decision carries traceable rationale, full logs of decisions, user actions and system access are captured and exportable, a compliance feedback stage tracks claims and corrections live in production, and the analytics agent tracks more than fifty performance indicators surfacing root causes and revenue leakage drivers. A buyer therefore has both the case level and the population level view.
The governance question specific to this vendor is created by its own differentiator and is not addressed anywhere. The platform learns continuously from payer responses, denial patterns, reimbursements and underpayments. A model optimised against payer acceptance is being trained toward codes that get paid, which is related to but not identical with codes that are correct. Those objectives align most of the time and diverge exactly where the money is, and a system rewarded for reimbursement outcomes has a structural gradient toward the coding choices that maximise it. Published results are consistent with either reading: denials down 64 percent overall is the intended effect and is also what the divergence would look like from outside.
Nothing published distinguishes them. No distribution of assigned billing levels against an expected benchmark, no breakdown by specialty, payer or physician, no bias or fairness testing, and no external audit of coded output. Held at B because the monitoring is genuine and no result from it is published. Ask what the learning objective optimises and how the loop is constrained against upcoding.
A stated performance level exists and the autonomy boundary is not only disclosed but adjustable, which is more than any competitor in this lane offers. Accuracy is published at 97.2 percent, the confidence threshold governing human routing is customer configurable, and a four stage quality process with a live compliance feedback loop tracks claims and corrections in production. The commitment that the models make no irreversible decisions, combined with override and exception tracking, means a buyer retains a mechanism to catch and reverse errors before they reach a payer.
None of it is recourse. No service level agreement, warranty, indemnity or remediation commitment was located. No accuracy breakdown by coding element exists, which matters because evaluation and management level assignment is harder and more often challenged on audit than procedure code selection. No denial or reversal rate for codes the system assigned is published as a vendor figure, though customer studies report denial movement.
One published statement actively weakens the position rather than strengthening it. The company asserts that its continuous learning ensures fully compliant and accurate outputs while publishing an accuracy figure below 100 percent on the same site. A buyer should establish which figure the contract references, because a marketing claim of perfection sits badly next to a measured rate in any dispute about a miscoded claim. Ask what contractual commitment attaches to the published accuracy figure.
The broadest named record system coverage in this lane by a wide margin. Twelve systems are published by name across both hospital and ambulatory segments, including the two dominant enterprise vendors and a long tail of practice management and ambulatory platforms. Against Arintra's three and Maverick's radiology only partners, that breadth is a genuine advantage for a buyer running a mixed estate or a physician group on a smaller platform.
What holds it at B rather than A is the absence of the two things that lift the other A grades in this lane. No vendor marketplace or program listing was located, so the integrations are asserted by logo rather than evidenced by an external credential of the kind Arintra and XpertDox both hold. And the mechanism list is thinner: interfaces, health messaging standards and custom interfaces are named, with no modern interoperability standard specified.
The word custom is doing real work. The company directs prospective buyers to contact it to discuss their specific system and integration requirements, and describes its team working with the customer's technical staff, which suggests per customer integration engineering rather than productised connectors. That is not a defect, and it does change the implementation burden a buyer should plan for. Ask which of the twelve are productised and which were built bespoke.
More is gestured at here than anywhere else in the lane, and none of it resolves to an answer a buyer could act on.
Data residency controls appear as a named item in the data protection set, which at least acknowledges residency as a dimension the product addresses. Two major cloud and model platform logos sit on the trust page, indicating infrastructure relationships. Physical data center controls are referenced in the security material.
None of that states where clinical documentation is processed or stored. No region is named, no tenancy model is described, no customer controlled or single tenant option is offered, and whether residency controls are configurable by the customer or a default the vendor sets is unstated. A logo indicating a cloud relationship is not a hosting disclosure, and this record does not treat it as one.
Graded C rather than B because the axis asks what the deployment model is, and after a dedicated pass across three published pages it remains unanswered. Ask for the processing and storage regions, the tenancy model, whether residency is configurable, and which of the two named platform relationships touches protected data.
Nothing about cost is published anywhere. A dedicated pass across the home page, the trust center and the frequently asked questions found no pricing page, no unit of charge, no range, no implementation or onboarding fee position, no minimum commitment, no trial terms, and no percentage cost reduction against a buyer's existing coding spend. The absence is recorded after looking rather than assumed.
The frequently asked questions cover integration, training, support, security and technology in detail and never reach cost, which is a choice rather than an oversight given the depth of everything else on that page. Ongoing support is described qualitatively as dedicated technical support, account management and regular performance reviews, with no statement about whether any of it carries separate charge.
The contrast within the lane is stark and worth recording: this vendor publishes the most detailed security and data stewardship material of any competitor and the least commercial information. Ask for the pricing mechanism, whether charge is per chart, per claim or per agent, how the four agents beyond coding are priced, and what happens commercially to charts routed to human review.
Coverage is broader than the outpatient framing the company's earlier material suggested, and the customer roster is the evidence. Named and described deployments span emergency departments heavily, including two emergency centers, a 150 provider emergency physician group and a 400 bed emergency focused hospital, alongside a 500 bed hospital, a regional hospital, a multi specialty clinic, an occupational health provider, a value based care organisation, a community health center and anesthesia billing teams.
The emergency department concentration is the notable pattern. The flagship comparative study was run on emergency charts, and several references are emergency settings, which suggests genuine depth in a high volume specialty where documentation is comparatively structured and coding is repetitive. Anesthesia appears as a second area with its own published workflow.
Coding scope covers the full element set rather than a subset. Specialty pages exist beyond these. Held at B rather than A because the breadth claimed across the wider revenue cycle is not matched by evidence in every setting named, and because nothing addresses coding regimes outside the United States. A buyer outside emergency or anesthesia should ask for a reference in their own setting.
Pricing
Vendor-published figures are labeled as such. Figures labeled “Estimated” are derived from third-party sources and have not been confirmed by the vendor.
| Entry Price | Pricing Basis | BAA Tier | Implementation | Source |
|---|---|---|---|---|
|
Not published
|
Not disclosed. No pricing page exists and no unit of charge is described. The agent lineup makes the unit question sharper than usual, since it is unstated whether the coding agent is licensed independently of the billing, accounts receivable, appeals and analytics agents. | Not disclosed. No business associate agreement posture, template, negotiation stance or execution requirement was located, despite the company contracting with hospitals and processing complete charts. The privacy compliance badge is correctly framed as compliance rather than certification, unlike XpertDox, and the underlying control programme is enumerated in detail on the trust center. | Not disclosed. The company states its team works with customer technical staff on integration and directs prospective buyers to contact it to discuss requirements, which suggests per customer integration work, and no statement addresses whether that work carries separate cost. Staff training is described as a programme tailored by role with no fee position stated. | Vendor Published |
Nothing about cost is published anywhere. A dedicated pass across the home page, the trust center and the frequently asked questions returned no pricing page, no unit of charge, no range, no implementation or onboarding fee position, no minimum commitment, no trial terms, and no percentage cost reduction against existing coding spend. The absence is recorded after looking rather than assumed.
The frequently asked questions page is otherwise unusually thorough, covering integration, staff training, ongoing support, security and technology stack in detail, and it never reaches cost, which reads as a deliberate choice rather than an oversight. Ongoing support is described qualitatively as dedicated technical support, account management and regular performance reviews, with no statement on whether any of it is separately charged.
The contrast within this lane is worth recording: CombineHealth publishes the most detailed data stewardship and security material of any autonomous coding vendor in the index and the least commercial information. Pricing is further complicated by the agent model, since coding is one of five named agents and nothing indicates whether they are licensed separately, bundled, or priced as a platform.
Ask for the pricing mechanism, whether charge is per chart, per claim or per agent, how the four non coding agents are priced, and what happens commercially to charts the configurable confidence threshold routes to human review.