Comparisons

Medical coding AI compared

A dated capability matrix for AI medical-coding tools — what each vendor documents about autonomy and human review, set beside the peer-reviewed evidence on code-assignment accuracy. No winner declared. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • This is a capability matrix, not a ranking: attributes side by side, every cell cited, no single best tool named.
  • The peer-reviewed reality check is stark — one benchmark found large language models highly error-prone at mapping medical codes, and in another the best of five LLMs reached only 45.3% accuracy on ICD-9 prediction, with none meeting clinical standards.
  • That is why every deployed system pairs a model with rules, curated code sets, and human review of complex cases; the products differ mainly in how much they automate before a human steps in.
  • Vendor accuracy figures (for example, a stated 70% reduction in manual coding) are self-reported and lack published independent audits — read them as claims, not findings.
  • Coding automation is administrative revenue-cycle software and generally sits outside the FDA device pathway. Autonomy models, code coverage, and routing change — confirm against the vendor.

AI that turns a clinical note into billing codes has moved from suggestion engine to, in some products, an autonomous step that assigns and submits codes on its own. This page compares the leading approaches as a capability matrix: it sets their documented attributes side by side, ties every cell to the vendor's own documentation or to peer-reviewed evidence, and names no single best tool. As of July 2026. The whole category turns on one design decision — how much to automate before a human reviews — so the human-in-the-loop concept is the lens to read it through.

How to use this comparison

Coding automation is unusual among AI categories in that the marketed accuracy numbers and the peer-reviewed accuracy numbers come from different worlds. Vendors report high-nineties percentages; independent studies of the underlying models report much lower. A ranking would have to pick one of those worlds and pretend it settles the question. A matrix keeps both in view: it records what each product documents about autonomy and routing, and it places that beside what independent research has actually measured, leaving the judgment to the reader who knows the setting. Read every figure against the method behind it, using our guide on how to read an AI validation study.

What the tools are

The products cluster along a single axis: how much they code without a human.

  • CodaMetrix documents an autonomous, "AI-Powered Contextual Coding Automation Platform" that "automatically applies codes across service lines" and routes complex accounts to human coders 3.
  • Nym documents autonomous coding built on "Clinical Language Understanding" plus "a rules-based approach," producing "fully transparent audit trails for every code assigned" 4.
  • Solventum 360 Encompass documents a computer-assisted coding model whose "expert-guided clinical AI" distinguishes "routine and complex cases," auto-submitting routine visits and routing complex ones to coders 5 — and, separately, an autonomous coding option 6, showing that one platform can span the assisted-to-autonomous range.

The capability matrix

Every filled cell cites its source. A dash means the attribute was not stated in the sources cited here — read it as "confirm with the vendor," not as "absent."

AttributeCodaMetrixNymSolventum 360 Encompass (CAC)Solventum 360 Encompass (Autonomous)
Automation modeAutonomous 3Autonomous 4Assistive / computer-assisted 5Autonomous 6
Core approachContextual coding AI 3Clinical Language Understanding + rules 4Expert-guided clinical AI 5AI-powered autonomous coding 6
Complex cases routed to human codersYes 3Yes 4Yes 5Yes 6
Audit trail / explainability documentedYes 4
Published independent accuracy audit

The last row is the one to sit with: no vendor here has a published, independent accuracy audit tied to its product. That empty row is the single most important fact on the page, and the next section explains why.

What independent research actually measured

Set the vendor claims aside and look at what peer-reviewed work has found when it tests the underlying technology directly. A benchmark that drew more than 27,000 real diagnosis and procedure codes and prompted models from three major developers concluded flatly that large language models "are highly error-prone when mapping medical codes" 1. A second study tested five models on ICD-9 code prediction and found the best of them — a reasoning model — reached just 45.3% accuracy, with the weakest at 14.6%, and concluded that "none of the models met clinical standards across all tasks" 2.

StudyWhat was testedResult
Code-querying benchmark 127,000+ real codes, models from three developersModels "highly error-prone when mapping medical codes"
ICD-9 prediction study 2Five LLMs, zero-shot promptingBest 45.3%, weakest 14.6%; "none met clinical standards"

This is the gap the whole product category exists to close. A raw model cannot be trusted to assign codes, so every deployed system wraps it — with curated code sets, deterministic rules, and, crucially, human review of the cases the system is unsure about. Nym's pairing of language understanding with "a rules-based approach" 4 and Solventum's routing of complex cases to coders 5 are two expressions of the same lesson. The characteristic failure of an unwrapped model — a confident but wrong code — is the coding-specific face of AI hallucination in clinical contexts, and it is why the human-review row in the matrix is filled for every product.

Why the wrapper is the whole game

If the raw model is unreliable and the finished product works, the value lives almost entirely in the engineering between them. That wrapper does three jobs. It constrains the output to valid, current code sets, so the system cannot invent a code that does not exist — a discipline a free-running model lacks. It routes by confidence, coding the routine cases automatically and diverting the ambiguous ones to a person, which is what Solventum's split between "routine and complex cases" 5 and Nym's fallback to human review 4 both describe. And it records why each code was chosen, so an auditor can retrace the decision — the purpose of Nym's "fully transparent audit trails" 4. When you compare products, you are comparing wrappers, and the peer-reviewed base rates 12 are the reminder that the raw model underneath is the same weak starting point for everyone.

Reading vendor accuracy numbers

Deployed platforms advertise accuracy figures well above what independent studies of raw models report — CodaMetrix, for instance, states a "70% reduction in manual coding" 3. Three cautions apply to every such number. First, it is self-reported, with no published independent audit behind it. Second, the denominator is often unstated: a reduction in manual coding, an automation rate, and a code-level accuracy rate answer different questions, and a headline number rarely says which it is. Third, results come from the vendor's best deployments, which are the ones that get published. None of this means the numbers are wrong — it means they are claims awaiting independent confirmation, and should be read as such until a peer-reviewed evaluation of the specific product lands.

The denominator problem is worth making concrete. Suppose a system autonomously codes 60% of charts and sends 40% to human coders, and that of the autonomous 60%, some fraction carries a coding error. A vendor could truthfully report a high "automation rate," a high "accuracy on autonomously coded charts," and a large "reduction in manual coding" — three impressive figures that still leave open the question a compliance officer actually cares about: what is the error rate across all charts, including the ones a human touched, and how many of those errors are under-coding that loses legitimate revenue versus over-coding that invites a clawback. A single percentage cannot answer that. This is why the only number worth acting on is one you generate yourself, by auditing a sample of the system's output against your own coders — the local test the peer-reviewed base rates 12 make non-negotiable.

The oversight frame

Coding automation is administrative revenue-cycle software, and it generally sits outside the FDA's medical-device pathway — the agency's decision-support framing concerns software that informs clinical care, and turns on whether a clinician can independently review the basis for an output 7. Coding tools inform billing rather than treatment, so their oversight comes from a different direction: coding audits, revenue-integrity review, and compliance controls. The stakes remain high — a wrong code is a compliance and reimbursement exposure — but the safeguard is the audit trail and the human reviewer, rather than a clearance. This is why Nym's documented "transparent audit trails" 4 is a substantive attribute rather than a marketing line: in an unregulated category, the ability to trace why a code was assigned is the control that makes autonomy accountable. For the regulated side of healthcare AI by contrast, see our tally of FDA-cleared AI devices.

Matching the mode to the risk

The assisted-versus-autonomous choice is a risk decision as much as a productivity one. Computer-assisted coding keeps a human on every chart and trades some speed for a standing safety net; the fully autonomous mode removes the coder from routine charts and concentrates human attention on the cases the system flags. The right point on that spectrum depends on the service line, the tolerance for a coding error, and the maturity of the audit process behind it — which is why a single vendor now documents both an assisted and an autonomous configuration 56. For a high-risk service line, the defensible default is to begin assisted and earn the move toward autonomy with your own audit data, rather than adopting the most automated mode on the strength of a vendor figure.

How to choose for your setting

Ask questions, not for a ranking.

  • How much does it automate before a human reviews, and how are uncertain cases caught? The autonomy mode 35 defines your residual coder workload and your error exposure.
  • Which code systems and service lines does it actually cover? Coverage of ICD-10, CPT, and risk-adjustment coding varies, and a gap becomes manual work.
  • Can you audit every assigned code? A documented, traceable audit trail 4 is the accountability backbone in a category without a clearance.
  • Will the vendor support a blinded accuracy test on your charts? Given the peer-reviewed base rates 12, a local audit against your own coders is the number that should decide the purchase.
  • How are denials and downstream corrections handled? Automation that shifts work to appeals has moved the cost rather than removed it.

How to read this comparison

Three cautions travel with the matrix. First, the accuracy evidence splits in two — high vendor-stated numbers with no independent audit, and lower peer-reviewed numbers for raw models — and an honest reading holds both, rather than choosing the flattering one. Second, the peer-reviewed studies test underlying models, not the wrapped commercial systems, so they set a floor for what the technology can do alone rather than a ceiling for an engineered product. Third, autonomy models, code coverage, and routing change on the vendors' timelines, so every cell is dated July 2026, and availability changes — reconfirm before you rely on any single cell.

Sources and method

This comparison pairs two peer-reviewed evaluations of code-assignment accuracy 12 with each vendor's own product documentation 3456 and the FDA's clinical decision support guidance for the oversight frame 7. Every filled matrix cell is tied to one of these; the empty independent-audit row is itself a sourced finding. We revisit this page on a 180-day cycle and whenever a peer-reviewed accuracy audit of a named product, or a documented change in autonomy or coverage, lands. For the tools that draft the note the coder works from, see our AI scribe head-to-head comparison.

Questions & answers

  • How accurate is AI medical coding?

    For a general large language model used on its own, the published evidence is unflattering: one benchmark called such models highly error-prone at mapping codes, and in another the best of five models reached only about 45% accuracy on ICD-9 prediction. Deployed commercial systems report far higher numbers, but those are vendor-stated and lack published independent audits, so they are best read as claims until a peer-reviewed evaluation confirms them.

  • What is the difference between autonomous coding and computer-assisted coding?

    Computer-assisted coding suggests codes for a human coder to confirm. Autonomous coding assigns and submits codes for routine cases without a coder touching them, routing only complex or uncertain cases to a human. Several platforms now offer both modes; the practical question is how much they automate before a person reviews, and how uncertain cases are caught.

  • Are AI coding tools FDA-regulated?

    Generally not. Coding automation is administrative revenue-cycle software rather than a clinical decision tool, so it typically sits outside the FDA's medical-device pathway. That does not lower the stakes — coding errors carry compliance and reimbursement risk — but it means oversight comes from audit and revenue-integrity processes rather than from a device clearance.

Sources

  1. Soroush A, Glicksberg BS, Zimlichman E, et al. Large Language Models Are Poor Medical Coders — Benchmarking of Medical Code Querying. NEJM AI. 2024. doi.org/10.1056/AIdbp2300040
  2. Evaluating the Reasoning Capabilities of Large Language Models for Medical Coding and Hospital Readmission Risk Stratification: Zero-Shot Prompting Approach. Journal of Medical Internet Research. 2025. doi.org/10.2196/74142
  3. CodaMetrix — coding automation platform documentation (accessed July 2026). www.codametrix.com/
  4. Nym Health — autonomous coding documentation (accessed July 2026). www.nym.health/
  5. Solventum. 360 Encompass Computer-Assisted Coding System — product documentation (accessed July 2026). www.solventum.com/en-us/home/health-information-technology/solutions/360-encompass-cac/
  6. Solventum. 360 Encompass Autonomous Coding System — product documentation (accessed July 2026). www.solventum.com/en-us/home/health-information-technology/solutions/360-encompass-autonomous/
  7. US Food and Drug Administration. Clinical Decision Support Software — Guidance for Industry and FDA Staff. www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software