Where Autonomous Medical Coding Actually Works in 2026
Physician adoption of AI is no longer the question. The AMA's 2026 Physician AI Sentiment Report, fielded in January and February 2026 with 1,692 respondents, found 81% of physicians using AI professionally, up from 38% in 2023. The question is which workflows can carry autonomy and which cannot — and in coding, the honest answer is narrower than the marketing.
Autonomous medical coding is real in 2026, and it is much narrower than the category name suggests. The service lines where practices are genuinely letting an agent assign CPT and ICD-10 codes without a coder touching every chart are the same three they were eighteen months ago: diagnostic radiology, pathology, and the emergency department. What those three share is not a technology property. It is a documentation property.
In each, the volume is high, the documentation is structured or templated, the code space for a given study or presentation is comparatively bounded, and — this is the decisive part — the coding decision is largely a transformation of something the report already states explicitly. A chest CT with contrast says so. A synoptic pathology report names the specimen. That is extraction with rules on top, and language models are good at extraction with rules on top.
Complex office visits fail every one of those tests. The level of service for a 99214 versus a 99215 turns on medical decision making, and medical decision making is not stated in the note — it has to be inferred from narrative, from what the physician considered and rejected, from risk that is often implicit. An agent that infers MDM from narrative is not extracting; it is judging. That is a different reliability regime, and it is why the honest deployments in 2026 keep office E/M in assisted mode with a human deciding the level.
| Service line | Documentation profile | Defensible 2026 posture | Where it breaks |
|---|---|---|---|
| Diagnostic radiology | Structured report, templated impression, bounded code space per study type | Autonomous is defensible for routine studies behind a conservative threshold | Modality/contrast/laterality errors; unsigned or amended reports |
| Pathology | Highly templated synoptic reporting; specimen-driven code selection | Autonomous is defensible for routine specimens | Multi-specimen accessions; add-on stains ordered after the fact |
| Emergency department | High volume, structured templates, repeatable presentation patterns | Autonomous viable for lower-acuity, routed for critical care and procedures | Critical care time, procedure bundling, observation status |
| Urgent care | Semi-structured; narrow but not bounded code space | Assisted with human sign-off; autonomy only after a clean audit cycle | E/M level drift; injection and procedure add-ons |
| Office / outpatient E/M | Narrative documentation; level of service turns on medical decision making | Assisted only. Human coder decides the level | MDM inference; time-based billing; social determinants capture |
| Surgical / procedural | Operative notes vary by surgeon; heavy modifier logic | Assisted only; global period and modifier rules need a human | Modifier 22/25/59, global period, staged procedures |
| Inpatient / DRG | Long, multi-author records; sequencing drives payment | Assisted only; CDI and coder judgment dominate | Principal diagnosis sequencing; POA indicators; query workflow |
| Behavioral health | Narrative-heavy; time and modality dependent | Assisted only; consent and modality rules are practice-specific | Time thresholds; telehealth modality; interactive complexity |
Service-line viability for autonomous coding, 2026. Frenchy Digital deployment guidance.
One clarification before any of the rest of this is useful. Everything in this article concerns administrative and documentation workflows under human governance. A coding agent assigns billing codes from documentation a clinician already authored. It does not diagnose, it does not treat, and it does not practice medicine. The AMA's framing of "augmented intelligence" is the correct one, and it is also the one that keeps you inside the lines.
The Evidence Problem: No Independent Benchmark of AI Coding Accuracy Exists
Here is the fact that should shape every purchasing conversation in this category, stated as plainly as we can state it: as of August 2026 there is no public, independent benchmark of AI medical coding accuracy. Not a peer-reviewed head-to-head. Not a neutral third-party test set. Not a standards body evaluation. Every accuracy figure circulating in this market is vendor-published, generated on a vendor-selected sample, scored against a vendor-defined notion of correctness.
That is not an accusation of dishonesty. It is a structural observation. Ambient documentation, an adjacent category, has accumulated real independent evidence — randomized trials, matched-cohort studies, multi-site outcome reports. Coding has not. The category grew fast, the buyers were operators rather than researchers, and nobody built the benchmark.
For completeness, the figures vendors most commonly publish are these, and we label them as what they are:
- 92–97% accuracy on structured, high-volume encounters (vendor claim): Typically cited for emergency department and outpatient radiology. The measurement conditions are almost never disclosed.
- 82–90% accuracy on complex inpatient coding (vendor claim): Note that inpatient accuracy is a different construct — sequencing and present-on-admission indicators drive payment as much as code selection does.
- 95–98% accuracy for human coders post-QA (vendor-published comparator): This is the benchmark vendors compare themselves against. It is also vendor-published, and 'post-QA' means after a QA process that may not resemble yours.
A practice cannot verify any of those numbers, and the reason is worth spelling out because it is the whole argument. An accuracy rate is meaningless without four disclosures that vendors essentially never make in writing:
- 1.The case mix: Which specialties, which payers, which acuity distribution. A 96% on routine screening mammography and a 96% on multi-trauma ED are not the same 96%.
- 2.The denominator: Primary code only, or every line? Are modifiers scored? Secondary diagnoses? Were low-confidence encounters excluded from the sample before the number was computed?
- 3.The definition of correct: Exact match to a reference coder, or 'payable and defensible'? Those diverge sharply on E/M levels and on diagnosis specificity.
- 4.The adjudication method: One reviewer, or two blinded reviewers with a third adjudicating disagreement? A single-reviewer ground truth measures agreement with one person's habits.
| Claim you will hear | What it actually is | What is missing | What to do instead |
|---|---|---|---|
| AI 92–97% accuracy on structured high-volume encounters | Vendor-published | Case mix, denominator, definition of correct, adjudication method | Run the same measurement on your own claims |
| AI 82–90% accuracy on complex inpatient | Vendor-published | Whether sequencing and POA were scored, or only code presence | Score sequencing separately from code selection |
| Human coders 95–98% accurate post-QA | Vendor-published comparator | Whether post-QA means after the same QA you actually run | Measure your own coders on the same blinded sample |
| Denials reduced from 18% to 3% | Vendor marketing | Baseline period, payer mix, whether policy changes were concurrent | Treat as marketing; do not put it in a board deck |
| 4–7x coder productivity | Vendor marketing | Whether review time and rework were counted in the denominator | Measure charts-per-coder-hour including rework |
| Confidence score of 0.94 | Vendor-internal | Whether the score is calibrated against observed accuracy | Demand a calibration curve from your own audit data |
Vendor claims in the AI coding market and the disclosures that would make them checkable, 2026.
An unfalsifiable number cannot support a compliance decision. If you cannot describe how a figure could be shown to be wrong, do not put it in a board deck, a payer conversation, or a policy document.
— Frenchy Digital evaluation principle
The practical consequence is not "do not buy." It is that the accuracy number you rely on has to be one you generated. That is the rest of this article's real subject, and it is a solvable problem — the audit costs a few thousand dollars in coder time and takes a couple of weeks. What it buys you is a defensible figure and, more importantly, a defensible process.
How to Run Your Own Sampled Accuracy Audit
This is the part of the engagement that most practices skip and that pays for itself the first time a payer asks a question. The design below is the one Frenchy Digital runs before we let any coding agent post a claim. It is deliberately unglamorous.
| Design parameter | Frenchy Digital default | Why it is designed this way |
|---|---|---|
| Unit of analysis | The claim line, rolled up to the encounter | Encounter-level agreement hides modifier and secondary-diagnosis errors |
| Ground truth | Two credentialed coders working blinded from the documentation; a third adjudicates disagreements | A single reviewer measures agreement with one person, not accuracy |
| Blinding | Reviewers never see the AI output or its confidence score | Seeing the suggestion is exactly the automation bias you are trying to measure |
| Human baseline | Score your own coders on the same sample in the same pass | Without a baseline you cannot tell whether the agent is better or worse than today |
| Stratification | Sample from every confidence band, every service line, and every top payer | Auto-accepted encounters are the ones nobody is otherwise looking at |
| Initial sample | Frenchy Digital default: 30 charts per stratum as a screening probe | Cheap enough to run pre-deployment; large enough to surface systematic errors |
| Confirmatory sample | Expand to a statistically valid sample, sized with compliance counsel, if the probe fails | Extrapolation and repayment exposure are legal decisions, not engineering ones |
| Error taxonomy | Wrong primary code; wrong E/M level; missing secondary dx; modifier error; laterality; bundling; medical-necessity linkage | A single accuracy percentage tells you nothing about what to fix |
| Directional metric | Net RVU delta versus adjudicated truth, reported separately from raw accuracy | Two systems can be 94% accurate with opposite financial risk profiles |
| Cadence | Baseline pre-deployment, monthly for the first quarter, then quarterly | Model, prompt, and code-set changes all invalidate a stale audit |
| Retention | Every sampled chart, every reviewer decision, every adjudication, versioned | This file is your evidence that review actually happened |
Sampled accuracy audit design for an AI coding agent — Frenchy Digital methodology, 2026.
The two steps everyone leaves out
Measure your own coders on the same sample. Practices routinely audit the AI against a reference standard and never audit their humans against the same standard in the same pass. Without that, you have no idea whether an 89% agent is a downgrade or an upgrade. In several engagements the human baseline came in materially below the number the practice assumed it was running at — which changed the decision entirely.
Sample from the auto-accepted bucket. The encounters that never reach a human are, by definition, the ones no control is examining. If your QA sample is drawn only from the human-reviewed queue, you have built a monitoring program that structurally cannot see the risk you created. Every monthly sample must be stratified across confidence bands, with the top band included.
Report accuracy and net RVU delta as two separate numbers. Two agents can both be 94% accurate while one drifts revenue upward and the other drifts it downward. Only one of those creates enforcement exposure, and a single accuracy percentage cannot tell them apart.
On sample sizing, a note about scope. The 30-charts-per-stratum screening probe above is our engineering convention for detecting systematic error cheaply, not a legal standard. If a probe surfaces a real error pattern, the next step — how large a confirmatory sample, whether findings get extrapolated, and what any of it obligates you to do — is a question for healthcare counsel and your compliance officer, not for your development partner. We build the pipeline and hand over the file; we do not size a repayment exposure.
Confidence Routing and the Threshold Trap
Essentially every serious 2026 architecture in this space works the same way. The agent produces a code set and a confidence score. Output above a threshold posts automatically. Output below it queues for a human coder. This is a good design — it concentrates scarce coder attention on the encounters where judgment actually pays.
The trap is what the threshold is. Raise it and more work routes to humans: higher labor cost, lower error exposure. Lower it and the reverse. That means the threshold is a dial that trades money against compliance risk, and in most implementations it is set once during onboarding by someone in operations, never written down, and quietly adjusted later when the coding queue backs up.
There is a second, subtler problem: confidence scores are usually uncalibrated. A score of 0.94 is a ranking signal from inside the model. It is not a claim that 94% of such outputs are correct — unless someone has measured observed accuracy per confidence bucket and shown that the curve tracks. Almost no vendor publishes that curve. Your own audit can produce it: bucket your sampled encounters by the confidence the agent assigned, compute adjudicated accuracy within each bucket, and plot it. If observed accuracy in the 0.90–0.95 band is 78%, your threshold is not where you think it is.
Finally, set the threshold asymmetrically. A code that increases reimbursement relative to your historical distribution should face a higher bar than a code that decreases it. This is not conservatism for its own sake — it directly reduces the pattern described in the enforcement section below.
| Routing tier | Trigger | Control |
|---|---|---|
| Auto-accept | High confidence, low-risk service line, no financial upside vs. adjudicated baseline | Sampled QA at a fixed monthly rate, no exceptions |
| Auto-accept with hold | High confidence but code increases reimbursement versus the prior distribution | Second-pass review before the claim posts |
| Route to coder | Below threshold, or any modifier logic, or any add-on procedure | Full human coding; the AI suggestion is visible only after the coder commits |
| Route to CDI / query | Documentation does not support any defensible code | Physician query workflow; never let the agent infer the missing element |
| Hard stop | New or revised code effective within the last 90 days; unsigned note; amended note | Human only until the next revalidation cycle clears it |
Confidence routing tiers for a coding agent — Frenchy Digital reference configuration, 2026.
CPT Appendix S: A Taxonomy, Not a Payment Schedule
CPT Appendix S, introduced by the AMA in 2021, is the correct vocabulary for this category, and it is routinely misdescribed by vendors. Appendix S classifies AI-enabled medical services and procedures into three categories — assistive, augmentative, and autonomous — corresponding to Levels I, II, and III, distinguished by the degree to which a physician must interpret the output and can override it.
| Appendix S category | What it describes | Level | What it means for a coding agent |
|---|---|---|---|
| Assistive | Software detects clinically relevant data without independent analysis or conclusions; the work product requires physician interpretation and action | Level I | Most computer-assisted coding sits here in practice |
| Augmentative | Software analyzes or quantifies data to yield clinically meaningful output; the physician still interprets and acts | Level II | Where a coding agent that proposes and explains codes belongs |
| Autonomous | Software interprets data and independently generates clinically meaningful conclusions; conclusions may be subject to physician override | Level III | The category vendors invoke; the override rights inside it are what matter |
| What it is not | A payment schedule, a fee category, or a route to reimbursement for AI software | n/a | No general AI CPT payment category exists in 2026 |
| May 2026 revisions | CPT Editorial Panel accepted revisions replacing 'machine' with 'software output(s)' | n/a | Language now tracks software behavior rather than hardware |
| In development | Clinically Meaningful Algorithmic Analyses (CMAA) framework | n/a | Watch it, but do not build a business case on it yet |
CPT Appendix S taxonomy for AI-enabled services, with the May 2026 CPT Editorial Panel revisions. Source: AMA.
At the May 2026 CPT Editorial Panel meeting, revisions were accepted that replace the word "machine" throughout Appendix S with "software output(s)" — a small edit with a real effect, since it moves the taxonomy's language from hardware to software behavior, which is where this category actually lives. The AMA is separately developing a framework for Clinically Meaningful Algorithmic Analyses (CMAA). That work is worth tracking. It is not worth building a business case on today.
What Actually Changed in CPT 2026 — and Why It Is an Agent Problem
The CPT 2026 code set, effective January 1, 2026, carried 418 editorial changes: 288 new codes, 84 deletions, and 46 revisions (as reported on the AMA release). For a practice, that is an annual operational chore. For an AI coding agent, it is a versioning discipline that most implementations do not have.
| CPT 2026 item | Detail | Agent implication |
|---|---|---|
| Total editorial changes | 418 changes effective January 1, 2026 | Every agent configuration and code-set mapping needs annual revalidation |
| New codes | 288 | New codes are the highest-risk surface for an agent: no training precedent |
| Deleted codes | 84 | A deleted code that survives in a mapping table becomes a rejected claim |
| Revised codes | 46 | Revisions are the silent failure: the code still exists, the meaning moved |
| Office/outpatient E/M 99202–99215 | Unchanged | The hardest coding decision in the practice did not get easier |
| RPM/RTM restructuring | New 99445 and 99470; 99453/99454/99457/+99458 revised | If you bill remote monitoring, this is the section to re-verify first |
| 99453 setup threshold | Dropped from 16 days of data to 2 days | Materially changes eligibility logic that an agent may have hard-coded |
| 99470 | First 10 minutes of treatment management; requires at least one real-time audio/video interaction | An agent cannot infer the interactive requirement from device data alone |
CPT 2026 changes most relevant to an automated coding pipeline. Effective January 1, 2026.
Two things stand out. First, the core office and outpatient E/M codes 99202–99215 were unchanged, still driven by medical decision making or total time. The single hardest coding judgment in an ambulatory practice did not get any easier in 2026, which reinforces the point from the first section about where autonomy belongs.
Second, the notable restructuring landed in remote physiologic and remote therapeutic monitoring. New codes 99445 and 99470 arrived, 99453/99454/99457/+99458 were revised, and the 99453 setup threshold dropped from 16 days of data to 2 days. That last change is exactly the kind of parameter an integration hard-codes and then forgets. If your agent has a rule that counts to sixteen, it has been wrong since January.
The general lesson: revised codes are more dangerous than deleted ones. A deleted code produces a rejection you will notice. A revised code still exists, still passes edits, and now means something slightly different — and an agent configured against last year's definition will keep submitting it confidently. Build an annual revalidation into the operating cadence, and treat any code effective within the last 90 days as human-only until it clears.
What Gets Paid For AI in 2026: The Honest Answer
Every ROI conversation in this category eventually reaches the same question: does anyone pay for this? The answer for AI coding in 2026 is no, and the specific shape of that "no" is worth understanding because vendors work hard around its edges.
In the CY2026 Medicare Physician Fee Schedule final rule, CMS solicited comment on how software-as-a-service and algorithm-based services should be paid, noting that they are not well accounted for under current payment methodology. Soliciting comment is not finalizing a pathway. No general AI payment mechanism was established (summarized here).
| Item | 2026 status | What it actually is |
|---|---|---|
| A CPT payment category for AI services | Does not exist | CPT Appendix S is a taxonomy only |
| Medicare PFS payment for AI coding software | Does not exist | CY2026 PFS solicited comment on SaaS/algorithm payment; nothing finalized |
| New Technology Add-on Payment (NTAP) | Inpatient only | Irrelevant to outpatient practice economics |
| Digital Mental Health Treatment G-codes | G0552, G0553, G0554 exist; CMS finalized expansion to ADHD | Device supply and treatment management, not AI coding |
| Remote monitoring codes | 99445, 99470 and the revised RPM/RTM family | Paid for the clinical service, not for the software that codes it |
| AI coding, ambient scribes, AI denial tools | No CPT or CMS reimbursement | Cost-side ROI only. Build the business case on labor and rework |
What Medicare and CPT do and do not pay for in relation to AI, as of August 2026.
The adjacent codes that did arrive are for clinical services, not for software. The restructured remote monitoring family pays for monitoring a patient. The Digital Mental Health Treatment codes G0552, G0553 and G0554 pay for device supply and treatment management, with CMS finalizing expansion to ADHD. New Technology Add-on Payment is an inpatient mechanism and does not touch outpatient practice economics at all. None of these pay you for running an AI coder.
The Compliance Core: OIG, DOJ, and the Automation-Bias Theory
This is the section that should determine whether you deploy at all, and how.
In its Medicare Advantage Industry Segment-Specific Compliance Program Guidance, published February 3, 2026, the HHS Office of Inspector General identified as potentially abusive the practice of "querying physicians via electronic medical record platforms, including prompts generated by artificial intelligence algorithms, to add risk-adjusting diagnoses." Read that carefully. The object of concern is not a fraudulent diagnosis. It is the prompt — the mechanism by which a physician is nudged toward a revenue-relevant code. That is the clearest federal statement to date that AI inside the documentation workflow can itself be the compliance problem.
On the enforcement side, the Department of Justice's 2026 National Health Care Fraud Takedown foregrounded AI and data analytics, and the Health Care Fraud Unit's Data Fusion Center — operating with HHS-OIG and the FBI — is now standing infrastructure rather than a pilot. The asymmetry is worth internalizing: the government is doing pattern detection at scale on claims data. If your agent creates a pattern, the pattern is visible.
Which brings us to the theory that should shape your controls. Practices deploying these tools almost universally cite human-in-the-loop review as their mitigating control. The enforcement counter-argument is straightforward: a reviewer accepting roughly 85% of AI suggestions at roughly half a second per chart is not reviewing anything. That is automation bias, and it converts your stated control into documented evidence that the control was not operating. Your own telemetry becomes the exhibit.
Human-in-the-loop is a compliance control, not a courtesy. If you cannot produce per-encounter review time and per-reviewer override rates, you do not have a human in the loop — you have a human in the log.
— Frenchy Digital compliance principle
| Control | What it produces | Why enforcement cares |
|---|---|---|
| Documented review time | Per-encounter dwell time for every human reviewer, logged automatically | The OIG theory of the case is that instantaneous acceptance is not review |
| Override and rejection rates | Per reviewer, per service line, per confidence band, trended monthly | A reviewer at a 99% acceptance rate is a finding, not a top performer |
| Immutable audit logs | Who saw what, when, what the AI proposed, what was submitted, what changed | Reconstructing a claim two years later is the entire ballgame |
| Periodic sampled QA | Blinded re-coding of a monthly sample including auto-accepted encounters | Auto-accepted work is invisible to every other control you have |
| Distribution monitoring | E/M level mix, RVU per encounter, diagnosis density, HCC capture rate | Enforcement analytics look at distributions, so you should look first |
| Physician query integrity | No AI-authored prompt that suggests a specific risk-adjusting diagnosis | Named in the OIG's February 2026 Medicare Advantage guidance |
| Change control | Version and date every model, prompt, threshold, and code-set update | An unversioned threshold change is an unexplainable revenue shift |
| Vendor due diligence file | BAA, retention terms, training-use terms, subprocessors, incident history | Your compliance posture inherits your vendor's |
Minimum control set for an AI coding deployment — Frenchy Digital implementation standard, 2026.
Two structural notes. Any vendor touching protected health information is a business associate under HIPAA and requires a signed BAA before a chart moves; there is no such thing as a "HIPAA-certified" vendor, only a compliant posture with evidenced Security Rule safeguards. And the HIPAA Security Rule NPRM published January 6, 2025 — proposing mandatory MFA, encryption at rest and in transit, asset inventories, annual audits, and 72-hour restoration — remains proposed, not final; the Unified Agenda now projects final action in 2027. Build toward it as a design target, but do not let anyone sell it to you as a current obligation.
If your coding agent is embedded in certified health IT, one more artifact is available free: the ASTP/ONC Decision Support Intervention transparency data. Predictive DSIs in certified products must publish 31 source attributes covering development details, fairness, external validation, quantitative performance, and maintenance. That is ready-made vendor due diligence a practice can read before it signs anything.
Two boundaries worth knowing so you do not over- or under-scope your compliance work. A coding agent is administrative software; it is generally not a medical device, and the FDA's Clinical Decision Support Software final guidance re-issued January 29, 2026 — which added an enforcement-discretion policy for software producing a single, clinically appropriate output — governs clinical decision support, not billing automation. Separately, CMS-0057-F binds payers rather than practices, and its FHIR API obligations do not arrive until January 1, 2027. For governance structure in the meantime, the NIST AI Risk Management Framework is the most usable non-binding scaffold we have found for documenting how you govern a model you did not build.
Upcoding: Note the Direction of the Risk
This deserves its own section because the risk has a direction, and the direction is not symmetric.
An AI coding agent tuned to "capture revenue that is being left on the table" is, by construction, tuned toward higher-level codes and denser diagnosis capture. That is the product working as specified. It is also, precisely, the pattern that payer analytics and federal data-mining programs are built to find. An agent optimized for revenue capture is an agent optimized toward the exact signature enforcement looks for.
The exposure is rarely any single code. Individual codes are usually defensible in isolation. The exposure is distribution shift — your 99214/99215 mix moving three points, your diagnoses-per-encounter rising fifteen percent, your HCC capture rate stepping up on a date that happens to be your go-live. Aggregate patterns are what get investigated, and "the software suggested it" is not a defense that has worked well historically.
| Metric to baseline | How to measure it | What should trigger review |
|---|---|---|
| E/M level distribution | Share of 99213 / 99214 / 99215 by provider | Any provider whose level-4+ share moves more than a few points post-deployment |
| RVU per encounter | Work RVUs per completed encounter by service line | A step change that starts on the deployment date |
| Diagnosis density | Average diagnoses submitted per encounter | Density rising without a change in panel acuity |
| Risk-adjusting diagnosis capture | New HCC-mapped diagnoses added at the coding step | Diagnoses appearing at coding that never appear in the assessment |
| Modifier frequency | Rate of modifier 25 and 59 use | Modifier use rising to clear edits rather than to describe the encounter |
| Override rate by reviewer | Percentage of AI suggestions changed before submission | Rates near zero, and rates that fall over time as reviewers habituate |
Distribution-drift monitoring for an AI coding deployment. Baseline before go-live; trend monthly.
Two contract-level protections follow. First, do not price the vendor on a share of incremental coding-level lift. Contingency on collections is an established RCM structure; contingency specifically on moving your codes upward pays someone else to take your compliance risk. Second, require that the agent never authors a physician query that names a specific risk-adjusting diagnosis — the mechanism the OIG called out in February. An agent may flag that documentation is insufficient. A human decides what to ask.
Reference Architecture for a Coding Agent
What we actually build, in the order we build it. None of this is exotic; the discipline is in refusing to skip steps under delivery pressure.
- Baseline audit first: Blinded dual-coder audit of your current output, with the human baseline and distribution snapshot captured before any automation exists. This artifact is your before-picture and you only get one chance to take it.
- Scoped EHR integration: Read access limited to the documentation the agent needs, write access limited to the coding fields it owns. Minimum necessary is an architecture constraint, not a policy sentence.
- Deterministic rules ahead of the model: Payer edits, NCCI-style bundling logic, code-set validity, and effective-date checks run as code. Models are for extraction and inference; rules are for things that must be right.
- Confidence scoring with an owned threshold: Threshold stored as a versioned configuration value with an owner, a rationale, and an approval trail. Not a hidden vendor default.
- Asymmetric routing: Higher bar for codes that increase reimbursement relative to your baseline distribution than for codes that decrease it.
- Immutable audit log: Append-only record of the source documentation version, the model and prompt version, the proposal, the confidence, the reviewer, the dwell time, and the submitted result.
- Sampling pipeline built on day one: Monthly stratified sample generated automatically across confidence bands and service lines, routed to blinded reviewers. If QA is manual, QA stops happening in month three.
- Distribution monitoring and alerting: Automated trending of E/M mix, RVU per encounter, and diagnosis density, with thresholds that page a human rather than a dashboard nobody opens.
- Code-set version control: Annual revalidation against the new CPT and ICD-10 releases, with any code effective in the last 90 days routed to humans until cleared.
- Rollback path: A single switch that returns a service line to full human coding without a deployment. You will use it, and the day you need it is not the day to build it.
On stacking an ambient scribe under a coding agent
A large share of 2026 deployments feed an AI-generated note into an AI coding agent. That composition deserves explicit attention, because the coding agent inherits every documentation defect upstream of it.
A published review of 356 ambient-generated notes found omissions in 18% and hallucinations in 11.5%, with most errors mild to moderate but 5.3% of notes containing errors rated potentially seriously harmful. A 2026 JMIR research letter separately documented ambient scribes propagating interpreter errors. A coding agent reading those notes cannot distinguish a hallucinated finding from a real one — it will code what the note says.
If you run both, the physician attestation on the note is doing more work than anyone acknowledges, and your audit sample should include a step that compares the note against the encounter rather than only the code against the note.
Building the Cost-Side ROI Model Honestly
Since nothing pays for AI coding directly, the business case has to be built on cost. Here is how to build one that survives scrutiny — and which inputs are solid versus soft.
The solid inputs: the U.S. Bureau of Labor Statistics puts median pay for medical records specialists, which includes billing and coding roles, at $50,250 in 2024. MGMA reports that automation is the top cost-cutting move groups planned for 2026, cited by 36%, well ahead of hiring freezes at 18%. On the denial side, Premier Inc. found claims adjudication cost providers $25.7 billion in 2023, with roughly $18 billion potentially wasted on claims that should have paid on submission, at an average adjudication cost of $57.23 per claim — 2023 data with no newer primary update we could locate.
The soft inputs, labeled: the commonly quoted fully loaded cost of $60,000–$80,000 per in-house biller including benefits and software is a vendor-relayed figure, not a survey result — use it as a planning range, not a citation. Outsourced RCM pricing of 4–10% of net collections is likewise vendor-quoted, though consistent enough across sources to plan against. The frequently cited HFMA "cost to collect under 2% of net collections" benchmark reaches us through secondary sources rather than a primary HFMA publication; treat it as a convention, not a standard. And the widely circulated rework-cost figures attributed to MGMA could not be traced to a primary MGMA publication at all, so we do not use them.
One more honesty check on the revenue side. In KLAS Arch Collaborative research on ambient documentation, over 80% of providers declined to see more patients after time was freed up, with only 18% wanting added volume. If your model converts saved minutes into incremental encounters, it is assuming a behavior most clinicians explicitly reject. Model the labor saving. Do not model the volume lift unless you have committed capacity and a physician who has agreed to it in writing.
On the denial lever specifically, use primary data and name the denominator. KFF's March 2026 analysis of 2024 ACA Marketplace claims found 19% of in-network claims denied, with administrative reasons accounting for 25% of denials and missing prior authorization or referral another 9% — the two buckets a coding and documentation agent can actually influence. Separately, KFF's January 2026 Medicare Advantage analysis found 7.7% of prior-authorization determinations denied in full or part. That is a prior-authorization figure, not a claims figure; the two get conflated constantly in vendor material, and using the wrong one will get your model challenged.
Frenchy Digital Cost Bands and Timelines
What this costs to build properly, using the same scoping bands we apply across every healthcare automation engagement:
| Engagement | Range | Timeline | Typical scope |
|---|---|---|---|
| Discovery + workflow audit | $9k–$22k | 2–4 weeks | Baseline coding audit, service-line viability assessment, vendor due diligence, written build plan |
| Single-workflow agent | $25k–$65k | 4–9 weeks | One service line, shadow mode then limited autonomy, confidence routing, audit logging |
| Multi-workflow with EHR integration | $65k–$160k | 9–16 weeks | Multiple service lines, EHR read/write integration, QA sampling pipeline, distribution monitoring |
| Multi-site / regulated build | $160k–$400k+ | 14–24 weeks | HIPAA posture, human-in-the-loop controls, immutable audit logging, multi-entity governance |
Frenchy Digital 2026 cost bands for AI coding and healthcare automation engagements.
Senior-led delivery is priced at $150–$225 per hour, with ongoing retainers from $2,500 to $9,500 per month covering model and code-set revalidation, QA sample review, monitoring, and incident response. Every engagement carries a 30-day post-launch warranty, and you receive a written scope and fixed-price phased proposal within 5 business days of the discovery call.
Limitations and Failure Modes — Read This Before You Buy
An honest accounting of what does not work, including the findings that cut against the category.
- The evidence base is vendor-controlled: No independent benchmark of AI coding accuracy exists. Every figure in every deck is a claim. This is the single largest limitation and it does not have a technical fix.
- An agent cannot fix bad documentation: It inherits it. If the note does not support the level of service, the correct behavior is a query, not an inference — and configuring an agent to infer the missing element is how a tool becomes a liability.
- Adjacent AI results are mixed, not uniformly positive: In a randomized trial published in NEJM AI in December 2025 covering 238 physicians across 14 specialties, one ambient product reduced time-in-note by 9.5% versus control while another showed no significant reduction at all. Independent evidence in this family of tools is genuinely mixed.
- Large deployments have underdelivered on productivity: The Permanente Medical Group's large-scale ambient deployment reported roughly 18 seconds saved per appointment versus non-users, and an Intermountain Health matched-cohort analysis found no statistically significant productivity gains — both reported in npj Digital Medicine in 2026.
- Complex office E/M remains a human decision: MDM inference from narrative is the hardest part of ambulatory coding and the part most exposed to enforcement. Autonomy here is not defensible in 2026.
- Modifier logic is brittle: Modifiers 22, 25 and 59, global periods, and staged procedures depend on context that frequently is not in the operative note at all.
- Payer-specific rules move constantly: Coverage policies and edits change per payer, per quarter, and an agent trained on national conventions will be confidently wrong on regional plans.
- Training and change management are underfunded: KLAS Arch Collaborative 2026 research found under 25% of AI-adopting clinicians said they received adequate training, and satisfaction plateaued past roughly four AI tools. The fifth tool is not free.
- The control program is a permanent cost: Audits, QA sampling, monitoring, and annual revalidation recur forever. A vendor quote that omits them understates total cost of ownership.
- Stacked AI compounds error: An ambient-generated note feeding a coding agent means documentation defects propagate into claims. Audit the note-to-encounter link, not just the code-to-note link.
Sources for the mixed findings above: the December 2025 randomized trial is reported in NEJM AI; the Permanente and Intermountain results are reported in npj Digital Medicine; the training and tool-saturation findings come from KLAS Arch Collaborative. We cite them because they are the closest independent evidence available on AI in clinical documentation, and because the pattern they show — real benefit, unevenly distributed, smaller than marketed — is the prior you should carry into an AI coding evaluation.
None of this argues against deployment. It argues for deploying narrowly, measuring your own results, and funding the control program as a first-class part of the build rather than an afterthought bolted on when someone asks a hard question.
Red Flags When Buying an AI Coding Vendor
Take this list into the demo. Every item has cost a real practice real money.
| Red flag | Why it matters |
|---|---|
| Publishes an accuracy rate without case mix, denominator, or adjudication method | The number is unfalsifiable and cannot support a compliance decision |
| Will not let you run a blinded audit on your own claims before you sign | You are being asked to accept their measurement instead of yours |
| Describes CPT Appendix S as though it creates reimbursement for AI | Either they do not understand the taxonomy or they are counting on you not to |
| Sells 'autonomous coding' across all specialties | Autonomy is defensible in a few structured service lines, not everywhere |
| Confidence threshold is set by the vendor and not exposed to you | You own the compliance consequence of a dial you cannot see |
| No calibration data behind the confidence score | An uncalibrated score is a ranking, not a probability |
| Prices on a share of incremental coding-level lift | That contract structure pays the vendor to move your codes upward |
| No BAA at the tier you are buying, or vague model-training terms | Your PHI exposure is contractual before it is technical |
| Cannot produce per-encounter reviewer dwell time | Then you cannot evidence that human review happened |
| Promises reimbursement is coming for AI services | CMS solicited comment in CY2026 and finalized nothing |
Frenchy Digital buyer's red-flag checklist for AI medical coding vendors, 2026.
Ask one question and watch what happens: "Will you support a blinded audit of your output on our claims, scored by our coders, before we sign?" A vendor confident in their product says yes immediately. The answer tells you more than the entire demo.
— Frenchy Digital buyer's principle
Working with Frenchy Digital
Frenchy Digital is a senior-led, Black-owned Los Angeles agency that builds administrative AI systems for regulated businesses. In healthcare that means documentation and revenue-cycle workflows under human governance — never clinical decision-making, never a claim that our software is FDA-cleared, and never the phrase "HIPAA-certified," because no such certification exists.
- Audit before automation: Every coding engagement opens with a blinded baseline audit of your current output and a distribution snapshot. We will tell you if the answer is that you should not automate this service line yet.
- Senior engineers only: No junior staffing on regulated builds. The person designing your audit sampling is the person who has done it before.
- Fixed-price phased scope: Written scope and fixed-price phased proposal within 5 business days of the discovery call. Phases are independently cancellable.
- Compliance controls as deliverables: Audit logging, reviewer dwell-time capture, override-rate reporting, and the QA sampling pipeline ship as part of the build, not as a later add-on.
- Your counsel decides the legal questions: We build the evidence pipeline and hand over the file. Sample sizing for confirmatory audits, extrapolation, and disclosure decisions belong to your compliance officer and healthcare counsel.
- Full ownership transfer: Source code, prompts, configurations, infrastructure accounts, and IP transfer to your practice at delivery under clean work-for-hire terms. No vendor lock-in.
- 30-day warranty and optional retainer: Every engagement carries a 30-day post-launch warranty; retainers from $2,500 to $9,500 per month cover revalidation, QA review, monitoring, and incident response.
Book a discovery call at calendly.com/frenchydigital/discovery-call or call +1 (424) 272-5601. If you want to see the audit design before you talk to us about building anything, ask for it on the call — we will walk through it whether or not you hire us.
Audit Before You Automate
Book a free 60-minute discovery call with Frenchy Digital, a senior-led Black-owned LA agency. You leave with a written audit design for your own claims and a fixed-price phased proposal within 5 business days.
Audit Before You Automate
Book a free 60-minute discovery call with Frenchy Digital. You leave with a written audit design for your own claims and a fixed-price phased proposal within 5 business days.
1517 S Bentley Ave Unit 204, Los Angeles CA 90025
Frequently Asked Questions
Sources & References
- 1AMA — CPT Appendix S: Taxonomy for AI in Medical Services and Procedures↗
- 2HHS OIG — Industry Segment-Specific Compliance Program Guidance (Medicare Advantage ICPG, Feb 3, 2026)↗
- 3CMS — Medicare Physician Fee Schedule↗
- 4CMS — Interoperability and Prior Authorization Final Rule (CMS-0057-F)↗
- 5Federal Register — HIPAA Security Rule NPRM, 90 FR 898 (Jan 6, 2025)↗
- 6FDA — Clinical Decision Support Software, final guidance (re-issued Jan 29, 2026)↗
- 7HHS — HIPAA for Professionals↗
- 8ASTP/ONC — Decision Support Interventions certification criterion (HTI-1)↗
- 9KFF — Medicare Advantage prior authorization determinations, 2024 data (Jan 28, 2026)↗
- 10KFF — Claims denials and appeals in ACA Marketplace plans, 2024 (Mar 24, 2026)↗
- 11AMA — 2026 Physician AI Sentiment Report↗
- 12Premier Inc. — Claims adjudication costs providers $25.7B↗
- 13U.S. Bureau of Labor Statistics — Medical Records Specialists↗
- 14MGMA — Cost efficiency with medical group staffing↗
- 15NIST — AI Risk Management Framework↗
- 16NEJM AI — Randomized trial of ambient documentation tools (Dec 2025)↗
- 17npj Digital Medicine — Ambient documentation outcomes across large deployments (2026)↗
- 18Peer-reviewed review of 356 ambient-generated clinical notes — omissions and hallucinations↗
- 19JMIR Medical Informatics — Propagation of interpreter errors by ambient scribes (2026)↗
- 20KLAS Arch Collaborative — Ambient Speech Outcomes 2025↗
- 21AAPC — AMA releases CPT 2026↗
- 22Holland & Knight — CMS releases CY2026 Medicare Physician Fee Schedule final rule↗
- 23Ballard Spahr — DOJ's health care fraud takedown spotlights AI and data analytics↗

