Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    Medical Coding
    August 9, 2026
    26 min read

    AI Agents for Medical CodingCPT & ICD-10 in 2026

    Where autonomous coding genuinely works, where it does not, and how to deploy it without creating an enforcement problem — including the sampled audit that replaces vendor accuracy claims with a number you can actually defend.

    AI agents for medical coding — CPT and ICD-10 automation, accuracy auditing, and compliance controls in 2026
    0
    Independent benchmarks of AI coding accuracy
    No public independent benchmark located, Aug 2026
    418
    CPT 2026 editorial changes (288 new, 84 deleted, 46 revised)
    AMA CPT 2026
    Feb 3, 2026
    OIG names AI-generated EHR prompts as potentially abusive
    HHS OIG Medicare Advantage ICPG
    $25k–$160k
    Typical coding-agent build band
    Frenchy Digital scoping 2026

    Key Takeaways

    • No independent benchmark of AI coding accuracy exists as of August 2026. Every published figure — 92–97% on structured encounters, 82–90% on complex inpatient, 95–98% for human coders post-QA — is vendor-published on a vendor-selected sample. Treat all of them as claims, not measurements.
    • The only accuracy number that can support a compliance decision is the one your own blinded, sampled audit produces on your own claims, with a human baseline measured on the same sample.
    • Autonomous coding is genuinely viable in radiology, pathology, and the emergency department — high volume, structured, templated documentation. Complex office visits are not, because level of service depends on medical decision making inferred from narrative.
    • Confidence routing is the standard architecture, and the confidence threshold is a business decision with compliance consequences. It needs a named owner, written rationale, change control, and a calibration curve measured against your own audit.
    • CPT Appendix S is a taxonomy — assistive, augmentative, autonomous, Levels I–III by degree of physician override — not a payment schedule. There is no general AI CPT payment category, and CMS only solicited comment on paying for SaaS and algorithms in the CY2026 PFS without finalizing a pathway. AI coding is a cost-side play.
    • The OIG's February 2026 Medicare Advantage guidance names AI-generated EHR prompts to add risk-adjusting diagnoses as potentially abusive. Accepting ~85% of suggestions at ~0.5 seconds per chart is not meaningful review — it is documented automation bias.
    • An agent optimized for revenue capture is an agent optimized toward the exact distribution shift enforcement looks for. Baseline your level-of-service and diagnosis-density distributions before deployment and monitor drift monthly.
    • Frenchy Digital scopes this work from $9k–$22k for discovery through $160k–$400k+ for regulated multi-site builds, senior-led at $150–$225/hr, with full source-code and IP ownership transferred at delivery.

    Where Autonomous Medical Coding Actually Works in 2026

    Physician adoption of AI is no longer the question. The AMA's 2026 Physician AI Sentiment Report, fielded in January and February 2026 with 1,692 respondents, found 81% of physicians using AI professionally, up from 38% in 2023. The question is which workflows can carry autonomy and which cannot — and in coding, the honest answer is narrower than the marketing.

    Autonomous medical coding is real in 2026, and it is much narrower than the category name suggests. The service lines where practices are genuinely letting an agent assign CPT and ICD-10 codes without a coder touching every chart are the same three they were eighteen months ago: diagnostic radiology, pathology, and the emergency department. What those three share is not a technology property. It is a documentation property.

    In each, the volume is high, the documentation is structured or templated, the code space for a given study or presentation is comparatively bounded, and — this is the decisive part — the coding decision is largely a transformation of something the report already states explicitly. A chest CT with contrast says so. A synoptic pathology report names the specimen. That is extraction with rules on top, and language models are good at extraction with rules on top.

    Complex office visits fail every one of those tests. The level of service for a 99214 versus a 99215 turns on medical decision making, and medical decision making is not stated in the note — it has to be inferred from narrative, from what the physician considered and rejected, from risk that is often implicit. An agent that infers MDM from narrative is not extracting; it is judging. That is a different reliability regime, and it is why the honest deployments in 2026 keep office E/M in assisted mode with a human deciding the level.

    Service lineDocumentation profileDefensible 2026 postureWhere it breaks
    Diagnostic radiologyStructured report, templated impression, bounded code space per study typeAutonomous is defensible for routine studies behind a conservative thresholdModality/contrast/laterality errors; unsigned or amended reports
    PathologyHighly templated synoptic reporting; specimen-driven code selectionAutonomous is defensible for routine specimensMulti-specimen accessions; add-on stains ordered after the fact
    Emergency departmentHigh volume, structured templates, repeatable presentation patternsAutonomous viable for lower-acuity, routed for critical care and proceduresCritical care time, procedure bundling, observation status
    Urgent careSemi-structured; narrow but not bounded code spaceAssisted with human sign-off; autonomy only after a clean audit cycleE/M level drift; injection and procedure add-ons
    Office / outpatient E/MNarrative documentation; level of service turns on medical decision makingAssisted only. Human coder decides the levelMDM inference; time-based billing; social determinants capture
    Surgical / proceduralOperative notes vary by surgeon; heavy modifier logicAssisted only; global period and modifier rules need a humanModifier 22/25/59, global period, staged procedures
    Inpatient / DRGLong, multi-author records; sequencing drives paymentAssisted only; CDI and coder judgment dominatePrincipal diagnosis sequencing; POA indicators; query workflow
    Behavioral healthNarrative-heavy; time and modality dependentAssisted only; consent and modality rules are practice-specificTime thresholds; telehealth modality; interactive complexity

    Service-line viability for autonomous coding, 2026. Frenchy Digital deployment guidance.

    The framing that matters: autonomy is not a property of the model. It is a property of the documentation the model is reading. Ask "how bounded is the code space for this encounter type?" before you ask "how good is the AI?"

    One clarification before any of the rest of this is useful. Everything in this article concerns administrative and documentation workflows under human governance. A coding agent assigns billing codes from documentation a clinician already authored. It does not diagnose, it does not treat, and it does not practice medicine. The AMA's framing of "augmented intelligence" is the correct one, and it is also the one that keeps you inside the lines.

    The Evidence Problem: No Independent Benchmark of AI Coding Accuracy Exists

    Here is the fact that should shape every purchasing conversation in this category, stated as plainly as we can state it: as of August 2026 there is no public, independent benchmark of AI medical coding accuracy. Not a peer-reviewed head-to-head. Not a neutral third-party test set. Not a standards body evaluation. Every accuracy figure circulating in this market is vendor-published, generated on a vendor-selected sample, scored against a vendor-defined notion of correctness.

    That is not an accusation of dishonesty. It is a structural observation. Ambient documentation, an adjacent category, has accumulated real independent evidence — randomized trials, matched-cohort studies, multi-site outcome reports. Coding has not. The category grew fast, the buyers were operators rather than researchers, and nobody built the benchmark.

    For completeness, the figures vendors most commonly publish are these, and we label them as what they are:

    • 92–97% accuracy on structured, high-volume encounters (vendor claim): Typically cited for emergency department and outpatient radiology. The measurement conditions are almost never disclosed.
    • 82–90% accuracy on complex inpatient coding (vendor claim): Note that inpatient accuracy is a different construct — sequencing and present-on-admission indicators drive payment as much as code selection does.
    • 95–98% accuracy for human coders post-QA (vendor-published comparator): This is the benchmark vendors compare themselves against. It is also vendor-published, and 'post-QA' means after a QA process that may not resemble yours.

    A practice cannot verify any of those numbers, and the reason is worth spelling out because it is the whole argument. An accuracy rate is meaningless without four disclosures that vendors essentially never make in writing:

    1. 1.The case mix: Which specialties, which payers, which acuity distribution. A 96% on routine screening mammography and a 96% on multi-trauma ED are not the same 96%.
    2. 2.The denominator: Primary code only, or every line? Are modifiers scored? Secondary diagnoses? Were low-confidence encounters excluded from the sample before the number was computed?
    3. 3.The definition of correct: Exact match to a reference coder, or 'payable and defensible'? Those diverge sharply on E/M levels and on diagnosis specificity.
    4. 4.The adjudication method: One reviewer, or two blinded reviewers with a third adjudicating disagreement? A single-reviewer ground truth measures agreement with one person's habits.
    Claim you will hearWhat it actually isWhat is missingWhat to do instead
    AI 92–97% accuracy on structured high-volume encountersVendor-publishedCase mix, denominator, definition of correct, adjudication methodRun the same measurement on your own claims
    AI 82–90% accuracy on complex inpatientVendor-publishedWhether sequencing and POA were scored, or only code presenceScore sequencing separately from code selection
    Human coders 95–98% accurate post-QAVendor-published comparatorWhether post-QA means after the same QA you actually runMeasure your own coders on the same blinded sample
    Denials reduced from 18% to 3%Vendor marketingBaseline period, payer mix, whether policy changes were concurrentTreat as marketing; do not put it in a board deck
    4–7x coder productivityVendor marketingWhether review time and rework were counted in the denominatorMeasure charts-per-coder-hour including rework
    Confidence score of 0.94Vendor-internalWhether the score is calibrated against observed accuracyDemand a calibration curve from your own audit data

    Vendor claims in the AI coding market and the disclosures that would make them checkable, 2026.

    An unfalsifiable number cannot support a compliance decision. If you cannot describe how a figure could be shown to be wrong, do not put it in a board deck, a payer conversation, or a policy document.

    Frenchy Digital evaluation principle

    The practical consequence is not "do not buy." It is that the accuracy number you rely on has to be one you generated. That is the rest of this article's real subject, and it is a solvable problem — the audit costs a few thousand dollars in coder time and takes a couple of weeks. What it buys you is a defensible figure and, more importantly, a defensible process.

    How to Run Your Own Sampled Accuracy Audit

    This is the part of the engagement that most practices skip and that pays for itself the first time a payer asks a question. The design below is the one Frenchy Digital runs before we let any coding agent post a claim. It is deliberately unglamorous.

    Design parameterFrenchy Digital defaultWhy it is designed this way
    Unit of analysisThe claim line, rolled up to the encounterEncounter-level agreement hides modifier and secondary-diagnosis errors
    Ground truthTwo credentialed coders working blinded from the documentation; a third adjudicates disagreementsA single reviewer measures agreement with one person, not accuracy
    BlindingReviewers never see the AI output or its confidence scoreSeeing the suggestion is exactly the automation bias you are trying to measure
    Human baselineScore your own coders on the same sample in the same passWithout a baseline you cannot tell whether the agent is better or worse than today
    StratificationSample from every confidence band, every service line, and every top payerAuto-accepted encounters are the ones nobody is otherwise looking at
    Initial sampleFrenchy Digital default: 30 charts per stratum as a screening probeCheap enough to run pre-deployment; large enough to surface systematic errors
    Confirmatory sampleExpand to a statistically valid sample, sized with compliance counsel, if the probe failsExtrapolation and repayment exposure are legal decisions, not engineering ones
    Error taxonomyWrong primary code; wrong E/M level; missing secondary dx; modifier error; laterality; bundling; medical-necessity linkageA single accuracy percentage tells you nothing about what to fix
    Directional metricNet RVU delta versus adjudicated truth, reported separately from raw accuracyTwo systems can be 94% accurate with opposite financial risk profiles
    CadenceBaseline pre-deployment, monthly for the first quarter, then quarterlyModel, prompt, and code-set changes all invalidate a stale audit
    RetentionEvery sampled chart, every reviewer decision, every adjudication, versionedThis file is your evidence that review actually happened

    Sampled accuracy audit design for an AI coding agent — Frenchy Digital methodology, 2026.

    The two steps everyone leaves out

    Measure your own coders on the same sample. Practices routinely audit the AI against a reference standard and never audit their humans against the same standard in the same pass. Without that, you have no idea whether an 89% agent is a downgrade or an upgrade. In several engagements the human baseline came in materially below the number the practice assumed it was running at — which changed the decision entirely.

    Sample from the auto-accepted bucket. The encounters that never reach a human are, by definition, the ones no control is examining. If your QA sample is drawn only from the human-reviewed queue, you have built a monitoring program that structurally cannot see the risk you created. Every monthly sample must be stratified across confidence bands, with the top band included.

    Report accuracy and net RVU delta as two separate numbers. Two agents can both be 94% accurate while one drifts revenue upward and the other drifts it downward. Only one of those creates enforcement exposure, and a single accuracy percentage cannot tell them apart.

    On sample sizing, a note about scope. The 30-charts-per-stratum screening probe above is our engineering convention for detecting systematic error cheaply, not a legal standard. If a probe surfaces a real error pattern, the next step — how large a confirmatory sample, whether findings get extrapolated, and what any of it obligates you to do — is a question for healthcare counsel and your compliance officer, not for your development partner. We build the pipeline and hand over the file; we do not size a repayment exposure.

    Sequence that works: baseline audit before automation → shadow mode where the agent codes but nothing posts, scored weekly → limited autonomous pilot in one service line at a conservative threshold → monthly stratified QA → staged expansion. Any vendor whose implementation plan skips shadow mode is optimizing for their time-to-value, not your risk.

    Confidence Routing and the Threshold Trap

    Essentially every serious 2026 architecture in this space works the same way. The agent produces a code set and a confidence score. Output above a threshold posts automatically. Output below it queues for a human coder. This is a good design — it concentrates scarce coder attention on the encounters where judgment actually pays.

    The trap is what the threshold is. Raise it and more work routes to humans: higher labor cost, lower error exposure. Lower it and the reverse. That means the threshold is a dial that trades money against compliance risk, and in most implementations it is set once during onboarding by someone in operations, never written down, and quietly adjusted later when the coding queue backs up.

    State it plainly: the confidence threshold is a business decision with compliance consequences. It needs a named owner, a written rationale, version history, and a review cadence — the same treatment you would give a charge master change. An undocumented threshold adjustment that coincides with a revenue shift is very difficult to explain after the fact.

    There is a second, subtler problem: confidence scores are usually uncalibrated. A score of 0.94 is a ranking signal from inside the model. It is not a claim that 94% of such outputs are correct — unless someone has measured observed accuracy per confidence bucket and shown that the curve tracks. Almost no vendor publishes that curve. Your own audit can produce it: bucket your sampled encounters by the confidence the agent assigned, compute adjudicated accuracy within each bucket, and plot it. If observed accuracy in the 0.90–0.95 band is 78%, your threshold is not where you think it is.

    Finally, set the threshold asymmetrically. A code that increases reimbursement relative to your historical distribution should face a higher bar than a code that decreases it. This is not conservatism for its own sake — it directly reduces the pattern described in the enforcement section below.

    Routing tierTriggerControl
    Auto-acceptHigh confidence, low-risk service line, no financial upside vs. adjudicated baselineSampled QA at a fixed monthly rate, no exceptions
    Auto-accept with holdHigh confidence but code increases reimbursement versus the prior distributionSecond-pass review before the claim posts
    Route to coderBelow threshold, or any modifier logic, or any add-on procedureFull human coding; the AI suggestion is visible only after the coder commits
    Route to CDI / queryDocumentation does not support any defensible codePhysician query workflow; never let the agent infer the missing element
    Hard stopNew or revised code effective within the last 90 days; unsigned note; amended noteHuman only until the next revalidation cycle clears it

    Confidence routing tiers for a coding agent — Frenchy Digital reference configuration, 2026.

    CPT Appendix S: A Taxonomy, Not a Payment Schedule

    CPT Appendix S, introduced by the AMA in 2021, is the correct vocabulary for this category, and it is routinely misdescribed by vendors. Appendix S classifies AI-enabled medical services and procedures into three categories — assistive, augmentative, and autonomous — corresponding to Levels I, II, and III, distinguished by the degree to which a physician must interpret the output and can override it.

    Appendix S categoryWhat it describesLevelWhat it means for a coding agent
    AssistiveSoftware detects clinically relevant data without independent analysis or conclusions; the work product requires physician interpretation and actionLevel IMost computer-assisted coding sits here in practice
    AugmentativeSoftware analyzes or quantifies data to yield clinically meaningful output; the physician still interprets and actsLevel IIWhere a coding agent that proposes and explains codes belongs
    AutonomousSoftware interprets data and independently generates clinically meaningful conclusions; conclusions may be subject to physician overrideLevel IIIThe category vendors invoke; the override rights inside it are what matter
    What it is notA payment schedule, a fee category, or a route to reimbursement for AI softwaren/aNo general AI CPT payment category exists in 2026
    May 2026 revisionsCPT Editorial Panel accepted revisions replacing 'machine' with 'software output(s)'n/aLanguage now tracks software behavior rather than hardware
    In developmentClinically Meaningful Algorithmic Analyses (CMAA) frameworkn/aWatch it, but do not build a business case on it yet

    CPT Appendix S taxonomy for AI-enabled services, with the May 2026 CPT Editorial Panel revisions. Source: AMA.

    At the May 2026 CPT Editorial Panel meeting, revisions were accepted that replace the word "machine" throughout Appendix S with "software output(s)" — a small edit with a real effect, since it moves the taxonomy's language from hardware to software behavior, which is where this category actually lives. The AMA is separately developing a framework for Clinically Meaningful Algorithmic Analyses (CMAA). That work is worth tracking. It is not worth building a business case on today.

    The thing vendors obscure: Appendix S is a taxonomy. It is not a payment schedule, it does not create billing codes for AI, and being classified as "autonomous" under Level III does not make anything payable. There is no general AI CPT payment category in 2026. A pitch deck that presents Appendix S as evidence that reimbursement for AI has arrived is either misinformed or counting on you to be.

    What Actually Changed in CPT 2026 — and Why It Is an Agent Problem

    The CPT 2026 code set, effective January 1, 2026, carried 418 editorial changes: 288 new codes, 84 deletions, and 46 revisions (as reported on the AMA release). For a practice, that is an annual operational chore. For an AI coding agent, it is a versioning discipline that most implementations do not have.

    CPT 2026 itemDetailAgent implication
    Total editorial changes418 changes effective January 1, 2026Every agent configuration and code-set mapping needs annual revalidation
    New codes288New codes are the highest-risk surface for an agent: no training precedent
    Deleted codes84A deleted code that survives in a mapping table becomes a rejected claim
    Revised codes46Revisions are the silent failure: the code still exists, the meaning moved
    Office/outpatient E/M 99202–99215UnchangedThe hardest coding decision in the practice did not get easier
    RPM/RTM restructuringNew 99445 and 99470; 99453/99454/99457/+99458 revisedIf you bill remote monitoring, this is the section to re-verify first
    99453 setup thresholdDropped from 16 days of data to 2 daysMaterially changes eligibility logic that an agent may have hard-coded
    99470First 10 minutes of treatment management; requires at least one real-time audio/video interactionAn agent cannot infer the interactive requirement from device data alone

    CPT 2026 changes most relevant to an automated coding pipeline. Effective January 1, 2026.

    Two things stand out. First, the core office and outpatient E/M codes 99202–99215 were unchanged, still driven by medical decision making or total time. The single hardest coding judgment in an ambulatory practice did not get any easier in 2026, which reinforces the point from the first section about where autonomy belongs.

    Second, the notable restructuring landed in remote physiologic and remote therapeutic monitoring. New codes 99445 and 99470 arrived, 99453/99454/99457/+99458 were revised, and the 99453 setup threshold dropped from 16 days of data to 2 days. That last change is exactly the kind of parameter an integration hard-codes and then forgets. If your agent has a rule that counts to sixteen, it has been wrong since January.

    The general lesson: revised codes are more dangerous than deleted ones. A deleted code produces a rejection you will notice. A revised code still exists, still passes edits, and now means something slightly different — and an agent configured against last year's definition will keep submitting it confidently. Build an annual revalidation into the operating cadence, and treat any code effective within the last 90 days as human-only until it clears.

    What Gets Paid For AI in 2026: The Honest Answer

    Every ROI conversation in this category eventually reaches the same question: does anyone pay for this? The answer for AI coding in 2026 is no, and the specific shape of that "no" is worth understanding because vendors work hard around its edges.

    In the CY2026 Medicare Physician Fee Schedule final rule, CMS solicited comment on how software-as-a-service and algorithm-based services should be paid, noting that they are not well accounted for under current payment methodology. Soliciting comment is not finalizing a pathway. No general AI payment mechanism was established (summarized here).

    Item2026 statusWhat it actually is
    A CPT payment category for AI servicesDoes not existCPT Appendix S is a taxonomy only
    Medicare PFS payment for AI coding softwareDoes not existCY2026 PFS solicited comment on SaaS/algorithm payment; nothing finalized
    New Technology Add-on Payment (NTAP)Inpatient onlyIrrelevant to outpatient practice economics
    Digital Mental Health Treatment G-codesG0552, G0553, G0554 exist; CMS finalized expansion to ADHDDevice supply and treatment management, not AI coding
    Remote monitoring codes99445, 99470 and the revised RPM/RTM familyPaid for the clinical service, not for the software that codes it
    AI coding, ambient scribes, AI denial toolsNo CPT or CMS reimbursementCost-side ROI only. Build the business case on labor and rework

    What Medicare and CPT do and do not pay for in relation to AI, as of August 2026.

    The adjacent codes that did arrive are for clinical services, not for software. The restructured remote monitoring family pays for monitoring a patient. The Digital Mental Health Treatment codes G0552, G0553 and G0554 pay for device supply and treatment management, with CMS finalizing expansion to ADHD. New Technology Add-on Payment is an inpatient mechanism and does not touch outpatient practice economics at all. None of these pay you for running an AI coder.

    Conclusion, stated cleanly: AI coding is a cost-side play. The business case is coder labor, rework avoidance, days in A/R, and denial reduction — never a new revenue code. Any model that assumes reimbursement is coming should be rebuilt without that assumption, because as of August 2026 it does not exist.

    The Compliance Core: OIG, DOJ, and the Automation-Bias Theory

    This is the section that should determine whether you deploy at all, and how.

    In its Medicare Advantage Industry Segment-Specific Compliance Program Guidance, published February 3, 2026, the HHS Office of Inspector General identified as potentially abusive the practice of "querying physicians via electronic medical record platforms, including prompts generated by artificial intelligence algorithms, to add risk-adjusting diagnoses." Read that carefully. The object of concern is not a fraudulent diagnosis. It is the prompt — the mechanism by which a physician is nudged toward a revenue-relevant code. That is the clearest federal statement to date that AI inside the documentation workflow can itself be the compliance problem.

    On the enforcement side, the Department of Justice's 2026 National Health Care Fraud Takedown foregrounded AI and data analytics, and the Health Care Fraud Unit's Data Fusion Center — operating with HHS-OIG and the FBI — is now standing infrastructure rather than a pilot. The asymmetry is worth internalizing: the government is doing pattern detection at scale on claims data. If your agent creates a pattern, the pattern is visible.

    Which brings us to the theory that should shape your controls. Practices deploying these tools almost universally cite human-in-the-loop review as their mitigating control. The enforcement counter-argument is straightforward: a reviewer accepting roughly 85% of AI suggestions at roughly half a second per chart is not reviewing anything. That is automation bias, and it converts your stated control into documented evidence that the control was not operating. Your own telemetry becomes the exhibit.

    Human-in-the-loop is a compliance control, not a courtesy. If you cannot produce per-encounter review time and per-reviewer override rates, you do not have a human in the loop — you have a human in the log.

    Frenchy Digital compliance principle
    ControlWhat it producesWhy enforcement cares
    Documented review timePer-encounter dwell time for every human reviewer, logged automaticallyThe OIG theory of the case is that instantaneous acceptance is not review
    Override and rejection ratesPer reviewer, per service line, per confidence band, trended monthlyA reviewer at a 99% acceptance rate is a finding, not a top performer
    Immutable audit logsWho saw what, when, what the AI proposed, what was submitted, what changedReconstructing a claim two years later is the entire ballgame
    Periodic sampled QABlinded re-coding of a monthly sample including auto-accepted encountersAuto-accepted work is invisible to every other control you have
    Distribution monitoringE/M level mix, RVU per encounter, diagnosis density, HCC capture rateEnforcement analytics look at distributions, so you should look first
    Physician query integrityNo AI-authored prompt that suggests a specific risk-adjusting diagnosisNamed in the OIG's February 2026 Medicare Advantage guidance
    Change controlVersion and date every model, prompt, threshold, and code-set updateAn unversioned threshold change is an unexplainable revenue shift
    Vendor due diligence fileBAA, retention terms, training-use terms, subprocessors, incident historyYour compliance posture inherits your vendor's

    Minimum control set for an AI coding deployment — Frenchy Digital implementation standard, 2026.

    Two structural notes. Any vendor touching protected health information is a business associate under HIPAA and requires a signed BAA before a chart moves; there is no such thing as a "HIPAA-certified" vendor, only a compliant posture with evidenced Security Rule safeguards. And the HIPAA Security Rule NPRM published January 6, 2025 — proposing mandatory MFA, encryption at rest and in transit, asset inventories, annual audits, and 72-hour restoration — remains proposed, not final; the Unified Agenda now projects final action in 2027. Build toward it as a design target, but do not let anyone sell it to you as a current obligation.

    If your coding agent is embedded in certified health IT, one more artifact is available free: the ASTP/ONC Decision Support Intervention transparency data. Predictive DSIs in certified products must publish 31 source attributes covering development details, fairness, external validation, quantitative performance, and maintenance. That is ready-made vendor due diligence a practice can read before it signs anything.

    Two boundaries worth knowing so you do not over- or under-scope your compliance work. A coding agent is administrative software; it is generally not a medical device, and the FDA's Clinical Decision Support Software final guidance re-issued January 29, 2026 — which added an enforcement-discretion policy for software producing a single, clinically appropriate output — governs clinical decision support, not billing automation. Separately, CMS-0057-F binds payers rather than practices, and its FHIR API obligations do not arrive until January 1, 2027. For governance structure in the meantime, the NIST AI Risk Management Framework is the most usable non-binding scaffold we have found for documenting how you govern a model you did not build.

    Upcoding: Note the Direction of the Risk

    This deserves its own section because the risk has a direction, and the direction is not symmetric.

    An AI coding agent tuned to "capture revenue that is being left on the table" is, by construction, tuned toward higher-level codes and denser diagnosis capture. That is the product working as specified. It is also, precisely, the pattern that payer analytics and federal data-mining programs are built to find. An agent optimized for revenue capture is an agent optimized toward the exact signature enforcement looks for.

    The exposure is rarely any single code. Individual codes are usually defensible in isolation. The exposure is distribution shift — your 99214/99215 mix moving three points, your diagnoses-per-encounter rising fifteen percent, your HCC capture rate stepping up on a date that happens to be your go-live. Aggregate patterns are what get investigated, and "the software suggested it" is not a defense that has worked well historically.

    Metric to baselineHow to measure itWhat should trigger review
    E/M level distributionShare of 99213 / 99214 / 99215 by providerAny provider whose level-4+ share moves more than a few points post-deployment
    RVU per encounterWork RVUs per completed encounter by service lineA step change that starts on the deployment date
    Diagnosis densityAverage diagnoses submitted per encounterDensity rising without a change in panel acuity
    Risk-adjusting diagnosis captureNew HCC-mapped diagnoses added at the coding stepDiagnoses appearing at coding that never appear in the assessment
    Modifier frequencyRate of modifier 25 and 59 useModifier use rising to clear edits rather than to describe the encounter
    Override rate by reviewerPercentage of AI suggestions changed before submissionRates near zero, and rates that fall over time as reviewers habituate

    Distribution-drift monitoring for an AI coding deployment. Baseline before go-live; trend monthly.

    Two contract-level protections follow. First, do not price the vendor on a share of incremental coding-level lift. Contingency on collections is an established RCM structure; contingency specifically on moving your codes upward pays someone else to take your compliance risk. Second, require that the agent never authors a physician query that names a specific risk-adjusting diagnosis — the mechanism the OIG called out in February. An agent may flag that documentation is insufficient. A human decides what to ask.

    Reference Architecture for a Coding Agent

    What we actually build, in the order we build it. None of this is exotic; the discipline is in refusing to skip steps under delivery pressure.

    • Baseline audit first: Blinded dual-coder audit of your current output, with the human baseline and distribution snapshot captured before any automation exists. This artifact is your before-picture and you only get one chance to take it.
    • Scoped EHR integration: Read access limited to the documentation the agent needs, write access limited to the coding fields it owns. Minimum necessary is an architecture constraint, not a policy sentence.
    • Deterministic rules ahead of the model: Payer edits, NCCI-style bundling logic, code-set validity, and effective-date checks run as code. Models are for extraction and inference; rules are for things that must be right.
    • Confidence scoring with an owned threshold: Threshold stored as a versioned configuration value with an owner, a rationale, and an approval trail. Not a hidden vendor default.
    • Asymmetric routing: Higher bar for codes that increase reimbursement relative to your baseline distribution than for codes that decrease it.
    • Immutable audit log: Append-only record of the source documentation version, the model and prompt version, the proposal, the confidence, the reviewer, the dwell time, and the submitted result.
    • Sampling pipeline built on day one: Monthly stratified sample generated automatically across confidence bands and service lines, routed to blinded reviewers. If QA is manual, QA stops happening in month three.
    • Distribution monitoring and alerting: Automated trending of E/M mix, RVU per encounter, and diagnosis density, with thresholds that page a human rather than a dashboard nobody opens.
    • Code-set version control: Annual revalidation against the new CPT and ICD-10 releases, with any code effective in the last 90 days routed to humans until cleared.
    • Rollback path: A single switch that returns a service line to full human coding without a deployment. You will use it, and the day you need it is not the day to build it.

    On stacking an ambient scribe under a coding agent

    A large share of 2026 deployments feed an AI-generated note into an AI coding agent. That composition deserves explicit attention, because the coding agent inherits every documentation defect upstream of it.

    A published review of 356 ambient-generated notes found omissions in 18% and hallucinations in 11.5%, with most errors mild to moderate but 5.3% of notes containing errors rated potentially seriously harmful. A 2026 JMIR research letter separately documented ambient scribes propagating interpreter errors. A coding agent reading those notes cannot distinguish a hallucinated finding from a real one — it will code what the note says.

    If you run both, the physician attestation on the note is doing more work than anyone acknowledges, and your audit sample should include a step that compares the note against the encounter rather than only the code against the note.

    Building the Cost-Side ROI Model Honestly

    Since nothing pays for AI coding directly, the business case has to be built on cost. Here is how to build one that survives scrutiny — and which inputs are solid versus soft.

    The solid inputs: the U.S. Bureau of Labor Statistics puts median pay for medical records specialists, which includes billing and coding roles, at $50,250 in 2024. MGMA reports that automation is the top cost-cutting move groups planned for 2026, cited by 36%, well ahead of hiring freezes at 18%. On the denial side, Premier Inc. found claims adjudication cost providers $25.7 billion in 2023, with roughly $18 billion potentially wasted on claims that should have paid on submission, at an average adjudication cost of $57.23 per claim — 2023 data with no newer primary update we could locate.

    The soft inputs, labeled: the commonly quoted fully loaded cost of $60,000–$80,000 per in-house biller including benefits and software is a vendor-relayed figure, not a survey result — use it as a planning range, not a citation. Outsourced RCM pricing of 4–10% of net collections is likewise vendor-quoted, though consistent enough across sources to plan against. The frequently cited HFMA "cost to collect under 2% of net collections" benchmark reaches us through secondary sources rather than a primary HFMA publication; treat it as a convention, not a standard. And the widely circulated rework-cost figures attributed to MGMA could not be traced to a primary MGMA publication at all, so we do not use them.

    One more honesty check on the revenue side. In KLAS Arch Collaborative research on ambient documentation, over 80% of providers declined to see more patients after time was freed up, with only 18% wanting added volume. If your model converts saved minutes into incremental encounters, it is assuming a behavior most clinicians explicitly reject. Model the labor saving. Do not model the volume lift unless you have committed capacity and a physician who has agreed to it in writing.

    On the denial lever specifically, use primary data and name the denominator. KFF's March 2026 analysis of 2024 ACA Marketplace claims found 19% of in-network claims denied, with administrative reasons accounting for 25% of denials and missing prior authorization or referral another 9% — the two buckets a coding and documentation agent can actually influence. Separately, KFF's January 2026 Medicare Advantage analysis found 7.7% of prior-authorization determinations denied in full or part. That is a prior-authorization figure, not a claims figure; the two get conflated constantly in vendor material, and using the wrong one will get your model challenged.

    A defensible model has four lines: coder hours redeployed (not headcount removed, in year one), rework avoided on cleaner first-pass submissions, days in A/R improvement valued at your cost of capital, and the cost of the control program itself — audits, QA sampling, and monitoring are recurring line items, not one-time project costs.

    Frenchy Digital Cost Bands and Timelines

    What this costs to build properly, using the same scoping bands we apply across every healthcare automation engagement:

    EngagementRangeTimelineTypical scope
    Discovery + workflow audit$9k–$22k2–4 weeksBaseline coding audit, service-line viability assessment, vendor due diligence, written build plan
    Single-workflow agent$25k–$65k4–9 weeksOne service line, shadow mode then limited autonomy, confidence routing, audit logging
    Multi-workflow with EHR integration$65k–$160k9–16 weeksMultiple service lines, EHR read/write integration, QA sampling pipeline, distribution monitoring
    Multi-site / regulated build$160k–$400k+14–24 weeksHIPAA posture, human-in-the-loop controls, immutable audit logging, multi-entity governance

    Frenchy Digital 2026 cost bands for AI coding and healthcare automation engagements.

    Senior-led delivery is priced at $150–$225 per hour, with ongoing retainers from $2,500 to $9,500 per month covering model and code-set revalidation, QA sample review, monitoring, and incident response. Every engagement carries a 30-day post-launch warranty, and you receive a written scope and fixed-price phased proposal within 5 business days of the discovery call.

    Included at every tier: the baseline blinded audit design, confidence-threshold governance documentation, immutable audit logging, the stratified QA sampling pipeline, distribution-drift monitoring, and full source-code and IP ownership transferred to your practice at delivery. No vendor lock-in.

    Limitations and Failure Modes — Read This Before You Buy

    An honest accounting of what does not work, including the findings that cut against the category.

    • The evidence base is vendor-controlled: No independent benchmark of AI coding accuracy exists. Every figure in every deck is a claim. This is the single largest limitation and it does not have a technical fix.
    • An agent cannot fix bad documentation: It inherits it. If the note does not support the level of service, the correct behavior is a query, not an inference — and configuring an agent to infer the missing element is how a tool becomes a liability.
    • Adjacent AI results are mixed, not uniformly positive: In a randomized trial published in NEJM AI in December 2025 covering 238 physicians across 14 specialties, one ambient product reduced time-in-note by 9.5% versus control while another showed no significant reduction at all. Independent evidence in this family of tools is genuinely mixed.
    • Large deployments have underdelivered on productivity: The Permanente Medical Group's large-scale ambient deployment reported roughly 18 seconds saved per appointment versus non-users, and an Intermountain Health matched-cohort analysis found no statistically significant productivity gains — both reported in npj Digital Medicine in 2026.
    • Complex office E/M remains a human decision: MDM inference from narrative is the hardest part of ambulatory coding and the part most exposed to enforcement. Autonomy here is not defensible in 2026.
    • Modifier logic is brittle: Modifiers 22, 25 and 59, global periods, and staged procedures depend on context that frequently is not in the operative note at all.
    • Payer-specific rules move constantly: Coverage policies and edits change per payer, per quarter, and an agent trained on national conventions will be confidently wrong on regional plans.
    • Training and change management are underfunded: KLAS Arch Collaborative 2026 research found under 25% of AI-adopting clinicians said they received adequate training, and satisfaction plateaued past roughly four AI tools. The fifth tool is not free.
    • The control program is a permanent cost: Audits, QA sampling, monitoring, and annual revalidation recur forever. A vendor quote that omits them understates total cost of ownership.
    • Stacked AI compounds error: An ambient-generated note feeding a coding agent means documentation defects propagate into claims. Audit the note-to-encounter link, not just the code-to-note link.

    Sources for the mixed findings above: the December 2025 randomized trial is reported in NEJM AI; the Permanente and Intermountain results are reported in npj Digital Medicine; the training and tool-saturation findings come from KLAS Arch Collaborative. We cite them because they are the closest independent evidence available on AI in clinical documentation, and because the pattern they show — real benefit, unevenly distributed, smaller than marketed — is the prior you should carry into an AI coding evaluation.

    None of this argues against deployment. It argues for deploying narrowly, measuring your own results, and funding the control program as a first-class part of the build rather than an afterthought bolted on when someone asks a hard question.

    Red Flags When Buying an AI Coding Vendor

    Take this list into the demo. Every item has cost a real practice real money.

    Red flagWhy it matters
    Publishes an accuracy rate without case mix, denominator, or adjudication methodThe number is unfalsifiable and cannot support a compliance decision
    Will not let you run a blinded audit on your own claims before you signYou are being asked to accept their measurement instead of yours
    Describes CPT Appendix S as though it creates reimbursement for AIEither they do not understand the taxonomy or they are counting on you not to
    Sells 'autonomous coding' across all specialtiesAutonomy is defensible in a few structured service lines, not everywhere
    Confidence threshold is set by the vendor and not exposed to youYou own the compliance consequence of a dial you cannot see
    No calibration data behind the confidence scoreAn uncalibrated score is a ranking, not a probability
    Prices on a share of incremental coding-level liftThat contract structure pays the vendor to move your codes upward
    No BAA at the tier you are buying, or vague model-training termsYour PHI exposure is contractual before it is technical
    Cannot produce per-encounter reviewer dwell timeThen you cannot evidence that human review happened
    Promises reimbursement is coming for AI servicesCMS solicited comment in CY2026 and finalized nothing

    Frenchy Digital buyer's red-flag checklist for AI medical coding vendors, 2026.

    Ask one question and watch what happens: "Will you support a blinded audit of your output on our claims, scored by our coders, before we sign?" A vendor confident in their product says yes immediately. The answer tells you more than the entire demo.

    Frenchy Digital buyer's principle

    Working with Frenchy Digital

    Frenchy Digital is a senior-led, Black-owned Los Angeles agency that builds administrative AI systems for regulated businesses. In healthcare that means documentation and revenue-cycle workflows under human governance — never clinical decision-making, never a claim that our software is FDA-cleared, and never the phrase "HIPAA-certified," because no such certification exists.

    • Audit before automation: Every coding engagement opens with a blinded baseline audit of your current output and a distribution snapshot. We will tell you if the answer is that you should not automate this service line yet.
    • Senior engineers only: No junior staffing on regulated builds. The person designing your audit sampling is the person who has done it before.
    • Fixed-price phased scope: Written scope and fixed-price phased proposal within 5 business days of the discovery call. Phases are independently cancellable.
    • Compliance controls as deliverables: Audit logging, reviewer dwell-time capture, override-rate reporting, and the QA sampling pipeline ship as part of the build, not as a later add-on.
    • Your counsel decides the legal questions: We build the evidence pipeline and hand over the file. Sample sizing for confirmatory audits, extrapolation, and disclosure decisions belong to your compliance officer and healthcare counsel.
    • Full ownership transfer: Source code, prompts, configurations, infrastructure accounts, and IP transfer to your practice at delivery under clean work-for-hire terms. No vendor lock-in.
    • 30-day warranty and optional retainer: Every engagement carries a 30-day post-launch warranty; retainers from $2,500 to $9,500 per month cover revalidation, QA review, monitoring, and incident response.

    Book a discovery call at calendly.com/frenchydigital/discovery-call or call +1 (424) 272-5601. If you want to see the audit design before you talk to us about building anything, ask for it on the call — we will walk through it whether or not you hire us.

    Audit Before You Automate

    Book a free 60-minute discovery call with Frenchy Digital, a senior-led Black-owned LA agency. You leave with a written audit design for your own claims and a fixed-price phased proposal within 5 business days.

    Audit Before You Automate

    Book a free 60-minute discovery call with Frenchy Digital. You leave with a written audit design for your own claims and a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Unit 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    1. 1AMA — CPT Appendix S: Taxonomy for AI in Medical Services and Procedures
    2. 2HHS OIG — Industry Segment-Specific Compliance Program Guidance (Medicare Advantage ICPG, Feb 3, 2026)
    3. 3CMS — Medicare Physician Fee Schedule
    4. 4CMS — Interoperability and Prior Authorization Final Rule (CMS-0057-F)
    5. 5Federal Register — HIPAA Security Rule NPRM, 90 FR 898 (Jan 6, 2025)
    6. 6FDA — Clinical Decision Support Software, final guidance (re-issued Jan 29, 2026)
    7. 7HHS — HIPAA for Professionals
    8. 8ASTP/ONC — Decision Support Interventions certification criterion (HTI-1)
    9. 9KFF — Medicare Advantage prior authorization determinations, 2024 data (Jan 28, 2026)
    10. 10KFF — Claims denials and appeals in ACA Marketplace plans, 2024 (Mar 24, 2026)
    11. 11AMA — 2026 Physician AI Sentiment Report
    12. 12Premier Inc. — Claims adjudication costs providers $25.7B
    13. 13U.S. Bureau of Labor Statistics — Medical Records Specialists
    14. 14MGMA — Cost efficiency with medical group staffing
    15. 15NIST — AI Risk Management Framework
    16. 16NEJM AI — Randomized trial of ambient documentation tools (Dec 2025)
    17. 17npj Digital Medicine — Ambient documentation outcomes across large deployments (2026)
    18. 18Peer-reviewed review of 356 ambient-generated clinical notes — omissions and hallucinations
    19. 19JMIR Medical Informatics — Propagation of interpreter errors by ambient scribes (2026)
    20. 20KLAS Arch Collaborative — Ambient Speech Outcomes 2025
    21. 21AAPC — AMA releases CPT 2026
    22. 22Holland & Knight — CMS releases CY2026 Medicare Physician Fee Schedule final rule
    23. 23Ballard Spahr — DOJ's health care fraud takedown spotlights AI and data analytics
    Chris Machetto - CEO & Founder of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2019 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.