Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    Revenue Cycle
    August 9, 2026
    23 min read

    AI Agents in Medical Billingand the Revenue Cycle

    A workflow-by-workflow assessment for practice administrators: where AI agents genuinely earn their keep across eligibility, claim scrubbing, denial triage, appeal drafting, patient balances and posting — plus the ROI math, the OIG compliance controls, and the vendor red flags that matter in 2026.

    AI agents supporting medical billing and revenue cycle management in an independent physician practice, 2026
    19%
    In-network ACA Marketplace claims denied (2024 data)
    KFF, March 2026
    <1%
    Denied in-network claims that were ever appealed
    KFF, March 2026
    80.7%
    Medicare Advantage prior-auth appeals overturned
    KFF, January 2026
    $0
    CMS reimbursement for AI revenue cycle tools
    CY2026 Medicare PFS final rule

    Key Takeaways

    • The strongest primary denial data in 2026 is KFF's: 19% of in-network ACA Marketplace claims denied in 2024, and fewer than 1% of those denials appealed. In Medicare Advantage, only 11.5% of denials were appealed but 80.7% of appeals were overturned. The appeal gap, not the denial rate, is the arbitrage.
    • Denominators are not interchangeable. KFF's 19% is a claims figure; KFF's 7.7% Medicare Advantage figure is prior-authorization determinations. There is no 2026 primary provider-side denial benchmark — the circulating 11.8% initial denial rate traces only to vendor blogs and should not be cited.
    • AI revenue cycle tools have no CMS or CPT reimbursement. CMS solicited comment on paying for software algorithms in the CY2026 PFS but finalized nothing. CPT Appendix S is a taxonomy, not a payment schedule. Every dollar of return has to come from labor or recovered cash.
    • Agents are strongest where the work is deterministic and repetitive — eligibility checks, pre-submission edits, denial classification, appeal drafting, statement follow-up, remittance reconciliation. They are weakest where the work is judgment: medical necessity, complex coding, and anything that changes a diagnosis.
    • The OIG's February 2026 Medicare Advantage compliance guidance explicitly names AI-generated EHR prompts to add risk-adjusting diagnoses as a potentially abusive practice. An AI feature whose measurable effect is to raise risk scores now sits inside a named enforcement risk area.
    • Human-in-the-loop only counts if it leaves evidence: documented review time, override and rejection rates, immutable audit logs, and sampled quality assurance. Automation bias — accepting almost everything, almost instantly — is a compliance failure, not a productivity win.
    • Frenchy Digital cost bands: discovery and workflow audit $9k–$22k; single-workflow agent $25k–$65k; multi-workflow with EHR integration $65k–$160k; multi-site or regulated build $160k–$400k+. Senior-led $150–$225/hr, retainers $2,500–$9,500/month, 30-day post-launch warranty, full IP transfer.

    The Denial Economics You Are Actually Buying Into

    Before any conversation about agents, get the denominator right. The strongest primary denial data available in 2026 comes from KFF's March 2026 analysis of 2024 ACA Marketplace claims: 19% of in-network claims were denied, with a range from 3% to 36% across individual insurers. Out-of-network denials ran 37%, and 20% of all claims were denied.

    The reason breakdown matters more than the headline. KFF found 36% of denials were coded "other" or unspecified, 25% were administrative, 13% were for an excluded service, 9% were for a missing prior authorization or referral, and only 5% were medical necessity. Read that as a buy-list: the administrative and missing-authorization categories are exactly where deterministic pre-submission checks work. The medical-necessity category is not, and no agent will change that.

    Then the number that should reorganize your priorities: fewer than 1% of denied in-network claims were appealed — roughly 262,982 appeals against approximately 85 million denials. Where patients or providers did appeal internally, insurers upheld their original decision 66% of the time, meaning about 34% were overturned.

    The Medicare Advantage picture is different, and the difference is a denominator, not a discrepancy. KFF's January 2026 analysis found 52.8 million prior-authorization determinations in 2024 — 1.7 requests per enrollee — of which 4.1 million (7.7%) were denied in full or in part. Only 11.5% of those denials were appealed, but 80.7% of the appeals that were filed were overturned. Traditional Medicare, for contrast, ran 625,000 prior-authorization reviews (about 0.02 per enrollee) with a 22.9% denial rate.

    Do not conflate the two figures. KFF's 19% is a claims denial rate in the ACA Marketplace. KFF's 7.7% is a prior-authorization determination denial rate in Medicare Advantage. The "Medicare Advantage denies 15–18% of claims" figures circulating in vendor material are a third denominator again, sourced to secondary summaries rather than KFF. Any vendor who quotes one number for "the denial rate" without naming its denominator has not read the underlying study.

    There is a bigger gap in the evidence base, and it is worth stating plainly because most RCM vendor decks paper over it: there is no 2026 primary provider-side denial benchmark. The most recent public edition of the Optum Revenue Cycle Denials Index is the 2024 edition, which found 41% of denials originate in patient access. The widely circulated "11.8% initial denial rate in 2026" traces only to vendor blogs with no primary publication behind it. Do not put it in a board deck, and be suspicious of anyone who does.

    For the cost of the adjudication process itself, the best primary figure is Premier Inc., though it is 2023 data with no published update: claims adjudication cost providers $25.7 billion in 2023, up 23% year over year, of which $18 billion was potentially wasted on claims that should have paid on submission. Average adjudication cost per claim rose to $57.23 from $43.84. Premier also reported that roughly 70% of denials are ultimately overturned, typically after three review rounds of 45 to 60 days each.

    FigureWhat it actually measuresSourceCaveat you must carry with it
    19% of in-network claims denied (3%–36% across insurers)ACA Marketplace claims, 2024 plan-year dataKFF, published March 2026Marketplace-specific. Not a proxy for your commercial or Medicare mix.
    37% of out-of-network claims denied; 20% of all claimsACA Marketplace claims, 2024KFF, published March 2026Out-of-network volume differs enormously by specialty.
    Denial reasons: 36% other/unspecified, 25% administrative, 13% excluded service, 9% no prior auth or referral, 5% medical necessityACA Marketplace denials, 2024KFF, published March 2026The largest bucket is literally 'other' — payer reason codes are poor data.
    Under 1% of denied in-network claims appealed (~262,982 appeals vs ~85M denials)ACA Marketplace, 2024KFF, published March 2026The single most actionable number in this table.
    66% of internal appeals upheld by the insurer (~34% overturned)ACA Marketplace internal appeals, 2024KFF, published March 2026Appealing more is not the same as recovering more.
    7.7% of prior-auth determinations denied in full or part (4.1M of 52.8M)Medicare Advantage prior-authorization determinations, 2024KFF, published January 2026Prior-auth determinations, NOT claims. Different denominator.
    11.5% of MA denials appealed; 80.7% of those appeals overturnedMedicare Advantage, 2024KFF, published January 2026The most defensible statistic in this space.
    Traditional Medicare: 625,000 PA reviews (~0.02 per enrollee), 22.9% deniedOriginal Medicare, 2024KFF, published January 2026Volume per enrollee is ~85x lower than Medicare Advantage.
    $25.7B provider cost of claims adjudication; $18B potentially wasted; $57.23 average cost per claimAll-provider, 2023 dataPremier Inc.2023 data. No 2025 or 2026 update has been published.
    41% of denials originate in patient accessProvider-side denialsOptum Revenue Cycle Denials Index, 2024 editionMost recent public edition is 2024. No 2026 provider-side benchmark exists.

    Denial and adjudication benchmarks available to a practice administrator in 2026, with the denominator and the caveat attached to each.

    The appeal gap is the arbitrage. A denial category where fewer than one in a hundred denials is contested, and where the contests that do happen win most of the time, is not an AI problem — it is a capacity problem that AI happens to be good at.

    Frenchy Digital revenue cycle principle

    The Workflow Map: What the Agent Does, What the Human Keeps, and How It Fails

    Every serious revenue cycle automation decision comes down to one question repeated ten times: for this specific workflow, is the work deterministic enough that a machine can do it, and is the failure mode cheap enough that a machine getting it wrong is recoverable? The table below is the version of that analysis we run in every healthcare discovery engagement.

    Note the fourth column. Most vendor material describes what a tool does well and stops there. The realistic error mode is the column that determines whether you need a human review queue, how big it needs to be, and what the compliance exposure looks like when it fails at 2 a.m. on a Saturday.

    WorkflowWhat an agent does wellWhat still needs a humanRealistic error mode
    Eligibility & benefits verificationBatch checks 3–5 days ahead, re-checks day-of, parses 271 responses and payer portals, flags plan changes, copay/deductible/coordination-of-benefits deltasResolving a conflict between payer sources, calling the payer when the response is ambiguous, telling the patient what they oweConfidently reports stale or partial coverage data because the payer returned stale or partial coverage data. Garbage in, confident garbage out.
    Prior-auth requirement detection & packet assemblyDetects when a CPT/payer pair requires authorization, assembles the clinical packet, tracks the submission and the decision clockThe clinical justification, the peer-to-peer call, and any decision about proceeding without authorizationMisses a payer policy that changed mid-quarter, or assembles a packet against the wrong plan variant.
    Claim scrubbing / pre-submission editsRuns NCCI and payer-specific edits, catches modifier and units errors, missing referral or authorization numbers, demographic and COB mismatchesAny edit that requires reading the note to decide whether the service was actually performed as codedOver-suppresses: auto-fixes a legitimately unusual claim into a wrong one, and nobody sees it because it never hit a work queue.
    Coding suggestion (CPT / ICD-10)Suggests codes from documentation, flags unspecified codes, checks specificity and laterality, surfaces documentation gapsFinal code selection, level-of-service judgment, and every diagnosis that affects risk adjustmentUpcoding drift under automation bias — and the OIG has now named AI-generated diagnosis prompts as a risk area.
    Claim status & payer follow-upPolls 276/277 and portals on a schedule, ages the worklist, escalates silent claims before timely-filing windows closeThe call that actually moves a stuck claim, and any negotiation with a payer representativePortal changes break the integration silently; the queue looks clean because nothing is being fetched.
    Denial triage & routingNormalizes CARC/RARC codes into practice-meaningful categories, clusters denials by payer, provider and root cause, prioritizes by recoverable dollars and deadlineDeciding which denials to fight, which to write off, and which reveal a contract problem rather than a claim problemMis-clusters the 36% 'other/unspecified' bucket and produces a tidy dashboard describing nothing real.
    Appeal draftingDrafts the letter, pulls the relevant note excerpts, cites the payer's own policy and the specific denial reason, versions each attemptReviewing and signing the letter, filing it, owning the deadline, and escalating to external reviewA perfectly drafted appeal that nobody files. The deadline passes and a recoverable denial becomes a write-off.
    Patient balance follow-upSequenced statements and reminders across SMS, email and voice, payment-plan offers within policy, hardship routing, self-service payment linksAny conversation about financial hardship, charity care, or a disputed balance — and every escalationDunning a patient the practice already agreed to write off, or contacting through a channel the patient revoked.
    Payment posting & ERA reconciliationAuto-posts clean 835s, matches remits to claims, splits contractual adjustments from true underpayments, flags variance against the contracted rateUnmatched cash, credit balances, refunds, takebacks, and anything touching the general ledgerSilently posts an underpayment as a contractual adjustment, so the revenue leak never appears in any report.
    A/R analytics & reportingDays in A/R, A/R over 90 days, first-pass yield, denial rate by payer and reason, appeal rate and win rate, all refreshed without a manual pullInterpreting the trend, deciding what to renegotiate, and challenging a number that looks too goodDashboards that report the agent's activity rather than the practice's cash. Activity is not collection.

    Revenue cycle workflow map — agent capability, retained human judgment, and the failure mode to design against. Frenchy Digital, 2026.

    The pattern: agents are strong where the input is structured, the rule is written down somewhere, and the output is checkable. They are weak where the work is reading a clinical note and forming a judgment. Every workflow in the top half of that table is administrative. Every one in the bottom half touches clinical meaning — which is where the AMA's "augmented intelligence" framing stops being a slogan and starts being an operating constraint.

    Eligibility and Benefits Verification — Start Here

    Eligibility is the highest-confidence place to spend first, for three reasons that have nothing to do with how impressive the technology is. The input is structured (270/271 transactions and payer portals). The rule set is explicit. And the failure mode is caught downstream at scrubbing or at the front desk rather than at a payer audit.

    It also targets the right denial categories. In KFF's Marketplace data, administrative denials (25%) plus missing prior authorization or referral (9%) account for about a third of all denials — and both are substantially upstream problems. Optum's 2024 index reaching the conclusion that 41% of denials originate in patient access points the same direction from the provider side.

    • Batch verification 3–5 days out: Runs the full upcoming schedule against payer sources, so a coverage termination surfaces while there is still time to reschedule or collect.
    • Day-of re-verification: Coverage changes between booking and visit more often than most practices measure. A second check the morning of the appointment catches the delta.
    • Structured deltas, not raw dumps: The useful output is 'this patient's plan changed, deductible remaining moved from $400 to $1,900, secondary is now active' — not a 271 payload pasted into a note.
    • Prior-auth requirement detection: Cross-references the scheduled CPT against the payer's policy so the authorization request starts days earlier, not at check-in.
    • Escalation to a human on ambiguity: Any conflict between two payer sources, any partial response, and any coordination-of-benefits question goes to a person. That queue is a feature, not a failure.

    The honest limitation: an agent cannot fix bad payer data. If the payer returns stale or incomplete coverage information, the agent will report it confidently, because it has no independent way to know it is wrong. Design your workflow so a surprising result triggers verification rather than action.

    Why eligibility beats coding as a first project

    Coding automation is the workflow practices ask about first and the one we recommend last. Coding touches clinical meaning, sits directly inside a named federal enforcement risk area, and has no independent accuracy benchmark. Eligibility touches none of that, produces a measurable result inside one billing cycle, and builds the audit-log and override-tracking habits you will need before you go near a code.

    If eligibility automation does not measurably move your first-pass acceptance rate within 90 days, that is important information — and you learned it for the price of a single-workflow build rather than a multi-workflow platform commitment.

    Claim Scrubbing and Coding Support — Where the Line Sits

    Claim scrubbing and coding get bundled together in vendor demos, and they should not be. Scrubbing is a rules problem: NCCI edits, payer-specific edits, modifier and unit validation, demographic and coordination-of-benefits consistency, missing authorization numbers. Those are checkable against a source of truth, so an agent can run them at volume and be audited afterward.

    Coding is a judgment problem. It requires reading documentation and deciding what was actually done. That is a different risk class, and the evidence base is far weaker than the marketing suggests.

    Be precise about what is known here: every published accuracy figure for autonomous coding is vendor-produced. No independent benchmark exists. Vendors commonly report roughly 92–97% accuracy on structured, high-volume encounters and 82–90% on complex work, against human coders at 95–98% post-QA — but those are supplier numbers, not neutral measurements. Claims such as "denials cut from 18% to 3%" or "4–7x coder productivity" should be read as marketing. The specialties where autonomous coding is consistently reported as viable are radiology, pathology and the emergency department: high-volume, templated, structured. A general outpatient practice does not look like that.

    The standard 2026 architecture is the right one: auto-accept high-confidence output, route everything below the threshold to a human. What matters is that you own the threshold, that the override rate is visible, and that the queue is staffed. A confidence threshold nobody tunes and a queue nobody works is a rubber stamp with extra steps.

    On the coding standards themselves: CPT Appendix S, introduced in 2021, classifies AI-enabled services as assistive, augmentative or autonomous by the degree of physician override, and at the May 2026 CPT Editorial Panel meeting revisions were accepted replacing "machine" with "software output(s)". The AMA is developing a framework for Clinically Meaningful Algorithmic Analyses. Appendix S is a taxonomy, not a payment schedule. CPT 2026 carried 418 editorial changes effective January 1, 2026 — 288 new, 84 deleted, 46 revised — but the core office and outpatient evaluation and management codes 99202–99215 were unchanged; the notable restructuring was in remote patient and therapeutic monitoring.

    Denial Triage and Appeal Drafting — The Highest-Yield Build

    If the numbers in section one are right, this is where the money is. Under 1% of denied in-network Marketplace claims get appealed. In Medicare Advantage, 11.5% of denials get appealed and 80.7% of those appeals win. The binding constraint in almost every independent practice is not appeal-worthiness — it is that nobody has the hours to work the queue.

    That is a capacity problem, and capacity problems are what agents are genuinely good at. Split the work into two distinct halves and the risk profile becomes manageable.

    1. 1.Triage — a good agent job: Normalize CARC and RARC codes into categories that mean something to your practice, cluster by payer, provider, CPT and root cause, and rank the worklist by recoverable dollars against days remaining on the appeal window.
    2. 2.Root-cause attribution — agent-assisted, human-confirmed: The agent proposes 'this cluster is a registration problem at the Tuesday clinic' or 'this payer changed its policy in March.' A human confirms it, because acting on a wrong root cause is expensive.
    3. 3.Drafting — a good agent job: Pull the relevant note excerpts, cite the payer's own published policy and the specific denial reason, assemble the attachments, version each attempt so the second-level appeal knows what the first one argued.
    4. 4.Review and signature — human, always: Nothing goes to a payer that a person has not read and approved. This is not a courtesy; it is the control that makes the rest of the workflow defensible.
    5. 5.Filing and deadline ownership — human, with a named owner: Every denial with an appeal deadline gets an owner, a countdown, and an escalation path. A flag in a dashboard is not an owner.
    6. 6.Escalation to external review — human judgment: Deciding to take a case past internal appeal, and deciding which denials reveal a contract problem rather than a claim problem, is administrator work.
    The failure mode that actually costs money is not a badly drafted appeal. It is a perfectly drafted appeal that nobody files, sitting in a queue until the payer's timely-filing window closes. That converts a recoverable denial into a permanent write-off, and it is invisible in every activity dashboard because the agent did its job. Design for the deadline, not for the draft.

    Two regulatory tailwinds help here, and both are narrower than they sound. CMS-0057-F (89 FR 8758) has required, since January 1, 2026, that impacted payers decide expedited prior authorizations within 72 hours and standard requests within 7 calendar days — those timeframes exclude qualified health plan issuers on the federal exchanges and exclude drugs — and that every denial state a specific reason. Impacted payers published prior-auth metrics for the first time on March 31, 2026. A specific denial reason is materially better raw material for an agent than a generic code.

    The limits: the rule binds Medicare Advantage organizations, state Medicaid and CHIP fee-for-service and managed care, and qualified health plan issuers on the federal exchanges — not commercial or ERISA plans and not Original Medicare fee-for-service. And the FHIR-based Prior Authorization, Provider Access and Payer-to-Payer APIs do not go live until January 1, 2027. Until then, most payer integration is still portals and clearinghouse transactions. The only provider-facing obligation is an Electronic Prior Authorization attestation measure for MIPS eligible clinicians beginning with the CY2027 performance period and 2029 payment year.

    On the prior-auth burden itself, the primary source is the AMA's 2025 prior authorization survey (n=1,000, released May 2026): 13 hours per week of physician and staff time, 40 prior authorizations per physician per week, and 40% of physicians with staff working exclusively on prior authorization. Ninety-five percent report care delays, 79% report patients abandoning treatment, and 94% say prior authorization contributes to burnout. Insurers covering 257 million Americans pledged reform effective January 1, 2026, and AHIP self-reported an 11% reduction in prior authorizations — but that is payer self-reported, and only 33% of physicians in the AMA survey expect the pledge to make a meaningful difference.

    Patient Balances and Payment Posting — Real Money, Different Risks

    Patient balance follow-up and payment posting sit at opposite ends of the risk spectrum, which is why they are worth treating in the same section. One is low technical risk and high reputational risk. The other is the reverse.

    Patient balance follow-up is straightforward automation — sequenced statements and reminders across SMS, email and voice, payment-plan offers inside a policy you define, self-service payment links, and hardship routing. The technical failure modes are minor. The reputational ones are not: dunning a patient whose balance the practice already agreed to write off, or contacting through a channel the patient revoked, damages a relationship in a way no collection is worth. Every conversation about hardship, charity care or a disputed balance goes to a human immediately, and the agent's job is to route quickly rather than to persuade.

    Payment posting is the opposite. Auto-posting clean 835 remittances is nearly ideal agent work — structured input, checkable output, high volume. But the dangerous error is quiet: an agent that classifies an underpayment as a contractual adjustment makes a revenue leak disappear from every report you run. The control is a variance check against your contracted rate on every posted line, with anything outside tolerance routed to a person. Unmatched cash, credit balances, refunds and takebacks stay human — they touch the general ledger.

    • Contract-rate variance flags: Every posted line compared against the fee schedule you negotiated. Underpayments that post silently are the most common invisible leak in a practice's revenue cycle.
    • Credit balances stay human: Refunds, takebacks and credit balance resolution carry compliance obligations that do not belong in an automated queue.
    • Consent and channel state: The agent must respect revoked channels, do-not-contact flags, and any active financial-hardship status. Store that state where the agent reads it, not in a note field.
    • Write-off boundaries in policy, not in prompt: Thresholds for adjustment and write-off belong in a rules layer the practice controls and can audit — not embedded in model instructions.
    • Reconciliation reporting a human reads weekly: Posted versus expected, by payer, every week. If nobody reads it, the automation is running unsupervised regardless of what the org chart says.

    The ROI Math, Done Honestly

    Start with the fact most vendors skip: there is no reimbursement line for any of this. In the CY2026 Medicare Physician Fee Schedule final rule, CMS solicited comment on paying for software-as-a-service algorithms under the fee schedule, acknowledging that they are "not well accounted for" — and finalized no general AI payment pathway. CPT Appendix S is a taxonomy, not a payment schedule. The New Technology Add-on Payment program is inpatient-only and irrelevant to an outpatient practice's economics.

    So AI revenue cycle tools are cost-side plays. Every dollar of return has to come from labor you stop spending or cash you would otherwise have written off. There are five levers, and each one has a precondition that has to actually be true in your practice.

    LeverHow the return arisesWhat has to be trueHow you measure it
    Labor hours displacedFewer staff hours on eligibility calls, status checks, statement chasing and remit matchingThe freed hours are actually redeployed or removed — not absorbed as slackHours per week per workflow, before and after, measured for at least one full billing cycle
    Cash recovered that would have been written offMore denials worked, more appeals filed before deadline, more underpayments caught at postingAppeal capacity was previously the binding constraint, not appeal-worthinessAppeal rate, appeal win rate, and net recovered dollars per 1,000 claims
    Denials prevented at the front endEligibility and scrubbing catch administrative and missing-authorization errors pre-submissionYour denial mix actually skews administrative — check your own CARC data firstFirst-pass acceptance rate and denial rate by reason category
    Faster cash conversionShorter time from encounter to posted paymentThe bottleneck was internal throughput, not payer adjudication timeDays in A/R and percentage of A/R over 90 days
    Avoided outsourcing spendKeeping billing in-house rather than moving to a percentage-of-collections vendorYour in-house cost to collect stays below the outsourced percentage after toolingCost to collect as a percentage of net collections, fully loaded

    The five ways a revenue cycle agent can pay for itself, and the precondition attached to each.

    Now the arithmetic, with every input labeled by provenance. Treat the model below as a worked example against stated assumptions — not a benchmark, and not a promise.

    LineFigureProvenance and caveat
    Net collections (modeled practice)$4,000,000Stated modeling assumption — substitute your own
    Cost to collect at 3.5% of net collections$140,000 / yrHFMA best practice is commonly relayed as ≤2%, with >4% a red flag — figures relayed via secondary sources, not a confirmed HFMA primary publication
    Equivalent in-house billing staff2.0 FTE at $70,000 fully loadedFully-loaded biller range of $60k–$80k including benefits and software is a vendor-published figure; BLS median for medical records specialists was $50,250 in 2024
    Outsourced alternative at 6% of net collections$240,000 / yrOutsourced RCM commonly quoted at 4–10% of net collections (per claim $3–$12) — vendor-sourced but consistent across many providers
    Agent build, year one (single workflow)$25,000–$65,000Frenchy Digital band, 4–9 weeks
    Run rate, year one (retainer)$30,000–$114,000Frenchy Digital retainer $2,500–$9,500/month
    Year-one break-even requirement$55,000–$179,000 of freed labor or recovered cashBuild plus retainer. Roughly 0.8 to 2.6 fully-loaded FTE-equivalents
    Reimbursement offset available$0No CMS or CPT payment pathway exists for AI revenue cycle software

    Illustrative year-one model for a practice at $4M net collections. Substitute your own numbers — the structure is the point, not the totals.

    Read the break-even line honestly. In a two-biller practice, a $65,000 build plus a $9,500-per-month retainer does not pay back on displaced labor alone — you would need to remove roughly two and a half fully-loaded FTEs, which a two-FTE billing function cannot produce. At that scale the investment only clears if it moves recovered cash. That is a real conclusion, and it is why we recommend most practices under roughly $3M in net collections start with a discovery and workflow audit rather than a build.

    A few provenance notes, because these figures get laundered constantly. The fully-loaded in-house biller range of $60,000–$80,000 including benefits and software is a vendor-published figure, widely repeated and directionally reasonable, but not a government statistic. The primary anchor is the Bureau of Labor Statistics median for medical records specialists, which includes billing and coding roles: $50,250 in 2024. Model with the loaded figure and cite the BLS one.

    Similarly, the "cost to collect at or below 2% of net collections" best practice is universally attributed to HFMA but is relayed through secondary and vendor sources rather than a confirmed HFMA primary publication; the working convention is that 2–3% means room to improve and above 4% is a red flag, with most practices landing between 2% and 4% of net patient revenue. Outsourced RCM at 4–10% of net collections (most commonly 4–9%, with high-complexity specialties at 10–12%, or $3–$12 per claim) is vendor-sourced but consistent enough across providers to plan against. Days in A/R of 36 days for better performers and 47 days for the broader sample comes from the MGMA 2024 Cost and Revenue Survey as relayed by secondary sources. A clean claim rate above 95% is an industry convention, not a regulatory mandate, and A/R over 90 days is conventionally held under about 13–14%.

    One primary figure is worth keeping in front of a board: MGMA reports that support staff salaries and benefits run about 25% of total practice revenue — roughly half of overhead — with median investment to support one physician FTE at $343,128 in Q4 2025. MGMA also found automation is the top 2026 cost-cutting move at 36%, ahead of hiring freezes at 18%. That is the context in which this decision is being made across the market.

    Compliance: The OIG's February 2026 Guidance and the Automation-Bias Problem

    In February 2026 the HHS Office of Inspector General published its Medicare Advantage Industry Segment-Specific Compliance Program Guidance. It contains the single most citable federal statement to date on AI in the revenue cycle: the guidance explicitly names, as a potentially abusive practice, querying physicians via electronic medical record platforms — including prompts generated by artificial intelligence algorithms — to add risk-adjusting diagnoses.

    Querying physicians via electronic medical record platforms, including prompts generated by artificial intelligence algorithms, to add risk-adjusting diagnoses.

    HHS-OIG, Medicare Advantage Industry Compliance Program Guidance, February 2026

    Read that carefully, because it is narrower and more consequential than the headlines suggested. It does not prohibit AI prompts. What it does is place an AI feature whose measurable effect is to raise risk scores inside a named enforcement risk area — which means the question at audit stops being "did the diagnosis have clinical support?" and becomes "can you show how this suggestion was generated, who reviewed it, and how long they spent?"

    That connects directly to the second idea an administrator needs: automation bias is a compliance failure, not a productivity metric. The enforcement theory taking shape is that a "human in the loop" who accepts roughly 85% of AI suggestions at half a second per chart is not performing review in any meaningful sense. The defense is not a policy document asserting that review occurs. It is evidence.

    ControlWhat it means in practiceEvidence you must be able to produce
    Scope boundaryWritten statement that the agent performs administrative work only and does not determine medical necessity or practice medicineSigned scope document plus system prompts and tool allowlists that enforce it technically
    Business associate agreementExecuted BAA with the vendor, with flow-down to subprocessors including the model providerCountersigned BAA and a current subprocessor list
    Documented review timeActual seconds spent reviewing each AI-suggested claim, code or letter — not a checkboxPer-item review timestamps in the audit log
    Override and rejection ratesTracked per reviewer and per suggestion type, trended monthly, with a threshold that triggers investigationMonthly override-rate report retained with compliance records
    Immutable audit logsEvery submitted claim tied to the model version, the suggestion, the reviewer, and the timestampAppend-only log store with tested export
    Sampled quality assurancePeriodic random sample re-reviewed against source documentation by someone who did not do the original reviewWritten QA methodology, sample size, findings, and corrective actions
    Risk-adjustment guardrailNo AI feature may prompt a clinician to add a diagnosis in a way that changes risk score without independent clinical basisExplicit control mapped to the OIG's February 2026 Medicare Advantage guidance
    Payer-facing accuracyNothing goes to a payer that a human did not review and approveNamed signer on every appeal and every corrected claim

    The minimum control set for an AI-assisted revenue cycle in 2026, mapped to what an auditor will actually ask to see.

    The rest of the federal picture is more stable than the vendor discourse suggests. Any AI vendor whose product creates, receives, maintains or transmits electronic protected health information is a business associate, and a signed BAA is required — including flow-down to subprocessors, which in practice means the model provider behind the product. Note the language: no organization is "HIPAA-certified," because no such certification exists. The correct terms are a HIPAA-compliant posture, a signed BAA, and documented Security Rule safeguards.

    The HIPAA Security Rule NPRM (90 FR 898, RIN 0945-AA22, published January 6, 2025, comments closed March 7, 2025) is still proposed. It has moved to the Unified Agenda's long-term actions with a projected final action of July 2027 — the agency's own non-binding estimate. The 2003 Security Rule is what governs an AI agent touching ePHI today. As proposed, the update would eliminate the "addressable" versus "required" distinction, mandate multi-factor authentication and encryption of ePHI at rest and in transit, require an annually updated asset inventory and network map, an annual compliance audit, vulnerability scanning every six months, annual penetration testing, and 72-hour restoration of critical systems. Plan for it; do not represent it as a current obligation.

    Two more items belong on an administrator's radar. First, enforcement is getting better instrumented: the DOJ's 2026 National Health Care Fraud Takedown featured AI and data analytics prominently, and the Health Care Fraud Unit's Data Fusion Center, run with HHS-OIG and the FBI, is now standing infrastructure. Second, the professional-body position is unambiguous: the AMA's House of Delegates adopted policy in June 2026 holding that AI is an assistive tool rather than an autonomous decision-maker, requiring transparency and physician oversight wherever AI is used in patient care, and opposing autonomous or semiautonomous AI as a substitute for physician review in coverage determinations.

    The payer side is under the same constraint. CMS's WISeR Model — running January 1, 2026 through December 31, 2031 across six states with six technology participants — uses AI and machine learning to triage prior-authorization requests in Original Medicare for 13 service categories. Its operating guardrail is worth quoting verbatim in your own governance documents: "All recommendations for non-payment are determined by appropriately licensed clinicians." The algorithm triages. It does not issue the denial. Hold your own agents to the same standard.

    One practical due-diligence shortcut: if the AI feature is embedded in certified health IT, the ASTP/ONC Decision Support Interventions criterion (§170.315(b)(11)) requires developers to publish source attributes — 31 of them for predictive interventions, covering development details, fairness, external validation, quantitative performance and maintenance schedule. Those obligations bind certified EHR developers rather than practices, but the published transparency data is a ready-made vendor due-diligence artifact you did not have to ask for.

    What These Agents Still Cannot Do

    A limitations section is not a disclaimer. It is the part of a vendor evaluation that saves you the most money, because every item below is something a well-run pilot will surface and a well-run sales process will not.

    • Prove their own accuracy: No independent benchmark exists for autonomous coding or denial prediction. Every published figure is vendor-produced. The only honest test available to you is a shadow-mode run against 90 days of your own adjudicated claims.
    • Generalize outside structured specialties: Autonomous coding is consistently reported as viable in radiology, pathology and the emergency department because the work is high-volume and templated. A multi-specialty outpatient practice does not share those properties.
    • Fix bad payer data: An eligibility agent reports what the payer returns. If the 271 response is stale or partial, the agent will state it with full confidence, because it has no independent source to check against.
    • Survive payer portal changes: Most payer integration in 2026 is still portals and clearinghouse transactions. CMS-0057-F's FHIR APIs do not arrive until January 1, 2027, and only for impacted payer types — not commercial or ERISA plans, not Original Medicare fee-for-service.
    • Own a deadline: An agent can surface, prioritize and draft. It cannot be accountable. A denial queue with no named human owner and no escalation path will lose appeals to timely-filing windows regardless of how good the drafting is.
    • Turn appeal volume into recovery: KFF found insurers upheld 66% of internal appeals in the ACA Marketplace. Filing more appeals is not the same as recovering more money — you need win rate by payer and by denial reason, not appeal count.
    • Resolve medical-necessity denials: Only 5% of Marketplace denials were medical necessity, and that 5% requires clinical judgment, documentation the agent did not create, and frequently a peer-to-peer conversation.
    • Replace review: Automation bias is the failure mode that turns a compliant workflow into an enforcement exposure. If your override rate is near zero and your review time is near zero, you have automated the rubber stamp, not the work.
    • Practice medicine: These are administrative and documentation workflows operating under human review. An agent does not determine medical necessity, does not make clinical decisions, and is not a licensed professional. That boundary belongs in your scope document, your system prompts, and your tool allowlists.
    The evidence gap, stated plainly: there is no 2026 primary provider-side benchmark for denial rates, no independent accuracy benchmark for AI coding, no primary MGMA publication behind the widely quoted "50–65% of denials are never reworked" and "$25–$118 per rework" figures, and no CMS reimbursement pathway for any of these tools. A vendor who presents all four as settled facts is telling you something useful about how they will handle your data.

    Red Flags When Evaluating RCM AI Vendors

    Revenue cycle AI is a crowded 2026 market with unusually weak evidence standards, which makes vendor selection more consequential than platform selection. These are the questions we put in front of every healthcare client during discovery — including the ones who ultimately hire someone else.

    Red flagWhy it matters
    Pure contingency fee on recovered dollars, uncappedIt pays the vendor for appeal volume. The cheapest denial is the one that never happens — pure contingency has no financial reason to improve your first-pass yield. Pair any contingency with a contracted first-pass acceptance rate.
    No written answer on who owns the denial and appeal dataYour denial corpus is the most valuable asset the engagement produces. Get ownership, export rights, and a training carve-out in writing before the first claim is processed.
    No named owner or SLA for appeal deadlinesAsk directly: who files, by when, and what happens if the window closes? If the answer is 'the system flags it,' a flag is not an owner and a missed deadline is a permanent write-off.
    Accuracy claims with no neutral basisThere is no independent benchmark for autonomous coding or denial prediction. Ask for the denominator, the specialty mix, the date range, and who computed the number. If the answer is 'our internal data,' treat it as marketing.
    Claims new AI CPT codes or payer reimbursementThere is no general AI CPT payment category. CPT Appendix S is a taxonomy, and CMS finalized no AI payment pathway in the CY2026 PFS. A vendor who gets this wrong is wrong about your economics.
    Says 'HIPAA-certified'No such certification exists. Correct language is a HIPAA-compliant posture, a signed BAA, and documented Security Rule safeguards. This one sentence tells you how carefully the vendor reads regulation.
    Auto-submit on by default with no confidence thresholdAutonomous submission without a low-confidence human queue is how a coding drift becomes a repayment obligation. Confidence thresholds and routing rules should be yours to set, not theirs.
    Cannot export an override-rate report or an audit logIf the platform cannot evidence who reviewed what and when, you cannot evidence human review — which is exactly what an audit will ask for.
    Refuses a shadow-mode pilot against your own historical claimsThe only honest accuracy test is the vendor's tool run against 90 days of your own adjudicated claims, scored against what actually happened. A refusal is an answer.
    No first-pass yield, denial rate or days-in-A/R commitments in the contractIf the metrics that define success are not in the agreement, the engagement will be measured on activity dashboards instead of cash.
    Integration only through the vendor's own middleware, with no worklist exportThat is lock-in. You should be able to terminate on a Friday and keep operating your billing on a Monday.
    Markets denial-reduction percentages as a guaranteeTreat headline claims such as 'denials cut from 18% to 3%' or '4–7x coder productivity' as marketing. No neutral source substantiates them.

    The Frenchy Digital red-flag checklist for revenue cycle AI vendor evaluation, 2026.

    The four contract questions that separate a real vendor from a demo

    1. What is the fee aligned to? If the vendor is paid on recovered dollars, they profit from denials existing. Pair contingency with a contracted first-pass acceptance rate measured against your pre-engagement baseline. 2. Who owns the denial data? Ownership, export rights, and an explicit carve-out from shared model training — in writing, before the first claim.

    3. Who owns appeal deadlines? A named role, a service level, and a defined consequence when a window closes. 4. What is the neutral basis for the accuracy claim? Ask for the denominator, the specialty mix, the date range, and who computed it — then ask for a shadow-mode pilot against your own historical claims. The willingness to run that pilot is the answer.

    If the metrics that define success are not in the contract, the engagement will be measured on activity dashboards. Activity is not collection.

    Frenchy Digital buyer's principle

    Frenchy Digital Cost Bands for Revenue Cycle Agents in 2026

    Frenchy Digital is a senior-led, Black-owned Los Angeles agency that builds administrative automation for healthcare practices. Here are the bands we scope against, with the assumptions behind each tier.

    EngagementRangeTimelineTypical scope
    Discovery + workflow audit$9k–$22k2–4 weeksBaseline first-pass yield, denial mix by payer and reason, A/R aging, staff time study, integration feasibility, written build recommendation
    Single-workflow agent$25k–$65k4–9 weeksOne workflow end to end — eligibility, denial triage, appeal drafting or statement follow-up — with human review queue and audit logging
    Multi-workflow practice automation with EHR integration$65k–$160k9–16 weeksThree or more connected workflows, EHR/PM integration, unified worklist, override tracking, reporting layer
    Multi-site / regulated build, HIPAA posture + HITL + audit logging$160k–$400k+14–24 weeksMulti-entity rollout, formal HIPAA posture documentation, human-in-the-loop controls, immutable audit logs, QA sampling program, staff training

    Frenchy Digital 2026 cost bands for healthcare revenue cycle automation.

    Senior-led delivery is priced at $150–$225 per hour, with ongoing retainers from $2,500 to $9,500 per month depending on the number of live workflows and the depth of the compliance program. Every engagement carries a 30-day post-launch warranty, and you receive a written scope and fixed-price phased proposal within 5 business days of the discovery call.

    Included at every tier: a documented human-review design for every automated workflow, audit logging with tested export, override-rate reporting, a written scope statement that limits the agent to administrative work, and full source-code and IP ownership transferred to your practice at delivery. No vendor lock-in. We do not claim FDA clearance, and we do not claim HIPAA certification — because neither exists for this class of software.

    A note on sequencing spend: the discovery and workflow audit exists precisely so you do not buy a build you do not need. In a meaningful minority of engagements, the audit's recommendation is to fix a payer contract, a registration script or a scheduling policy first — and to revisit automation once the underlying process produces data an agent can act on.

    Where to Spend First: A 12-Month Sequence

    The order matters more than the tooling. This sequence puts the most deterministic, lowest-clinical-risk workflows first, so the practice builds the audit-log and override-tracking discipline on low-stakes work before applying it anywhere near a diagnosis code.

    WindowFocusWhy here
    Months 0–1Baseline measurementFirst-pass acceptance rate, denial rate by payer and CARC, appeal rate, appeal win rate, days in A/R, A/R over 90 days, cost to collect, staff hours by task
    Months 1–3Eligibility verification agentMost deterministic workflow, zero clinical exposure, hits the administrative and missing-authorization denial categories directly
    Months 3–5Denial triage and routingNormalize CARC/RARC into practice-meaningful categories, prioritize by recoverable dollars and deadline, build the worklist a human actually works
    Months 5–8Appeal drafting with human filingTargets the appeal gap. Agent drafts and cites; a named person reviews, signs, files, and owns the deadline clock
    Months 8–10Patient balance follow-upSequenced multichannel outreach with hardship routing and human escalation. Revenue impact is real; reputational risk is the constraint
    Months 10–12Posting and underpayment detectionAuto-post clean remits, flag variance against contracted rates, route unmatched cash and credit balances to a human
    OngoingCoding support — last, not firstSuggestion-only, confidence-thresholded, with override tracking and sampled QA. Never autonomous in a general outpatient practice

    A 12-month revenue cycle automation sequence for an independent practice — Frenchy Digital, 2026.

    Month zero is not optional. Without a documented baseline, you cannot distinguish an agent that improved your collections from a quarter with a favorable payer mix, and you will have no defensible answer when your board asks what the $65,000 bought. Capture first-pass acceptance rate, denial rate by payer and CARC, appeal rate, appeal win rate, days in A/R, A/R over 90 days, cost to collect, and staff hours by task — before anything is switched on.

    The four metrics to put in front of your board every quarter

    First-pass acceptance rate tells you whether the front end is working. Net recovered dollars per 1,000 claims tells you whether the denial workflow is working — and unlike appeal count, it cannot be gamed by filing weak appeals.

    Cost to collect as a percentage of net collections, fully loaded to include the tooling and the retainer, tells you whether the investment is actually cheaper than the labor it replaced. Override rate on AI suggestions tells you whether human review is real. A rate trending toward zero is a warning, not a win.

    For context on adoption: the AMA's 2026 Physician Survey on Augmented Intelligence (n=1,692, fielded January 15 to February 2, 2026) found 81% of physicians now use AI in practice, up from 66% in 2024 and 38% in 2023, with roughly 70% seeing AI's value in automating the drivers of burnout — documentation, chart review and prior authorization. The question in 2026 is no longer whether to adopt. It is which workflow, in what order, and under what controls.

    The Administrator's Summary

    If you take four things from this piece: the appeal gap is the largest single opportunity in your revenue cycle and the numbers supporting that are primary and recent. There is no reimbursement for any of these tools, so the entire business case is labor and recovered cash. The OIG has named AI-generated diagnosis prompts as a risk area, which makes documented human review a control rather than a courtesy. And the evidence base for vendor accuracy claims is thin enough that a shadow-mode pilot against your own claims is the only test worth trusting.

    Start with eligibility. Measure before you build. Keep a named human on every deadline. And treat any vendor who cannot name the denominator behind their headline statistic as the answer to a question you were about to spend six figures on.

    Decide Where to Automate Your Revenue Cycle First

    Book a free discovery call with Frenchy Digital — a senior-led, Black-owned Los Angeles agency. You leave with a baseline measurement plan and a fixed-price phased proposal within 5 business days.

    Decide Where to Automate Your Revenue Cycle First

    Book a free discovery call with Frenchy Digital. You leave with a baseline measurement plan and a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Unit 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    Chris Machetto - CEO & Founder of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2019 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.