The question behind the question
Okay, and what does it cost me every month after that?
That's the second question on almost every discovery call I take about an agent. The first one is always the build price. The second one comes a few minutes later, usually after the person has done some quiet math about whether they're buying a piece of software or signing up for a utility bill.
It's the right question. It's also the one most pricing pages and most "how much does an AI agent cost" articles answer badly, because they give you a build range with no source and then wave at "ongoing costs" without a single number.
So here's the whole bill. On September 28, 2026 I went through the list prices for the models, voice platforms, databases, hosting and tracing tools an agent actually runs on, wrote down what each page says, and did the arithmetic for two realistic setups: a support chat agent handling 3,000 conversations a month, and a phone receptionist handling 1,000 minutes a month.
The short answer, which the rest of this article earns: the build is the biggest single check, the tokens are the smallest recurring line, and the maintenance is the line nobody budgets for.
Disclosure.Frenchy Digital builds AI agents, and I quote our own prices in this article. Those are ours and I say so each time. Every third-party price is the vendor's own list price on the date checked, attributed, and linked in the sources at the bottom. Negotiated prices will differ.
The two wrong models
There are two ways people get AI agent cost wrong, and they're opposites.
The first wrong model is the website model. You pay someone to build the thing, it goes live, and after that the cost is basically hosting. That's how a brochure site works, and it's how a lot of people budget an agent. Then the first model invoice arrives and they feel ambushed.
The second wrong model is the meter model. This is the more sophisticated buyer, the one who has read about tokens. They assume the model bill is the big ongoing cost, so they spend weeks comparing per-million-token prices and choosing between two models that differ by a few dollars a month at their volume.
Both are wrong in a way that costs real money.
The real model looks more like owning a car than buying a website. There's a purchase price (the build), there's fuel (tokens and minutes), there's insurance and registration (hosting, tracing, the database), and there's servicing (evaluation, fixes, and migrating when the model underneath you is retired). For most small and mid-sized agents I've priced, the fuel is the cheapest of those four. The servicing is the one that quietly outgrows everything else.
Why does this matter? Because the two wrong models lead to opposite mistakes. The website buyer under-budgets and abandons a working agent the first time it needs attention. The meter buyer over-optimizes a line that barely moves the total and skips the evaluation work that actually keeps the agent correct. Therefore the useful budget has four lines, not one, and I'll build each of them below.
What the build costs
The build is a one-time cost, and it ranges widely because "an agent" covers everything from a single booking flow to a multi-system platform.I'll start with our own numbers because they're the ones I can stand behind, then show you the one market figure I found with a stated method.
Frenchy Digital prices bespoke agent work in four bands. These are the same figures we publish everywhere, and they don't change from article to article:
| Engagement | Price | Timeline |
|---|---|---|
| Discovery and workflow audit | $9k-$22k | 2-4 weeks |
| Single-workflow agent | $28k-$70k | 4-9 weeks |
| Multi-workflow platform with system integration | $70k-$180k | 9-16 weeks |
| Enterprise, multi-site or regulated build | $180k-$420k+ | 14-24 weeks |
Work is senior-led at $150 to $225 an hour, every engagement comes with a 30-day post-launch warranty, we send a fixed-price phased proposal within 5 business days of discovery, and full source code and IP ownership transfer to the client. The full scope of what we build sits on the AI agent creation service page.
Then there's the packaged route. We sell eight pre-scoped agents, listed on the AI agents hub, at a fixed $5,000 for the agent plus a one-time $5,000 setup fee. That's $10,000 at checkout, taken as one Stripe payment, not a subscription. The setup fee covers discovery, the build, connecting the agent to your calendar and phone or web channel, the guardrail work, and tuning after launch. The pages state that most packaged builds run 2 to 4 weeks, and the custom one 2 to 6 weeks depending on how many systems are involved.
To be clear, $10,000 sits well below the $28k-$70k single-workflow band, and a careful reader should ask why. The answer is scope. A packaged agent does a recognized job on systems you already use, like an existing calendar, inbox or CRM, so nothing gets migrated and nothing gets invented. The bespoke bands price the work beyond that: several systems to integrate, custom write paths into your own software, regulated data, or several agents coordinating. If your job fits a package, you shouldn't pay for bespoke. If it doesn't, a package will frustrate you.
One thing the package pages don't do is publish a monthly fee, or say who pays for model tokens, telephony and hosting after launch. Those are scoped on the discovery call. I'm not going to invent a number for them here, which is exactly why the rest of this article prices those lines from the vendors' own pages instead.
What about the market? The one dataset I found with a stated method is Clutch's. Its artificial intelligence pricing guide, built from verified client reviews and updated in September 2026, reports an average AI project cost of about $120,595, a most common project band of $10,000 to $49,999, and a typical timeline of roughly 10 months. Clutch also says most listed AI firms charge $25 to $49 an hour. That's the downside of the dataset: it skews toward offshore providers, so it's a market average, not a US senior-rate benchmark.
If you're still deciding who should build it, I ranked ten third-party firms on attributes you can check in the top AI agent development companies for 2026, and the pricing models those firms use (retainer, project, productized) are pulled apart in how AI automation agencies price their work. This article stays on the agent itself: what it costs to build and keep alive.
Model tokens, line by line
Tokens are the fuel, and at small-business volume they're cheaper than almost anyone expects.Every figure below is a list price checked September 28, 2026, per million tokens, on each provider's own pricing page. OpenAI's rows are its Standard, short-context rates; long context costs roughly twice as much.
To make the prices comparable I fixed one workload. Scenario 1 is a support chat agent handling 3,000 conversations a month, with each conversation using about 8,000 input tokens (system prompt, policy text, retrieved documents, the customer's messages) and about 1,000 output tokens (the agent's replies). That's an assumption, not a measurement. Your conversations could be half that or triple it, and the table scales linearly.
The multiplication, so you can check it: 3,000 conversations times 8,000 tokens is 24 million input tokens a month. 3,000 times 1,000 is 3 million output tokens. Every row below is 24 times the input price plus 3 times the output price.
| Model | Input per 1M tokens | Output per 1M tokens | Cached input | Monthly cost, scenario 1 | Per conversation |
|---|---|---|---|---|---|
| Claude Fable 5.1 (Anthropic) | $10 | $50 | $0.25 | $390 | $0.13 |
| Claude Opus 5.5 (Anthropic) | $4 | $20 | $0.20 | $156 | $0.052 |
| Claude Sonnet 5.5 (Anthropic) | $2 | $10 | $0.20 | $78 | $0.026 |
| Claude Haiku 4.5 (Anthropic) | $1 | $5 | $0.10 | $39 | $0.013 |
| gpt-6-astra (OpenAI) | $10 | $50 | $1 | $390 | $0.13 |
| gpt-6-sol (OpenAI) | $2 | $10 | $0.20 | $78 | $0.026 |
| gpt-6-luna (OpenAI) | $0.10 | $0.50 | $0.01 | $3.90 | $0.0013 |
| Gemini 3.1 Pro Preview (Google) | $2 | $12 | Not used here | $84 | $0.028 |
| Gemini 3.8 Flash (Google), to Dec 31, 2026 | $0.75 | $3.75 | $0.075 plus storage | $29.25 | $0.00975 |
| Gemini 3.8 Flash (Google), from 2027 | $1.50 | $7.50 | Not re-checked | $58.50 | $0.0195 |
Take the Claude Sonnet 5.5 row as the worked one. Input is 24 times $2, which is $48. Output is 3 times $10, which is $30. Total, $78 a month. Divide by 3,000 conversations and you get $0.026, about two and a half cents a conversation.
Now look at the spread. The flagships, Claude Fable 5.1 and OpenAI's gpt-6-astra, both list at $10 input and $50 output, which is exactly 5x Sonnet 5.5 on this workload: $390 a month. At the other end, gpt-6-luna does the same token volume for $3.90. That's a 100x range across models that are all, on paper, able to hold a support conversation.
So is the answer to use the cheapest model? No, and this is where the meter model breaks. The difference between Sonnet 5.5 and Haiku 4.5 on this workload is $39 a month. If the smaller model mishandles even a handful of refund questions that a person then has to clean up, you've spent the $39 in the first hour of cleanup. Pick the model on how it performs on your evaluation set, then look at the price. Not the other way round.
Three pricing details that change the math
Caching. Anthropic's pricing page prices a cache read at 0.1x the base input rate for most models, a five minute cache write at 1.25x and a one hour write at 2x. If 6,000 of the 8,000 input tokens in each conversation are the same system prompt and policy text, they can be cache hits. Then uncached input is 6 million tokens at $2, or $12; cached input is 18 million at $0.20, or $3.60; output is still $30. Total about $45.60 a month before cache write charges, down from $78. That's roughly a 42% cut on the token line.
Batch. Anthropic's Batch API is 50% off input and output, and OpenAI's pricing page lists Batch and Flex at exactly half the Standard rate (gpt-6-astra drops to $5 and $25). The catch: batch work isn't real time, so it helps with nightly summaries, ticket tagging and evaluation reruns, not with the customer waiting on a reply.
Price changes you can see coming. Google's Gemini pricing page lists Gemini 3.8 Flash at $0.75 input and $3.75 output through December 31, 2026, then $1.50 and $7.50. Same agent, same traffic, and the token line doubles from $29.25 to $58.50 on January 1. Gemini 3.1 Pro is listed at $2 and $12 for prompts up to 200k tokens, and it's a preview model, which matters for the maintenance section below.
Two smaller multipliers are worth knowing about. Anthropic charges 1.1x for US-only inference if you need data processed in the US, and its page notes that the tokenizer used by Claude 4.7 and later models produces roughly 30% more tokens for the same text. That second one is sneaky: migrate to a newer model at the same per-token price and your bill can still rise by about a third, because the same words count as more tokens.
Anthropic's own page includes an example worth borrowing as a sanity check: about 3,700 tokens a conversation on Haiku 4.5 comes to about $37 per 10,000 tickets. That's under half a cent a ticket. My scenario uses more than twice the tokens per conversation because retrieval stuffs documents into the prompt, which is typical for a support agent that has to quote your policies.
Voice and telephony minutes
A voice agent is priced by the minute, and the minute is a bundle of four things: the phone line, speech to text, the model, and text to speech.Some vendors sell the bundle, some sell the parts. You need to know which you're looking at before you compare.
Scenario 2 is a phone receptionist handling 1,000 minutes a month. For a small practice or a trades business that's a realistic month. For a busy one it's a quiet week.
Start with the line itself. Twilio's US voice pricing lists inbound local calls at $0.0085 a minute, outbound local at $0.0140, a local number at $1.15 a month and a toll-free number at $2.15, with recording at $0.0025 a minute and transcription at $0.05. So 1,000 inbound minutes is $8.50, plus $1.15 for the number: $9.65. That's the phone company part only. Nothing is listening yet.
Now the voice platforms, all attributed to their own pricing pages:
- Retell: Retell's pricing page itemizes infrastructure at $0.055 a minute, text to speech at $0.015, a language model from $0.0016 (GPT 5 nano) to $0.32 (GPT 6 Astra, standard tier), and Retell telephony at $0.015. With GPT 5.6 Terra, which the page labels Recommended, at $0.064, those add up to about $0.149 a minute. That total is my sum, not Retell's: the page's own calculator default shows $0.11 a minute, with the model at $0.04 and telephony at $0.00, so a phone line has to be added to that figure. A Retell number is $2 a month.
- Deepgram: Deepgram's pricing page lists its Voice Agent API at $0.050 to $0.163 a minute, streaming speech to text from $0.0048 a minute on a limited-time promotion (regular $0.0077), and a $200 free credit.
- ElevenLabs: ElevenLabs lists its agents product at $0.08 a minute, with a burst rate of $0.16. I could not confirm from the page whether the language model is included in the $0.08, so treat it as possibly excluding the model.
| Stack | Published rate | 1,000 minutes | Plus Twilio line | Monthly total |
|---|---|---|---|---|
| Twilio alone (phone line only, no speech or model) | $0.0085/min inbound, $1.15 number | $8.50 | Included | $9.65 |
| Deepgram Voice Agent API, low end | $0.050/min | $50 | $9.65 | $59.65 |
| ElevenLabs agents | $0.08/min (model inclusion unconfirmed) | $80 | $9.65 | $89.65 |
| Retell, page calculator default (telephony shown as $0.00) | $0.11/min | $110 | $9.65 | $119.65 |
| Retell, our sum with GPT 5.6 Terra and Retell telephony | $0.149/min | $149 | Included, plus $2 number | $151 |
| Deepgram Voice Agent API, high end | $0.163/min | $163 | $9.65 | $172.65 |
So the voice line for 1,000 minutes runs from about $60 to about $173 a month on these stacks. Divide the ends: $172.65 over $59.65 is about 2.9x. The whole spread is smaller than a single day of a person's wages, which is why the voice line is rarely what decides whether a receptionist agent pays.
Here's the downside, in the same breath. The voice line scales with every minute, including the minutes you didn't want: robocalls, wrong numbers, the caller who stays on for twenty minutes. Text agents cost per conversation; voice agents cost per minute of patience. Put a maximum call length and a transfer rule in the design before launch, not after the first bill.
If you want to set the voice line against a person at the front desk, that comparison is worked in the AI receptionist versus human receptionist cost breakdown. Our own AI receptionist package is the same $5,000 plus $5,000 as the others; its running costs are not published on the page and are scoped on the call.
Retrieval, hosting, observability
The plumbing is three small subscriptions, and at small scale they cost more than the tokens. That surprises people, so here are the list prices.
Retrieval. A support agent that quotes your policies needs somewhere to search them. Supabase's pricing page lists Pro from $25 a month including $10 of compute credits, enough for one Micro instance, and pgvector runs inside the same Postgres database you'd use for everything else. Pinecone's pricing page lists Builder at $20 a month, Standard at a $50 monthly minimum with pay as you go above it, and Enterprise at a $500 minimum. For a few thousand documents, the Postgres route is usually enough. A dedicated vector database earns its keep at a scale most small businesses never reach.
Hosting. Vercel's pricing page lists Pro at $20 a month with $20 of included usage credit, and $20 per developer seat. For an agent that answers a web chat, that's the whole hosting line at low volume.
Observability. This is the one people skip, and it's the one I'd cut last. It's how you read what the agent said to your customers yesterday. Langfuse's pricing page lists a free Hobby tier with 50k units, Core at $29 a month with 100k units and $8 per additional 100k, Pro at $199, Enterprise at $2,499, and a $300 Teams add-on. LangChain's LangSmith lists Plus at $39 per seat a month with 10k base traces included, then pay as you go. I didn't capture LangSmith's exact overage rate, so I won't print one.
Add up the entry tiers for scenario 1: Supabase Pro $25, Vercel Pro $20, Langfuse Core $29. That's $74 a month, which is almost exactly the $78 token bill. In other words, at 3,000 conversations a month, the plumbing roughly doubles the fuel.
That ratio flips at scale. Tokens grow with every conversation; the plumbing grows in steps. At 30,000 conversations the Sonnet token line is ten times bigger, $780, while the plumbing might move up one tier. It's a useful way to think about where your bill is going: below a few thousand conversations, you're paying mostly for infrastructure; above a few tens of thousands, mostly for the model.
What does a real multi-agent build's plumbing look like? The one agent build we've published, Beyond Points, runs a Claude-based multi-agent system over MCP with Gemini Flash chat, a Puppeteer browser agent in Docker on a VPS that completes card-to-partner transfers behind an explicit confirmation gate, and Supabase with 71 edge functions and about 70 migrations. Those are build counts, not outcome claims, but they make the point: every one of those pieces is a line on someone's monthly bill and a thing someone has to maintain.
Two worked scenarios
Put the lines together and a typical small agent costs somewhere between $150 and $250 a month to run, before anyone's time. Here are both scenarios assembled, so you can see which line dominates.
Scenario 1: support chat agent, 3,000 conversations a month
Consider a small e-commerce or services business whose support inbox handles about 3,000 conversations a month. The agent answers order, policy and booking questions from a knowledge base and hands anything else to a person.
- Tokens, Claude Sonnet 5.5 at list: $78 a month ($0.026 a conversation).
- Retrieval, Supabase Pro: $25.
- Hosting, Vercel Pro: $20.
- Observability, Langfuse Core: $29.
- Total: $78 + $25 + $20 + $29 = $152 a month, or about $0.051 a conversation.
Swap the model and the total moves less than you'd think. On Haiku 4.5 it's $39 + $74 = $113. On Gemini 3.8 Flash before the price change it's $29.25 + $74 = $103.25. On a flagship at $390 it's $464. The cheapest and the most expensive realistic stacks differ by about 4.5x, and the middle options sit within about $50 of each other.
Scenario 2: voice receptionist, 1,000 minutes a month
Suppose a two-location clinic or trades business routes 1,000 minutes a month of inbound calls to an agent that answers questions, books appointments into an existing calendar and transfers anything clinical, urgent or angry to a person.
- Voice stack, Retell with GPT 5.6 Terra and Retell telephony (our sum): $149, plus a $2 number, $151.
- Or Deepgram Voice Agent API plus a Twilio line: $59.65 to $172.65.
- Or ElevenLabs agents plus a Twilio line: $89.65, possibly plus model charges.
- Hosting and observability can often reuse the tiers from scenario 1 if you run both; alone, add roughly $49 for Vercel Pro and Langfuse Core.
So a receptionist runs about $110 to $220 a month all in on list prices. Per minute, that's roughly 11 to 22 cents. Per call, if the average call lasts three minutes, it's 33 to 66 cents, which is my assumption about call length, not a measured figure.
Notice what's missing from both cards: anybody's time. Neither scenario includes the hours spent reading transcripts, fixing a prompt when a policy changes, or rerunning tests when the model is swapped. That's deliberate, because it's the line I want to look at on its own, after the platform comparison.
Per-resolution platforms compared
The alternative to building is renting an agent priced per resolution or per conversation, and the arithmetic favors renting at low volume and building at high volume.Here's where the line crosses for scenario 1.
Three published price models, from each vendor's own page on September 28, 2026:
- Intercom Fin: Intercom lists Fin at $0.99 per outcome, with helpdesk seats at $29 (Essential), $85 (Advanced) or $132 (Expert) per seat a month.
- Salesforce Agentforce: Salesforce lists $2 per conversation as one option, or Flex Credits at $500 per 100,000 credits, with the page's example of one action at 20 credits, $0.10. Agentforce 1 Editions start at $550 per user a month and add-ons at $125 per user a month.
- Microsoft Copilot Studio: Microsoft lists a $200 monthly pack of 25,000 Copilot Credits, plus pay as you go. I did not extract how many credits a typical conversation consumes, so I can only tell you the pack works out to $0.008 a credit, not what a conversation costs.
| Option | Published price | 3,000 conversations a month | Multiple of the custom run cost |
|---|---|---|---|
| Custom text agent on Claude Sonnet 5.5, full stack | Tokens plus $74 of tools | About $152 | 1x |
| Intercom Fin, half the conversations count as outcomes | $0.99 per outcome, plus seats | $1,485 plus seats | About 9.8x |
| Intercom Fin, every conversation counts as an outcome | $0.99 per outcome, plus seats | $2,970 plus seats | About 19.5x |
| Agentforce, per conversation option | $2 per conversation | $6,000 | About 39.5x |
| Agentforce Flex Credits, assuming 3 actions per conversation | $0.10 per 20-credit action | $900 | About 5.9x |
Walk the Fin row. 3,000 outcomes times $0.99 is $2,970 a month, before seats. Divide by the custom stack's $152 and you get about 19.5x. If only half the conversations count as outcomes, it's $1,485, about 9.8x. Agentforce's per-conversation option at $2 is $6,000, about 39.5x. The Flex Credits row assumes three actions per conversation, which is my assumption, not Salesforce's.
Now the honest part. Those multiples compare a platform's price against my run cost with nobody's time in it. The platform price includes the vendor's engineering, its evaluation work, its model migrations and its uptime. The custom number doesn't. That's not a small asterisk. It's most of the reason the platforms can charge what they charge.
Read the definition of a resolution before you do any math.A per-outcome price is only as meaningful as the vendor's definition of an outcome, and the vendor writes that definition. Does a customer who leaves without replying count? Does a conversation that later reopens as a ticket count? The answer moves your bill up or down by as much as 2x at the same traffic. Ask for the definition in writing and for a monthly export of every conversation that was billed.
Here's the break-even, done out loud, as a scenario that assumes you pay the list running costs above and leaves maintenance out for now. A $10,000 packaged build replacing Fin at 3,000 billed outcomes saves $2,970 minus $152, or $2,818 a month on list prices. $10,000 divided by $2,818 is about 3.5 months. At 1,500 billed outcomes the saving is $1,333 a month and the payback stretches to about 7.5 months. At 300 billed outcomes Fin costs $297 a month, the saving is $145, and the payback is about 69 months, which is longer than any model version I'd plan around. Below a few hundred resolutions a month, rent.
And if you already live in one of those ecosystems, the integration you'd be paying us to build might already exist inside the platform. That's a real reason to rent even when the arithmetic says build. The customer support agent package makes the most sense when your helpdesk isn't one of those platforms, or when the per-outcome bill has started to feel like rent on your own customers.
Maintenance is the real bill
After launch, the biggest cost of an agent is usually people time: evaluation, fixes and model migrations. This is the line both wrong models miss.
Start with the part you can't opt out of. Models get retired on the provider's schedule, not yours. Anthropic's deprecations page commits to at least 60 days' notice for publicly released models, and the recent record shows what that looks like in practice:
- Claude Sonnet 3.7: deprecated October 28, 2025, retired February 19, 2026.
- Claude Haiku 3.5: deprecated December 19, 2025, retired February 19, 2026.
- Claude Sonnet 4 and Opus 4: notice April 14, 2026, retired June 15, 2026.
- Claude Opus 4.1: notice June 5, 2026, retired August 5, 2026.
OpenAI's deprecations page shows the same rhythm. The gpt-4o-realtime-preview models were announced for shutdown on September 15, 2025 and shut down May 7, 2026. gpt-3.5-turbo-0125 and gpt-4-0613 were announced April 22, 2026 for shutdown October 23, 2026. whisper-1 and gpt-4o-transcribe were announced August 26, 2026 for shutdown February 26, 2027. And preview models can get as little as about two weeks' notice, which is why I flagged Gemini 3.1 Pro as a preview model earlier.
Count the Anthropic list: four retirement events inside about ten months. If your agent sits on one model, expect to migrate it roughly once or twice a year. Each migration means rerunning your evaluation set, reading the failures, adjusting prompts, and checking the token count, since the newer tokenizer alone can add about 30% to the bill.
What does an evaluation rerun cost? The tokens are trivial. Suppose your test set is 200 recorded conversations at the scenario 1 sizes. That's 1.6 million input tokens at $2 ($3.20) and 200,000 output tokens at $10 ($2) on Sonnet 5.5: about $5.20 a run, and half that through the Batch API. The people cost is not trivial. Someone has to read the failures and decide what to change.
So here's the assumption I use for a small, stable agent: about 24 hours of senior time a year, meaning a quarterly review of transcripts and evaluations plus one model migration. At $150 to $225 an hour that's $3,600 to $5,400 a year. Compare it with scenario 1's Sonnet token bill of $78 times 12, or $936 a year. Maintenance is about 3.8x to 5.8x the token line. That's the multiplier to remember.
For an agent that touches several systems or changes monthly, 24 hours is far too low, and that's what a retainer is for. Ours run $2,500 to $9,500 a month, or $30,000 to $114,000 a year. Hiring someone in-house is the other route. The Bureau of Labor Statistics puts the median software developer wage at $135,980 for May 2025, and its compensation survey shows wages at about 70% of total private industry compensation in June 2026. Dividing $135,980 by 0.70 gives an estimated loaded cost of about $194,000 a year. That's my estimate from an all-industry benefit share, not an occupation-specific figure, and it excludes recruiting, equipment and equity.
Then there's the security work, which never finishes. Any agent that reads text from customers is exposed to prompt injection, and the UK's National Cyber Security Centre has argued there's a good chance it will never be mitigated the way SQL injection was, because models don't separate instructions from data. So the job isn't to solve it; it's to shrink the blast radius, keep anything that commits you in deterministic code, and review the logs. That's a recurring cost, and it belongs in the budget. How to set up the tracing and test sets that make this cheap is covered in AI agent evaluation and observability.
Sticker price versus ownership
Car buyers have a name for this problem: total cost of ownership.It's the purchase price plus everything it costs to keep the car on the road for as long as you own it, divided over the years you drive it.
The concept exists because sticker prices lie by omission. A cheap car with expensive parts and poor fuel economy can cost more over five years than a pricier one that sips fuel and rarely breaks. Nobody thinks the dealer is wrong about the sticker. They just know it's the first number, not the number.
Map it onto an agent and every line from this article has a slot:
- Sticker price: The build: $10,000 for a package, $28k-$70k and up for bespoke.
- Fuel: Tokens and voice minutes, which scale with every conversation.
- Insurance and registration: Hosting, retrieval and observability, which you pay whether the agent is busy or not.
- Servicing: Evaluation reruns, prompt fixes and transcript review.
- Recall notices: Model retirements, which arrive on the manufacturer's schedule: at least 60 days' notice from Anthropic for public models, and as little as about two weeks for OpenAI preview models.
And the car-buying lesson carries over directly: the cheapest fuel doesn't make the cheapest car. Here's year one for three scenarios, assembled from every line above.
| 12 months | Build | Run (12 months) | Maintenance (assumed) | Total year one |
|---|---|---|---|---|
| A. Packaged support chat agent, 3,000 conversations/month | $10,000 | $1,824 | $3,600 to $5,400 | $15,424 to $17,224 |
| B. Packaged voice receptionist, 1,000 minutes/month (Retell stack) | $10,000 | $1,812 | $3,600 to $5,400 | $15,412 to $17,212 |
| C. Bespoke single-workflow agent, 30,000 conversations/month | $28,000 to $70,000 | $12,588 | $30,000 to $114,000 | $70,588 to $196,588 |
| Comparison: Intercom Fin, 3,000 outcomes/month | $0 | $35,640 plus seats | Vendor's, inside the price | $35,640 plus seats |
How I got each row. The package pages don't say who pays running costs after launch, so the run column is my assumption at vendor list prices, whoever ends up paying them. Scenario A: the $10,000 package, $152 times 12 is $1,824 of run cost, plus the $3,600 to $5,400 maintenance assumption. Scenario B: the $10,000 package, $151 times 12 is $1,812 for the Retell stack alone (Retell hosts the call, so I left out the optional $49 of Vercel and Langfuse), plus the same maintenance assumption. Scenario C is a bespoke single-workflow agent at ten times scenario 1's traffic: Sonnet tokens of $780, Pinecone Standard at $50, Vercel Pro at $20 and Langfuse Pro at $199 make $1,049 a month, or $12,588 a year, with a retainer for maintenance. All three are scenarios with stated assumptions, not client results.
Now read the table the way a car buyer would. In scenarios A and B, the build is about 60% of year one and the run is about 11%. In scenario C, the run is between about 6% and 18% of year one depending on where in the range you land, and maintenance can be larger than the build. Nowhere in that table are tokens the biggest line. That's the meter model, disproved on its own numbers.
Year two is where ownership pays. The build doesn't recur. Scenario A drops to roughly $5,424 to $7,224 a year, run plus maintenance. Fin at the same 3,000 outcomes is still $35,640 plus seats. Over two years, that's about $20,800 to $24,500 owned against about $71,300 rented, roughly a 3x gap, and it's only real if your resolution count is actually 3,000 and the maintenance actually gets done.
When it is not worth it
An agent isn't worth building when the payback runs longer than the model it sits on.That's my working test, and it rules out more projects than it rules in.
Do the division before you do anything else. Take the build cost, divide by the monthly saving you can measure today (not the one you hope for), and compare the result with how long the model is likely to stay live. With a retirement cadence of roughly a year, a payback of three or four months is comfortable, a year is marginal, and five years is a hobby.
Beyond the arithmetic, these are the situations where I tell people to wait:
- Low volume: A few hundred conversations a month. A per-resolution product, or a well-written FAQ page, is cheaper until volume grows.
- A process that changes weekly: Every change is maintenance time. Stabilize the process first, then automate it.
- Expensive, uncatchable mistakes: If a wrong answer costs real money and no deterministic check can catch it before it reaches the customer, a person should own that step.
- No one to own it: An agent nobody reads the transcripts of drifts quietly. If there's no owner, budget for a retainer or don't launch.
- The platform you already pay for does it: If your helpdesk or CRM already sells the agent at a price that works at your volume, building a parallel one is vanity.
Admitting this costs us work, and I'd rather lose a project at the discovery stage than build something that gets switched off in month five.
Red flags in a cost quote
A good AI agent quote separates the build from the run and shows its assumptions; a bad one gives you a single number you can't check.These are the red flags I'd look for in any proposal, including ours.
- One blended monthly number: If the quote says a monthly figure with no volume assumption, you can't tell whether it covers 500 conversations or 50,000. Ask for the per-conversation or per-minute math.
- No named model: If the quote doesn't name the model and its list price, you can't check the token line or plan for its retirement.
- Silence on who pays usage: Tokens, telephony and hosting have to be paid by someone. If the quote doesn't say whether that's you directly, the agency at cost, or the agency with a markup, ask before signing.
- No evaluation plan: If nothing in the quote describes a test set and how it gets rerun, the maintenance line is either missing or hidden.
- No retirement plan: Ask what happens when the model is retired. The answer should name who does the migration and how it's billed.
- Savings percentages without a source: A quote that promises the agent will cut costs by a round percentage should be able to show where the percentage came from. Usually it can't.
- No IP transfer: If you don't own the code, every future change is on the builder's terms, and the switching cost goes into your total cost of ownership.
The broader vetting checklist for the builder, not the quote, lives in the sibling article on red flags when hiring an AI agent developer, linked at the bottom of this page.
Limitations and refusals
Every price in this article is a list price on one day, and several lines rest on my assumptions.Here's what that means for how much weight to put on them.
- Prices move. Google has already published a doubling for Gemini 3.8 Flash on January 1, 2027, Deepgram's lowest speech rate is a limited-time promotion, and ElevenLabs' text to speech rates carry promotions ending in October. Recheck every line before you commit.
- List is not negotiated. A committed or negotiated price can differ from list on models and platforms; I can only print what's public.
- The token volumes are assumptions. 8,000 input and 1,000 output tokens per conversation is a plausible support workload, not a measurement of yours.
- I could not confirm whether ElevenLabs' $0.08 agent minute includes the language model, or how many Copilot Studio credits a conversation uses, or LangSmith's exact overage rate.
- The maintenance figure of about 24 hours a year is my planning assumption for a small, stable agent. It's the least certain number in the article and the most important.
- Clutch's averages come from reviews of listed firms and skew offshore, and its page blocks automated fetches, so I could not re-verify it on the day of writing.
And here's what I refuse to print, because the numbers don't survive a source check.
"AI agents cut support costs by 30%"and its cousins at other round percentages. I looked for a primary source and found none. The percentage appears in vendor copy; it doesn't appear in a study with a method.
"AI agents cost $5k to $500k to build"style ranges. Several development shops publish ranges like this with no cited source and no sample. Our own bands are ours, labeled as ours, and the only market figure above is Clutch's, with its method and its skew stated.
Gartner's $80 billion.Gartner did predict, in an August 31, 2022 press release, that conversational AI would reduce contact center agent labor costs by $80 billion in 2026. That was a forecast made four years ago, not a measured 2026 result, and I've found nothing that measures whether it came true. If a vendor cites it as a fact about this year, that's a tell.
Failure-rate statistics, all refused."95% of GenAI pilots fail", "85% of AI projects fail", "87% never reach production" and the Gartner "40% of agentic projects cancelled" line all trace to thin or misread sources with no traceable methodology for the claim as quoted, and I don't use any of them to scare anyone into or out of a budget.
I also don't compare an AI conversation against a "human cost per conversation" figure, because I haven't found one with a source I trust. The per-conversation numbers in this article come from my own arithmetic on list prices, which you can check.
Three things this week
You can have a defensible AI agent budget by Friday.This is the order I'd do it in.
- 1.Pull last month's numbers: total support conversations or inbound call minutes, and the average length. Those two figures drive the whole run-cost line, and most people are guessing at them.
- 2.Build the four-line budget from this article in a spreadsheet: build, run (tokens or minutes plus the three plumbing subscriptions), maintenance hours at a real hourly rate, and a model migration once a year. Then run it again for a per-resolution platform at your actual volume.
- 3.Divide the build cost by the monthly saving you can measure today. If the payback is under about a year, get two quotes and check them against the red flags above. If it's longer, wait, rent, or fix the process first.
If the math says build, the discovery call is where we price the run cost at your volume, not ours. Book it at calendly.com/frenchydigital/discovery-call. Either way, do the division first.
Want the Whole Bill Before You Build?
Book a discovery call with Frenchy Digital, a senior-led Black-owned Los Angeles agency. We estimate your agent's run cost at your real volume and send a fixed-price phased proposal within 5 business days. Call +1 (424) 272-5601.
Want the Whole Bill Before You Build?
Book a discovery call. We map the workflow, estimate the monthly run cost at your real volume, and send a fixed-price phased proposal within 5 business days.
1517 S Bentley Ave Apt 204, Los Angeles CA 90025
Frequently Asked Questions
Sources & References
- 1Anthropic, Claude API pricing (checked September 28, 2026)↗
- 2OpenAI, API pricing (checked September 28, 2026)↗
- 3Google, Gemini API pricing (checked September 28, 2026)↗
- 4Twilio, US voice pricing↗
- 5Retell AI, Pricing↗
- 6ElevenLabs, API pricing↗
- 7Deepgram, Pricing↗
- 8Pinecone, Pricing↗
- 9Supabase, Pricing↗
- 10Vercel, Pricing↗
- 11LangChain, LangSmith pricing↗
- 12Langfuse, Pricing↗
- 13Intercom, Pricing including Fin↗
- 14Salesforce, Agentforce pricing↗
- 15Microsoft, Copilot Studio↗
- 16Clutch, Artificial intelligence development pricing guide (September 2026)↗
- 17Anthropic, Model deprecations↗
- 18OpenAI, Deprecations↗
- 19US Bureau of Labor Statistics, Software developers (15-1252), May 2025↗
- 20US Bureau of Labor Statistics, Employer Costs for Employee Compensation, June 2026↗
- 21UK NCSC, Prompt injection is not SQL injection (December 8, 2025)↗

