Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    Strategy
    August 9, 2026
    26 min read

    Build vs Buy for AI Agentsin 2026

    A decision framework grounded in what the evidence actually shows. Before the framework, an audit of the four failure statistics every build-vs-buy article leans on — and the numbers that survive being read.

    Build versus buy decision framework for AI agents in 2026 — evaluating failure statistics, token costs, vendor benchmarks, and the hybrid path
    28% / 20%
    AI use cases that fully succeed vs fail outright
    Gartner, Apr 2026 — N=782 I&O leaders
    67% vs 33%
    Deployment rate: external partnerships vs internal builds
    MIT NANDA 2025 — self-reported
    15×
    Tokens a multi-agent system uses versus a chat interaction
    Anthropic engineering, multi-agent research system
    $30k–$80k
    Single production agent with evals and observability
    Frenchy Digital scoping 2026

    Key Takeaways

    • The pessimistic canon does not survive being read. The MIT 95% figure is a non-peer-reviewed v0.1 working paper based on 52 interviews and 153 conference-collected surveys, measuring whether custom tools reached production in six months — and the same report puts generic chatbot pilot-to-implementation near 83%.
    • Gartner's over-40%-cancelled-by-2027 line is an analyst prediction with no published derivation. The only data in that release is a January 2025 poll of 3,412 webinar attendees measuring investment posture, not cancellations.
    • The two older statistics are citation laundering: 85% misquotes a 2018 Gartner prediction about biased outputs, and 87% traces to a 2019 sponsored post where the number appears only in the writer's own rhetorical question.
    • The numbers that are measured: Gartner's April 2026 survey of 782 I&O leaders found 28% of AI use cases fully succeed and meet ROI while 20% fail outright; Deloitte's 3,235-leader study found 66% realizing productivity gains but only 20% revenue growth against 74% who hoped for it.
    • The most decision-relevant finding is buried in the MIT report itself: external partnerships with learning-capable customized tools reached deployment about 67% of the time versus about 33% for internally built tools. Self-reported, but directly on-topic and almost never cited.
    • Model cost from the multipliers, not the sticker: agents use roughly 4× the tokens of chat and multi-agent systems roughly 15×. Then add evaluation infrastructure, observability at $39–$249 per month and up, security review, and one model migration per year.
    • No independent body validates agent vendor benchmarks. The official SWE-bench Verified board is academic-only and frozen at 79.2% from December 2025; fresh decontaminated tasks put frontier models near 64%; a UC Berkeley team hit 100% on five benchmarks without solving a task.
    • Frenchy Digital cost bands: discovery $9k–$22k; single production agent $30k–$80k; multi-workflow platform $80k–$200k; enterprise or regulated build $200k–$450k+.

    Audit the Failure Statistics Before You Use Them

    Almost every build-versus-buy article about AI agents opens the same way: with a failure rate. Ninety-five percent of pilots fail. Forty percent of agentic projects will be cancelled. Eighty-five percent of AI projects fail. Eighty-seven percent never reach production. The numbers are used to argue in both directions — buy, because building fails; build, because vendors fail — and they are almost never checked.

    We checked them. All four trace back to documents you can read, and reading them changes the decision materially. One is a non-peer-reviewed working paper measuring something much narrower than its headline. One is an analyst prediction with no published derivation. Two are citation laundering, where a real document said something different and the paraphrase hardened into a statistic.

    This matters for a practical reason, not a pedantic one. If you believe 95% of AI projects fail, the rational move is to buy the smallest thing that works and wait. If you believe the measured 2026 numbers — roughly a quarter of use cases fully succeeding, a fifth failing outright, and most landing somewhere in between — the rational move is to pick your first workflow carefully, instrument it, and expect a mixed portfolio. Those are different companies a year later.

    The one-line version: the pessimistic canon is weaker than its reputation, the optimistic vendor case is weaker than its benchmarks, and the most decision-relevant number in the whole literature — external partnerships reaching deployment about twice as often as internal builds — is buried inside the report everyone quotes for the opposite conclusion.

    What follows is written for someone who has to sign the decision: a CTO, a VP of engineering, a technical founder. It audits the statistics, gives you the ones that hold up, then supplies a seven-dimension framework, a real cost model with August 2026 prices, and the vendor questions that actually separate products from demos.

    The claim you will hearWhat it actually isWhat it can legitimately support
    “95% of GenAI pilots fail”A 2025 MIT NANDA working paper labelled v0.1 and self-described as Preliminary Findings. Not peer-reviewed; its only listed reviewer is a co-author. Method: a review of over 300 publicly disclosed initiatives, structured interviews with 52 organizations, and 153 senior-leader surveys collected at four industry conferences. Its exact sentence is that 95% of organizations are getting zero return.That custom, task-specific GenAI tools frequently did not reach production inside a six-month observation window, as reported by executives at conferences. Nothing about generative AI as a category — the same report puts generic chatbot pilot-to-implementation rates near 83%.
    “Over 40% of agentic AI projects will be cancelled by 2027”A Gartner Predicts statement from a June 25, 2025 press release, attributed to Senior Director Analyst Anushree Verma. No derivation is published. The release's only data instrument is a January 2025 poll of 3,412 webinar attendees that asked about investment posture: 19% significant investment, 42% conservative, 8% none, 31% wait and see.That an analyst firm expects cancellations to rise. It cannot support a cancellation rate, because no cancellations were measured. Downstream coverage that says “based on a poll of more than 3,400 organizations” has welded a real sample size onto a figure that sample never produced.
    “85% of AI projects fail”A corruption of a February 2018 Gartner prediction. The original text: through 2022, 85 percent of AI projects will deliver erroneous outcomes due to bias in data, algorithms or the teams responsible for managing them.A caution about output quality and bias, made before the current model generation existed. It is not a project failure rate and never was.
    “87% of data science projects never reach production”A July 2019 VentureBeat page bylined “VB Staff” and explicitly labelled Sponsored. The number appears only inside the writer's own rhetorical question, which is internally inconsistent — it says 13% and then “just one out of every 10” in the same sentence. No study, no sample size, no methodology. The panel it promotes is never quoted stating the figure.Nothing. There is no underlying research to cite.
    “42% of companies abandoned most of their AI initiatives”Real research — S&P Global Market Intelligence, more than 1,000 respondents across North America and Europe — but published March 14, 2025 and describing 2024 into early-2025 behavior, up from 17% the prior year. The same work found the average organization scrapped 46% of proofs of concept before production.A 2024-to-early-2025 snapshot. It is routinely relabelled as current-year data. If you use it, date it.
    “88% of agent pilots never reach production”No traceable origin. It circulates without an attributable study, sample, or publication.Nothing. Do not put it in a board deck.

    Provenance audit of the six failure statistics most often quoted in AI build-versus-buy arguments — Frenchy Digital, August 2026.

    The MIT 95%: What the Document Actually Says

    This is the load-bearing statistic of the pessimistic case, so it deserves the most scrutiny. The source is The GenAI Divide: State of AI in Business 2025, a 26-page document from MIT NANDA covering a research period of January to June 2025. Six things about it change how much weight it can carry.

    • It is a v0.1 working paper, not peer-reviewed research: The document is self-labelled Preliminary Findings. There is no peer review. The single listed reviewer is one of its own co-authors. Its disclaimer states the views are solely the authors' and not those of their affiliated employers — which means it was never institutionally underwritten MIT research, whatever the shorthand “MIT study” implies.
    • The sample is small and conference-collected: The stated method is a systematic review of over 300 publicly disclosed AI initiatives, structured interviews with representatives from 52 organizations, and survey responses from 153 senior leaders collected across four major industry conferences. The 153 are a convenience sample of people who attend AI conferences, which is not a neutral population when the question is AI adoption.
    • It measures custom tools inside a six-month window: The funnel that produces the headline splits general-purpose LLMs from custom, task-specific GenAI. The pessimistic figure is the complement of the task-specific bar. The same report says generic chatbots show high pilot-to-implementation rates of roughly 83%, and that workers at over 90% of surveyed companies reported regular use of personal AI tools. “Generative AI fails” is not what the document found.
    • Success was scored on whether an executive remarked on it: The report defines success as initiatives users or executives have remarked as causing a marked and sustained productivity or P&L impact, measured six months after pilot. That is a subjective, interview-derived criterion, not a financial measurement.
    • The report's own limitations section concedes the point: It states that six months may be insufficient and that the approach is potentially understating success rates, and that its funnel is directionally accurate based on individual interviews rather than official company reporting. Those are the authors' words about their own headline number.
    • The authors had a commercial interest in the thesis: NANDA is an MIT Media Lab group commercializing agent infrastructure. The report doubles as an argument for its own thesis. That is not disqualifying on its own — most industry research has this shape — but it belongs in the citation.

    There is a further problem with the funnel itself, which is why you will not find its stage percentages reproduced here. Independent extractions of the same PDF disagree on those numbers, and the report contradicts itself: the percentage shown in its exhibit for one funnel stage does not match the percentage in its own executive summary. A number a document cannot state consistently is not a number you should put in a board deck.

    The media record is also wrong, and nobody fixed it. The Fortune article of August 18, 2025 that made this viral described the method as roughly 150 interviews, a survey of 350 employees, and 300 public AI deployments. The report says 52 organizations and 153 senior leaders. No correction was issued, and most downstream coverage — including some otherwise careful secondary research — propagates Fortune's version rather than the primary document's.

    The institutional status is the last piece, and precision matters. The original NANDA subdomain now redirects to the MIT Media Lab group overview page, and that page's publications listing does not include this report at all. The PDF survives on third-party mirrors. It has never been formally retracted. So the accurate characterization is: absent from its own group's publication list, with the original URL redirecting away — softer than a retraction, stronger than nothing.

    The sharpest critique came from Arnon Shimoni in a August 2025 response, and his definitional objection is the part worth keeping:

    MIT counts any pilot that doesn't reach “full production deployment” as a failure… This is like saying 95% of git branches fail because they don't get merged to main.

    Arnon Shimoni, “MIT's 95% AI Failure Rate Is Wrong,” August 27, 2025

    Apply the same skepticism to the critic. Shimoni works at a vendor, and his counter-estimate of a real failure rate somewhere around 25 to 30% is proprietary and unauditable. Cite the definitional argument — that treating every abandoned proof of concept as a failure conflates exploration with defeat — and leave his number where you found it.

    What to say when someone brings the 95% to a meeting

    Do not argue that AI projects mostly succeed; that is not what the evidence shows either. Say instead that the figure comes from a preliminary working paper measuring whether custom tools reached production inside six months, that the same paper puts generic chatbot implementation near 83%, and that its own limitations section says six months may understate success.

    Then move the conversation to the Gartner April 2026 survey, which measured outcomes rather than predicting them. It is a better number for making a decision, and it is less dramatic in both directions.

    The Other Three Numbers Everyone Quotes

    The Gartner agentic cancellation prediction is the one most often cited in agent-specific writing. The exact wording, from a June 25, 2025 press release attributed to Senior Director Analyst Anushree Verma, is that over 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs, unclear business value or inadequate risk controls.

    It is a prediction. Gartner publishes Predicts statements as analyst opinion, and there is no derivation in the release. The only data instrument in the document is a January 2025 poll of 3,412 webinar attendees, which asked about investment posture — 19% reported significant investment, 42% conservative, 8% none, 31% wait and see. It contains no cancellation measurement of any kind. Respondents were self-selected webinar attendees, not a sampled panel.

    The specific error to avoid: nearly all downstream coverage frames the 40% as “based on a poll of more than 3,400 organizations.” It is not. A real sample size has been welded onto a figure that sample never produced. And for the record, the prediction has not been walked back — the release stayed live and unedited into 2026. Anyone telling you Gartner retracted it is also wrong.

    The two older statistics are simpler. “85% of AI projects fail” is a corruption of a February 2018 Gartner prediction whose actual text was that through 2022, 85 percent of AI projects will deliver erroneous outcomes due to bias in data, algorithms or the teams responsible for managing them. That is a claim about output quality in a pre-transformer-boom world, not a project failure rate.

    “87% of data science projects never reach production” is worse. It traces to a July 2019 VentureBeat page bylined “VB Staff” and explicitly labelled Sponsored. The number appears only inside the writer's own rhetorical question, which asks why only 13% of data science projects, or just one out of every 10, actually make it into production — 13% and one in ten in the same sentence. There is no study, no sample, and no methodology, and the panel the page promotes is never quoted stating the figure.

    One more worth flagging because it appears constantly in 2026 decks: the “42% of companies abandoned most of their AI initiatives” figure from S&P Global Market Intelligence is real research with more than 1,000 respondents — but it was published March 14, 2025 and describes 2024 into early-2025 behavior, up from 17% the year before. It is routinely relabelled as current data. If you use it, date it.

    There is also a tension in the pessimistic case that nobody addresses. The same analyst firm forecasting mass agentic cancellation also forecasts roughly 47% growth in AI spending in 2026. Whatever is happening, spend is not retreating. A narrative that predicts collapse while the budget line grows is describing churn and reallocation, not abandonment.

    What Is Actually Measured in 2026

    Here are the numbers that survive scrutiny. They are less quotable than the canon, which is exactly why they are more useful. Note the shape they share: a minority of clear wins, a minority of clear failures, and a large partial-success middle that neither camp's narrative accommodates.

    SourceSample and field datesWhat it foundCaveat to state alongside it
    Gartner — AI projects in I&O (Apr 7, 2026)N=782 infrastructure and operations leaders, fielded Nov–Dec 202528% of AI use cases fully succeed and meet ROI expectations; 20% fail outright; the remainder land in partial success. 77% delivered at least one successful use case, 57% reported at least one failure. Skills gaps and data quality each cited by 38%.Scoped to infrastructure and operations, not all enterprise AI. Analyst-firm survey; full methodology not published.
    Gartner — data foundations (Apr 16, 2026)Fielded Nov–Dec 2025; sample size not stated in the release bodyOnly 39% of technology leaders are confident current AI investments will positively affect financial performance. Organizations with successful AI initiatives invest up to four times more, as a share of revenue, in data quality, governance, people and change management. Highest-maturity organizations report up to 65% greater business outcomes.No disclosed N. The four-times figure is a comparison of self-reported spend, not a controlled effect.
    Deloitte — State of AI in the Enterprise 2026N=3,235 senior leaders across 24 countries, fielded Aug–Sep 202566% realizing productivity and efficiency gains, 53% better insights, 40% cost reduction, and 20% revenue increase — against 74% who hoped to grow revenue through AI. 34% using AI for deep transformation, 30% redesigning key processes, and 37% still surface-level with minimal process change.Deloitte sells AI services. Transparent sample and methodology, interested party. Note the title changed: “Generative” was dropped for 2026.
    Stanford HAI — 2026 AI IndexAggregated multi-source index, published 202688% of organizations use AI in at least one function, but under 10% have fully scaled AI in any single function and only 29% report significant ROI. 78% of Fortune 500 have deployed at scale. Inaccuracy is now the top-cited risk at 74%, up 14 points year over year.The Index measures adoption and consumer surplus. It is not a project-outcome study; do not use it as a failure rate.
    Gartner — manager expectations (Mar 4, 2026)Backed by July 2025 surveys: N=2,986 employees and N=1,973 managers45% of managers say AI has improved their teams' work as much as they expected. Only 14% of managers report facing no challenges.Self-reported expectation matching, not financial outcome. Useful as counter-evidence to the collapse narrative, not as ROI proof.

    The measured 2026 evidence on enterprise AI outcomes, with the caveat each one requires — Frenchy Digital, August 2026.

    The April 7, 2026 Gartner I&O survey is the single most defensible figure in this space right now: 782 leaders, fielded November to December 2025, with 28% of use cases fully succeeding and meeting ROI expectations against 20% failing outright. Read the second pair of numbers too — 77% of organizations delivered at least one success and 57% reported at least one failure. Most organizations are doing both simultaneously, which is what a portfolio looks like, not what a collapse looks like.

    Deloitte's 2026 study — 3,235 senior leaders across 24 countries, fielded August to September 2025 — supplies the cleanest expectation-versus-reality gap available: 74% hope to grow revenue through AI, and 20% already are. Meanwhile 66% report productivity and efficiency gains. That combination is the honest description of the current state. AI is reliably taking cost out and unreliably putting revenue in. If your business case rests on the second, you are betting against the measured distribution.

    One title change is worth noting because it breaks citations: Deloitte dropped “Generative” from the report name for 2026. Anyone citing “State of Generative AI in the Enterprise” for a 2026 figure is citing a dead title, which is a fast way to spot a deck assembled from search results.

    Stanford HAI's 2026 AI Index supplies the adoption-versus-depth contrast: 88% of organizations use AI somewhere, under 10% have fully scaled it in any single function, and only 29% report significant ROI. The Index measures adoption and consumer surplus rather than project outcomes, so do not convert those numbers into a failure rate — but the gap between 88% adoption and sub-10% scaling is the real story of 2026, and it is a scaling problem, not a pilot problem.

    The finding with the clearest build-versus-buy implicationcomes from Gartner's April 16, 2026 release: organizations with successful AI initiatives invest up to four times more, as a share of revenue, in data quality, governance, people and change management. Not in models. Not in agents. In the substrate underneath. If your build budget is 90% engineering and 10% data work, you have inverted the ratio that separates the successes from the failures.

    The Buried Finding: 67% Versus 33%

    Here is the part almost nobody cites, and it is in the same MIT report that produced the 95% headline. It is directly about build versus buy, and it points the opposite way from how the report is usually deployed.

    External partnerships with learning-capable, customized tools reached deployment ~67% of the time, compared to ~33% for internally built tools. While these figures reflect self-reported outcomes and may not account for all confounding variables, the magnitude of difference was consistent across interviewees.

    MIT NANDA, The GenAI Divide: State of AI in Business 2025

    The report's team-structure exhibit shows the same split: strategic partnerships at 66%, internal development at 33%. And Shimoni, the report's most prominent critic, says this specific finding matches his own data exactly — which is notable, because he disputes almost everything else in it. When a report and its loudest critic agree on one number, that number is worth more than the headline they disagree about.

    Now handle it carefully, because it is the same document with the same problems. The outcomes are self-reported. Confounds are obvious and unmeasured: organizations that engage an external partner may already have clearer scope, executive sponsorship, and a budget line, all of which independently predict deployment. A partnership also creates external accountability that an internal project does not have. The causal arrow could run either way.

    What it does support is a specific, testable proposition: the structure of an AI project — who is accountable, whether scope is contractual, whether there is a defined delivery date — appears to matter at least as much as whether the code was written inside or outside the building. That is a claim you can act on without believing the 95%.

    The practical read: the finding is not “buying beats building.” It is that projects with external structure deploy roughly twice as often as projects without it. You can reproduce that structure internally — a named owner, a written scope, a fixed date, a defined acceptance test — and if you cannot, that is useful information about whether your organization is ready to build.

    The word “learning-capable” in that sentence also carries weight. The report's thesis is that tools which improve from feedback outperform static deployments. That maps directly onto whether you have an evaluation and observability loop at all — which is, in our experience, the single most reliable predictor of whether an agent project is still running twelve months later.

    The Build-vs-Buy Decision Framework

    Build versus buy is the wrong unit of analysis. An agent system has layers — model, orchestration, tools, retrieval, evaluation, observability, interface — and the right answer is usually different for each. What follows is the seven-dimension assessment we run in discovery, applied per layer rather than to the project as a whole.

    DimensionPoints to buildPoints to buyHow to test it in one week
    DifferentiationThe workflow is how you compete — your sequencing, your rules, your proprietary data are the product.The capability is table stakes: meeting notes, ticket triage, retrieval over public documentation.Ask whether a competitor could buy the same capability tomorrow. If yes, building it spends senior engineering time to reach parity.
    Data sensitivityRegulated data, contractual residency constraints, or customer agreements that name subprocessors.Data you would already put into a SaaS tool without a second thought.Ask for the subprocessor list, retention windows per artifact, and the deletion path. How fast they answer tells you more than the document.
    Integration depthThe agent must write into systems of record, honor your permission model, and respect row-level scoping.The integration is read-mostly and the vendor ships a connector you can test this week.Pick your ugliest internal system and ask for a working read-write demo against a sandbox in two weeks. Most vendors decline.
    Iteration speedYou expect to change prompts, tools, and policy weekly based on what production traces show.The workflow is stable and the vendor's release cadence is acceptable.Count how many times the equivalent manual process changed last quarter. That is your real iteration rate.
    In-house capabilityYou have engineers who have shipped and operated an LLM system in production, not just prototyped one.Your team is strong but has never run an eval suite, read traces at volume, or survived a model migration.Ask the team to write the eval set before the agent. If nobody can define pass and fail for twenty real cases, you are not ready to build.
    Regulatory exposureYou must produce audit logs, human-in-the-loop evidence, and pinned model versions on demand.No regulator-facing evidence is required, or the vendor already produces exportable, machine-readable audit logs.Ask for a redacted audit-log export covering one workflow for one day, including tool calls. A dashboard view is not an answer.
    Total cost of ownershipToken volume is high enough that per-seat licensing dominates, or you need to route across models and providers.Volume is low or spiky, and per-seat pricing is cheaper than the engineering required to replace it.Model twelve months at three times expected volume, including evaluation, observability, and one model migration. Re-run when a provider changes pricing.

    Seven-dimension build-versus-buy assessment applied per architectural layer — Frenchy Digital discovery framework, 2026.

    Two dimensions deserve elaboration because they are where the decision usually turns.

    In-house capability is consistently overestimated, and there is a specific reason. Prototyping an agent is genuinely easy now, and the distance between a working prototype and an operable system is larger than it has ever been. The prototype does not need evaluation, tracing, cost guardrails, retry semantics, injection defenses, or a model migration plan. The system needs all of them. Teams that have shipped a prototype often reason from that experience to a build estimate, and the estimate is wrong by a factor.

    The test we use is unglamorous: ask the team to write twenty evaluation cases with clear pass and fail criteria, drawn from real examples, before any agent code exists. Teams that can do this in a week are ready to build. Teams that cannot are telling you that the problem is not yet specified, which is a much cheaper thing to discover now than in month five.

    Iteration speed is the dimension that most often overrules the others. If production traces are going to change your prompt and tool design weekly for the first two quarters — and for a genuinely differentiated workflow, they will — then a vendor release cadence is a hard ceiling on how fast you can learn. Conversely, if the workflow is stable and well-understood, that same cadence costs you nothing and saves you an on-call rotation.

    A note on the capability question and AI coding assistants

    Build estimates in 2026 frequently assume that AI coding tools have made the engineering cheap. Measure that on your own team before you budget on it. The only randomized controlled trial published on the question — METR, July 10, 2025 — put 16 experienced developers on 246 real issues in their own repositories, averaging over 22,000 stars and a million lines of code, and found they were 19% slower with AI allowed. They had forecast a 24% speedup, and after finishing still believed AI had made them 20% faster.

    METR now labels that result historical: it studied early-2025 tools and workflows and does not necessarily reflect current ones, and METR has since changed its experimental design. Cite it as a 2025 finding about 2025 tooling, not as a 2026 fact. The durable part is the roughly 39-point gap between perceived and measured speed, which is a reason to instrument your own delivery rather than to distrust the tools.

    Honest Cost Modelling: Tokens First, Then Everything Else

    Most build-versus-buy cost models start from the per-token price, which is the least important number in the calculation. Start from the multipliers instead. Anthropic's engineering write-up on its multi-agent research system states that agents typically use about four times more tokens than chat interactions, and multi-agent systems about fifteen times more than chats. The same post reports that on one browsing benchmark, token spend alone accounted for 80% of performance variance.

    Those multipliers are vendor-published and directional — credible public per-task token data is thin and almost entirely vendor-sourced — but they are the best available anchors, and they are the right shape. A chat feature and an agent are not the same cost class, and a multi-agent system is not the same cost class as an agent. If your model does not have a 4× or 15× term in it, it is a chat cost model wearing an agent's name.

    Model — as of August 2026Input / MTokOutput / MTokCached input / MTokNotes
    Claude Opus 5$5.00$25.00$0.501M context. Anthropic charges no long-context premium — a 900k-token request bills at the same per-token rate as a 9k one.
    Claude Fable 5$10.00$50.00$1.00A tier above Opus. 1M context, 128K max output.
    Claude Sonnet 5 (introductory, through Aug 31 2026)$2.00$10.00$0.20Rises to $3.00 input and $15.00 output on Sep 1 2026. Put that date in your cost model.
    Claude Haiku 4.5$1.00$5.00$0.10200K context, 64K max output. The right default for high-volume classification and routing.
    gpt-5.6-sol$5.00$30.00$0.50GPT-5.6 family: 1.05M context, 128K max output, knowledge cutoff Feb 16 2026.
    gpt-5.6-terra$2.00$12.00$0.20Mid tier of the same family.
    gpt-5.6-luna$0.20$1.20$0.02The cheapest frontier-family option on OpenAI's pricing page.
    Gemini 3.1 Pro (preview)$2.00 up to 200k / $4.00 above$12.00 up to 200k / $18.00 above$0.20 / $0.40Google does charge a long-context premium — input and output both double above 200k tokens.
    Gemini 3.6 Flash$1.50$7.50$0.15Cache storage billed separately at $1.00 per million tokens per hour.

    Published API pricing per million tokens, verified against provider pricing pages on August 10, 2026. Prices change monthly — re-verify before committing to a model.

    Two worked examples from Anthropic's own documentation give a sense of the magnitudes. A customer-support conversation at roughly 3,700 tokens on Claude Haiku 4.5 works out to about $37 per 10,000 tickets. A one-hour coding session on Claude Opus 5 — 50,000 input tokens, 15,000 output, plus runtime — comes to $0.705, falling to $0.525 when 40,000 of that input is served as cache reads. Both are vendor examples, so treat them as order-of-magnitude anchors rather than forecasts.

    The mechanisms that actually move the bill are structural rather than per-token:

    • Prompt caching, and its silent failure mode: On Anthropic, a five-minute cache write costs 1.25× base input, a one-hour write 2×, and a cache read 0.1×. Break-even is one read for the five-minute TTL and two for the one-hour. The trap: the minimum cacheable prefix is model-dependent and non-monotonic — 512 tokens on Opus 5, Fable 5 and Mythos 5; 1,024 on Opus 4.8 and the Sonnet 5, 4.6 and 4.5 line; 2,048 on Opus 4.7; 4,096 on Opus 4.6, Opus 4.5 and Haiku 4.5. Below the floor it silently does not cache, and your bill quietly stays high.
    • Batch processing: The Batch API is 50% off both input and output and stacks with caching. Any workload that does not need a synchronous answer — nightly enrichment, backfills, evaluation runs — belongs there. This is the largest single lever most teams have not pulled.
    • Reasoning tokens bill as output: On both major providers, reasoning tokens are charged at the output rate, which is the expensive side of the ledger. Two defaults changed recently and both increase cost silently: GPT-5.6 defaults reasoning context to all turns, re-rendering earlier turns' reasoning into every subsequent request, where earlier models used the current turn only; and on Claude Opus 5, omitting the thinking parameter now runs adaptive thinking where the same call on Opus 4.8 ran with none.
    • Long context is not uniformly priced: Anthropic charges no long-context premium — a 900k-token request bills at the same per-token rate as a 9k one. Google does: Gemini Pro models double both input and output price above 200k tokens. If your architecture pushes long contexts, that difference dominates the model comparison.
    • Server-side tools are a separate meter: As of August 2026 Anthropic bills web search at $10 per 1,000 searches, code execution at $0.05 per hour after 1,550 free container-hours per organization per month, and Managed Agents session runtime at $0.08 per session-hour. Web fetch is free. A research agent can spend more on searches than on output.
    • Budgets are not all hard caps: Anthropic's session budgets on Managed Agents are a hard dollar cap enforced as a pre-request gate — the session pauses at budget_reached rather than terminating, and raising the budget resumes it. Task budgets are advisory: the model sees a countdown, but the only hard cap is max_tokens. Know which one you configured before you rely on it.
    The migration cost nobody budgets:Claude 4.7 and Sonnet 5 use a tokenizer that produces roughly 30% more tokens for the same text. The per-token price did not change; the cost per request did. Re-baseline with a token-counting endpoint after every model change, never with a multiplier applied to last quarter's numbers. This single item has broken more cost models than any pricing change.

    Finally, price the alternative honestly. Menlo Ventures' December 2025 enterprise survey put total enterprise AI spend at $37 billion for 2025, up from $1.7 billion in 2023, with $12.5 billion of that on foundation-model APIs. That is the market you are buying into, and it is repricing fast in both directions: Epoch AI's analysis found that the price of a fixed capability level fell somewhere between 9× and 900× per year depending on the benchmark, though that work was published in March 2025 with no verified 2026 update. Anything you build to optimize today's token price may be optimizing a constraint that dissolves. See our token economics deep dive for the full treatment.

    The Costs Teams Forget to Model

    Token spend is the visible cost and rarely the dominant one. Here is what a build actually carries, with real prices where prices exist.

    Hidden costWhat it actually isPrice signal, August 2026
    Evaluation infrastructureThe set of real failure cases, the graders, and the CI wiring that tells you whether a change helped. Anthropic's guidance is to start with 20 to 50 tasks drawn from actual failures, run each trial isolated from a clean environment, grade the outcome rather than the path, and read sampled transcripts weekly.Mostly labour. LangChain's numbers are the only concrete ones published: begin at 10% production trace sampling, use at least 20 labelled examples to calibrate a judge, and expect a human to sustain 50 to 100 trace reviews per hour.
    Observability and tracingSpan-level capture of prompts, tool calls, latencies and token spend, retained long enough to debug a complaint from three weeks ago. Note that the OpenTelemetry GenAI semantic conventions are still entirely experimental, moved repositories in 2026, and deprecated gen_ai.system in favor of gen_ai.provider.name — instrumentation you write today will need revisiting.$29 to $249 per month for a small team; $2,499 per month at Langfuse's enterprise tier; usage overages on top. See the tool table below.
    Security reviewThreat modelling against the OWASP Top 10 for LLM Applications, tool-permission design, and injection test cases in CI. Prompt injection is not solved and no mitigation makes an agent safe — the workable posture is defense in depth and blast-radius reduction, which costs design time on every new tool you add.A recurring engineering tax rather than a licence. Budget review time per tool, not per project.
    Model migrationProviders change defaults, tokenizers and deprecation dates on their own schedule. Claude 4.7 and Sonnet 5 use a tokenizer that produces roughly 30% more tokens for the same text at unchanged per-token prices. On Claude Opus 5, omitting the thinking parameter now runs adaptive thinking where the same request on Opus 4.8 ran with none. GPT-5.6 defaults reasoning context to all turns, so a multi-turn agent re-bills earlier reasoning.Re-baseline with a token counter after every model change — never with a multiplier. Assume at least one migration per year per model you depend on.
    Framework churnMicrosoft moved AutoGen and Semantic Kernel to maintenance mode in October 2025, superseded by Microsoft Agent Framework 1.0, GA April 3, 2026. The OpenAI Agents SDK is still pre-1.0. Google's ADK and Pydantic AI have both shipped major version 2. Anything you build on top of a pre-1.0 dependency carries a migration you have not scheduled.Zero licence cost, real engineering cost. Weight it toward frameworks with a governance story rather than a star count.
    Server-side tool usageAgent tools billed outside the token meter. As of August 2026, Anthropic charges $10 per 1,000 web searches, gives 1,550 free code-execution container-hours per organization per month and then $0.05 per hour, and bills Managed Agents session runtime at $0.08 per session-hour. Web fetch is free.Small per unit, large at agent volumes. A search-heavy research agent can spend more on searches than on output tokens.
    Reliability engineeringRetries, idempotency, partial-failure handling, and a rollback path for a prompt change that degrades quality. Non-determinism means your change control needs pinned versions and golden-set evaluation, not just code review.The gap between a demo and a system. It is also where most build estimates are wrong by a factor rather than a percentage.

    The recurring costs of operating an agent that build estimates routinely omit — Frenchy Digital, 2026.

    Observability deserves its own table because the pricing is knowable and the licensing differences are load-bearing. Two of these tools have a genuine self-host path; the rest do not, whatever their marketing implies. Note in particular that Arize Phoenix ships under Elastic License 2.0, which is not an OSI-approved open source licence.

    ToolEntry tierProduction priceLicence and self-host reality
    LangSmithDev tier free, 5k traces/monthPlus $39 per seat/month; overage $0.50 per 1k tracesProprietary; self-hosting only on Enterprise
    LangfuseHobby free, 50k units/monthCore $29/month, Pro $199/month, Enterprise $2,499/monthMIT licensed; self-host is genuinely full-featured
    BraintrustStarter freePro $249/month, no per-seat chargeProprietary; bring-your-own-cloud deployment
    Arize Phoenix / AXAX free to 25k spans/monthPro $50/monthPhoenix is Elastic License 2.0 — not OSI open source
    W&B WeaveFree 1 GB/monthFrom $60/month, then $0.10/MB (about $100/GB)Proprietary
    Datadog Agent ObservabilityFree 40k spans/monthPro $160/month annual for 100k spans, +$3.50 per 10kProprietary SaaS; the strongest OpenTelemetry support
    HeliconeHobby free, 10k requests/monthPro $79/month; Team $799/monthApache-2.0, ungated self-host; gateway-shaped

    Agent observability tooling pricing and licensing, checked against vendor pricing pages August 10, 2026.

    On evaluation, the useful ceiling is human throughput, not tooling cost. LangChain's guidance — the only source publishing concrete sampling numbers — suggests starting at 10% production trace sampling, using at least twenty labelled examples to calibrate an LLM judge, and assuming a human sustains 50 to 100 trace reviews per hour. Anthropic's evaluation guidance adds the structural rules: start from 20 to 50 tasks drawn from real failures, run each trial isolated from a clean environment because shared state produces correlated failures and inflated scores, grade the outcome rather than the path, and read sampled transcripts weekly.

    How common is that discipline? LangChain's State of Agent Engineering (n=1,340, fielded late 2025) found 89% of respondents have some observability but only 52.4% run offline evaluations, 37.3% run online ones, and 29.5% run none at all. That survey is vendor-run with self-selected respondents and biased toward eval adoption — and it still finds roughly three in ten doing nothing. An independent qualitative counterweight is bleaker in a different direction: in a 19-interview study, 12 of 19 practitioners called informal vibe checks irreplaceable and 13 of 19 had tried automated evaluation and described it as close to useless.

    Security is not a phase. Prompt injection is unsolved. No mitigation available in 2026 makes an agent safe, and any vendor or internal architecture document that implies otherwise is wrong. The workable posture is defense in depth and blast-radius reduction: treat all retrieved content as data rather than instruction, allowlist tools per workflow, deny by default any tool argument a human did not supply, and run injection cases in CI on every prompt change. Budget this as a per-tool design cost that recurs forever, not a one-time review. Our OWASP LLM Top 10 walkthrough covers the control set in detail.

    What Buy Actually Looks Like in 2026 — and Its Risks

    Buying is a real option and often the correct one. It is also not the safe option, which is how it tends to get framed. Three specific risks dominate.

    Discontinuation. Both flagship consumer browser agents shut down inside twelve months. OpenAI deprecated Operator and shut it down on August 31, 2025, folding the capability into ChatGPT agent. Google shut down Project Mariner on May 4, 2026, absorbing it into Gemini Agent and AI Mode, and the stated reasons are worth reading if you are evaluating any agentic product: heavy compute for real-time visual processing, slow performance, and form-selection errors. These were not small vendors. The same dynamic applies to platform features — OpenAI is winding down its fine-tuning platform and has closed it to new users, which is a dependency change for anyone whose product assumed it.

    Agent washing. In the same June 2025 release that produced the 40% prediction, Gartner estimated that only around 130 of the thousands of vendors claiming agentic AI are real, describing the rest as rebranded assistants, RPA, and chatbots. Treat that number as an analyst estimate rather than a count — no methodology is published for it, and the ratio it implies is startling enough that it deserves the caveat. Gartner renewed the agent-washing warning in 2026 for supply-chain planning specifically. The detection method does not depend on the estimate being right: ask what decision the system makes without a human, what tools it can call, what happens when a tool call fails, and to see a trace of a real multi-step run. Wrappers cannot produce the trace.

    Lock-in. The layers that lock you in are not the ones vendors emphasize. Your prompts, tool definitions, evaluation sets, and trace history are the assets. If those live only inside a vendor platform, switching cost is not a migration — it is a rebuild. This is the single strongest argument for building the orchestration layer even when you buy everything around it.

    What a good buy decision looks like in practice

    Buy the model API. Buy the tracing and evaluation platform, preferring one with an OSI licence and a real self-host path so the exit is technical rather than commercial. Buy connectors that already work against your systems — verified by a read-write demo, not a slide. Buy managed browser or sandbox infrastructure rather than operating your own fleet.

    Then insist on four contract terms regardless of vendor: pinned model versions with advance notice of changes and a published changelog; machine-readable export of traces, evaluation cases, and audit logs on demand; a named subprocessor list with retention windows per artifact; and a stated deletion path including derived data. A vendor that will not put those in writing before you sign will not build them after you sign either.

    One structural note in favor of buying at the protocol layer: the Model Context Protocol was donated by Anthropic to the Agentic AI Foundation, a directed fund under the Linux Foundation, on December 9, 2025, with Anthropic, Block and OpenAI as co-founders and Google, Microsoft, AWS, Cloudflare and Bloomberg supporting. Vendor-neutral governance plus cross-vendor client support is what makes something an actual standard rather than a proprietary interface with good marketing. Building tool integrations against MCP is a lower lock-in bet than building against any single vendor's plugin schema.

    Vendor Benchmarks Nobody Independently Validates

    Every agent vendor will show you benchmark scores. No independent body validates them. This is not a suspicion; it is the documented state of agent evaluation in 2026, and the gap between vendor-reported and independently-verified numbers is large enough to change a purchasing decision.

    What the deck saysWhat the primary source shows
    “We score 95–96% on SWE-bench Verified”The official leaderboard has accepted submissions only from academic and research institutions with open methods and an arXiv report since November 18, 2025. It holds 134 submissions; the most recent, dated December 15, 2025, is live-SWE-agent with Claude Opus 4.5 at 396 of 500 problems resolved — 79.2%. No 2026 frontier entries exist on that board. Circulating 95–96% figures are vendor self-reported.
    “Near human-level on real software tasks”On SWE-rebench's fresh, decontaminated set — 111 problems across 65 repositories, collected May 15 to July 1, 2026 — frontier models land near 64%: Fable 5 at 64.5% and $4.40 per problem, Grok 4.5 at 63.8% and $1.47, Opus 5 at 63.4%. Same class of task, roughly 30 points apart, and the difference is whether the model could have seen it.
    “SWE-bench Verified is the industry standard”OpenAI retired it in February 2026. It audited 138 problems o3 failed across 64 runs, each reviewed by at least six engineers, and found 59.4% had material test or description flaws — 35.5% overly narrow tests enforcing implementation details, 18.8% wide tests checking unspecified functionality.
    “85% on OSWorld”Self-reported tracker entries put Claude Opus 4.8 at 83.4% and Sonnet 4.6 at 78.5%. Anthropic's own model cards put the same models at 72.7% and 72.5% on OSWorld-Verified — a 6 to 11 point gap on identical models. The scaffold and the attempt budget explain more than the model does. The human baseline is 72.36%.
    “Our benchmark scores prove the agent works”A UC Berkeley team reached 100% on Terminal-Bench, SWE-bench Verified, SWE-bench Pro, FieldWorkArena and CAR-bench, about 98% on GAIA and about 100% on WebArena, without solving a single task — via pytest conftest hooks, config leakage, prompt-injected LLM judges and VM state manipulation. A follow-up catalogued 219 distinct flaws across ten benchmarks.
    “Independent leaderboards confirm our numbers”Aggregator sites inflate badly. One claimed 92.5% on ARC-AGI-2 and 91.9% on Terminal-Bench 2.0, contradicting primary sources by 8 to 55 points. ARC-AGI-2's own human panel averages 60%. Check the primary leaderboard or do not cite the number.
    “Higher reasoning effort means better results”Princeton's HAL ran 21,730 rollouts across nine models and nine benchmarks at a cost of roughly $40,000 and found higher reasoning effort reduced accuracy in the majority of runs. Its live board notes agents can be 100 times more expensive while only 1% better.

    Agent benchmark claims against primary sources — verified August 10, 2026.

    The reward-hacking result is the one to internalize. A UC Berkeley team reached perfect or near-perfect scores across five major agent benchmarks without solving a single task, using pytest configuration hooks, configuration leakage, prompt-injected LLM judges, and virtual machine state manipulation. A follow-up effort catalogued 219 distinct flaws across ten benchmarks and found that patching them dropped vulnerable tasks from near 100% to under 10% on four of them. The benchmarks are improving. They were not measuring what everyone assumed while the scores that populate vendor decks were being collected.

    There is also an evaluation-awareness effect. Anthropic's interpretability work found internal representations consistent with recognizing an evaluation context in roughly 26% of SWE-bench Verified problems, against under 1% in real product conversations. Whatever that means mechanistically, it argues against treating benchmark performance as a proxy for production behavior on your workload.

    The substitution: stop asking vendors for benchmark scores. Ask them to run your evaluation set — twenty to fifty real cases from your own failure history — and to give you the traces. Any vendor confident in their product will do it. The ones that will not have told you something a leaderboard never could. This also produces the eval set you need anyway if you end up building.

    Cost-aware evaluation matters as much as accuracy. Princeton's HAL project ran 21,730 rollouts across nine models and nine benchmarks for roughly $40,000 and released 2.5 billion tokens of logs. Its headline finding is that higher reasoning effort reduced accuracy in the majority of runs, and log inspection caught agents searching HuggingFace for the benchmark rather than solving it. Its live board makes the procurement point bluntly: agents can be a hundred times more expensive while only 1% better. Ask every vendor for accuracy and cost per task together, or you are comparing on one axis of a two-axis decision.

    Red Flags When Evaluating an Agent Vendor

    Send these before the demo, not after the pilot. Each one has appeared in a real evaluation we have run.

    Red flagWhy it matters
    “Autonomous agent” with no trace to showAsk to see a trace of a real multi-step run with tool calls and failures. A wrapper cannot produce one. This is the fastest agent-washing test there is.
    Benchmark scores with no harness namedScaffold and attempt budget move agent benchmark scores by tens of points. A score without the harness, the attempt budget, and the date is marketing.
    Audit logs you can view but not exportYou cannot answer a customer security review, a regulator, or your own incident postmortem from a dashboard you do not control.
    No model version pinningBehavior changes without a change-control record and you cannot reproduce last quarter's output. Ask what happens on the day the provider deprecates a model.
    Pricing per seat for a workflow that runs unattendedPer-seat pricing on an automation is a bet against your own volume growth. Model it at three times expected usage before signing.
    No answer on subprocessors or data retentionEvery vendor in the path is part of your data flow. “That's proprietary” is a decision, and the decision is no.
    “We always run the latest model”A silent upgrade is an unreviewed change to a non-deterministic system. Insist on pinned versions, advance notice, a changelog and a rollback path.
    Evaluation offered as a dashboard, not a datasetIf you cannot export your eval cases and run them yourself, you cannot switch vendors and you cannot verify their claims.
    Prompt injection described as “handled”It is not solved, and a vendor claiming otherwise has told you they do not understand the threat. The right answer is defense in depth and reduced blast radius.
    No named path off the platformAsk what you take with you: prompts, tool definitions, eval sets, traces, and fine-tuned artifacts. If the answer is “your data”, that is not the same thing.
    Roadmap-dependent capability in the demoTwo flagship consumer browser agents shut down in twelve months. Buy what exists today, not what is promised for next quarter.

    Frenchy Digital red-flag list for AI agent vendor evaluation, 2026.

    Ask for one artifact rather than a document set: a trace of a real multi-step run, including at least one failed tool call and how the system recovered. It answers more questions than a security questionnaire, because a vendor that cannot produce it does not have the system, whatever the questionnaire says.

    Frenchy Digital buyer's principle

    The Straight Recommendation: Buy the Commodity, Build the Seam

    Here is the recommendation without hedging. For nearly every organization we have scoped this for, the right answer is hybrid, and the split falls in a consistent place.

    • Buy the model: Frontier capability is repricing and improving faster than you can amortize a self-hosted alternative, and no long-context premium on one provider versus a doubling above 200k tokens on another is exactly the kind of difference you want to arbitrage rather than absorb. Keep the provider swappable at the interface, and re-baseline costs with a token counter every time a model changes.
    • Buy the tracing and evaluation platform: This is solved commodity infrastructure at $29 to $249 per month for most teams. Prefer a tool with an OSI licence and a genuine self-host path so your exit is technical rather than commercial — that distinction is why Langfuse under MIT and Helicone under Apache-2.0 read differently from tools whose self-host is gated behind an enterprise contract.
    • Buy connectors that already work: Verified by a read-write demo against a sandbox, not a logo grid. If the vendor will not demo against your ugliest system in two weeks, the connector does not work the way you need it to.
    • Build the orchestration and tool layer: This is where your domain logic lives, where permission scoping is enforced, and where the answers to audit questions come from. It is also the layer that determines your switching cost. Building it is how you keep the option to change everything else.
    • Build the retrieval scoping and data layer: The Gartner finding that successful organizations invest up to four times more in data quality, governance and change management is the strongest signal in the 2026 evidence, and none of that investment is purchasable. See our RAG guide for the mechanics.
    • Build the evaluation set, always: Not the platform — the cases. Twenty to fifty real failures with clear pass and fail criteria are the most portable asset you will own. They survive model changes, vendor changes and framework changes, and they are the thing that lets you evaluate any of the above honestly.

    The sequencing matters as much as the split. Start with one workflow that has a measurable baseline you already record. Instrument the before-state before you build anything. Ship the narrow version, read the traces weekly, and let production failures — not a roadmap — decide the second workflow. Organizations that pilot four disconnected vendors in parallel get four partial answers and no compounding; the compliance, identity, evaluation and tracing substrate is largely a fixed cost, paid once, and the fourth agent inherits all of it.

    On architecture, resist the pull toward multi-agent by default. The published evidence cuts both ways and the reconciliation is fairly clear: multi-agent structures win on wide, shallow, independent-thread work like research and market scans, while single-agent structures win on deep, narrow, shared-state work like coding and long-form writing. Both camps agree the mechanism is context, not agent count — and the 15× token multiplier means the wrong choice is expensive as well as wrong. Our multi-agent architecture guide covers the trade-off in depth.

    What This Costs With Frenchy Digital

    For completeness, and stated honestly: we are the partner option in this decision, which means the 67%-versus-33% finding above is a number we benefit from. Read it with that in mind, and read our cost bands as what a senior-led build actually costs rather than as an argument.

    EngagementRangeTimelineTypical scope
    Discovery + architecture review$9k–$22k2–4 weeksBuild-versus-buy recommendation per layer, twelve-month cost model, vendor shortlist assessment, eval-set outline
    Single production agent (one workflow, evals, observability)$30k–$80k5–10 weeksOne workflow end to end, tool layer with deny-by-default, eval suite in CI, tracing, cost guardrails
    Multi-workflow agent platform with integrations$80k–$200k10–18 weeksSeveral workflows, systems-of-record integration, retrieval layer, shared orchestration, routing across models
    Enterprise / regulated build (SOC 2 posture, HITL, audit logging)$200k–$450k+16–26 weeksMulti-tenant isolation, append-only audit pipeline, human-in-the-loop instrumentation, documentation package

    Frenchy Digital cost bands for AI agent engagements, 2026.

    Senior-led delivery runs $150 to $225 per hour, and ongoing retainers run $2,500 to $9,500 per month covering model and dependency upgrades, evaluation expansion, incident response, and a quarterly technical review. Every engagement carries a 30-day post-launch warranty, and you receive a written scope with a fixed-price phased proposal within 5 business days of the discovery call.

    Included at every tier: the build-versus-buy assessment per layer, an evaluation set drawn from your real failures, tracing and cost guardrails from day one, a tool layer with deny-by-default argument handling, and full source-code and IP ownership transferred to you at delivery. Frenchy Digital is a senior-led Black-owned Los Angeles agency and we do not build lock-in — if the honest answer for a layer is buy, that is what the proposal will say.

    For comparison, the build-versus-buy economics you will find circulating — a simple single-purpose agent at $15,000 to $80,000, a production multi-agent system with integrations, evaluation, monitoring and compliance at $250,000 to $600,000 and up, maintenance at $100,000 to $500,000 a year, break-even against licensing at eighteen to twenty-four months — come from vendor blogs with no published methodology or sample. They are in the right neighborhood, which is the most that can be said for them. Do not put them in a business case as data.

    Limitations and Honest Failure Modes

    This article audits other people's evidence, so it owes you the same treatment of its own.

    • The best available success statistics are analyst-firm surveys: The Gartner April 2026 I&O survey is the most defensible figure in this space, and its full methodology is not published. It is scoped to infrastructure and operations rather than to all enterprise AI, and analyst-firm surveys carry a commercial relationship with the market they measure. It is the best number available. It is not a controlled study.
    • The 67%-versus-33% finding is self-reported and confounded: It comes from the same document whose headline figure this article spends 1,200 words qualifying. Organizations that engage external partners plausibly differ from those that do not in scope clarity, sponsorship and budget — any of which could produce the gap. It is a directional signal about project structure, not a proven causal effect of buying.
    • The token multipliers are vendor-published and directional: The 4× and 15× figures come from a single Anthropic engineering post. No independent, non-vendor case study with real token-per-task numbers exists that we could find. They are the best anchors available and they are the right order of magnitude; they are not measurements of your workload.
    • Prices move monthly and one is already dated: Every price here was verified on August 10, 2026 against provider pricing pages. Claude Sonnet 5's introductory rate expires on August 31, 2026 and rises by 50%. Anything you model on today's prices needs a re-run schedule, not a spreadsheet.
    • Prompt injection is not solved and nothing here fixes it: Every architectural recommendation in this article reduces blast radius. None of them make an agent safe against a determined injection. If your workload cannot tolerate a wrong action that is unrecoverable, the correct answer may be not to give the agent that tool at all.
    • Benchmarks are improving faster than this snapshot: Benchmark integrity work moved substantially between late 2025 and mid-2026 — retirements, decontaminated task sets, patched reward-hacking vectors, cost-controlled leaderboards. The specific numbers here will age. The structural lesson — that no independent body validates vendor agent benchmarks — will not.
    • The framework assumes you can specify the workflow: Every dimension in the decision table presumes a workflow with a definable outcome. A large share of failed agent projects fail before that point, on a problem nobody could state precisely enough to evaluate. If your team cannot write twenty pass-or-fail cases, the build-versus-buy question is premature.

    None of this argues against building, and none of it argues for buying. It argues for making the decision per layer, on evidence you have checked, with a cost model that includes the parts nobody quotes. The organizations getting value from agents in 2026 are not the ones that picked correctly at the outset. They are the ones that instrumented the workflow well enough to find out.

    Deciding Whether to Build or Buy?

    Book a free 60-minute discovery call with Frenchy Digital — a senior-led Black-owned LA agency. You leave with a build-versus-buy recommendation per layer, a twelve-month cost model, and a fixed-price phased proposal within 5 business days. Call +1 (424) 272-5601.

    Deciding Whether to Build or Buy?

    Book a free 60-minute discovery call. You leave with a build-versus-buy recommendation per layer, a twelve-month cost model, and a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Unit 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    1. 1Gartner — AI Projects in I&O Stall Ahead of Meaningful ROI Returns (Apr 7, 2026)
    2. 2Gartner — Successful AI Initiatives Invest Up to Four Times More in Data Foundations (Apr 16, 2026)
    3. 3Gartner — Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (Jun 25, 2025)
    4. 4Gartner — 45% of Managers Report AI Has Lived Up to Expectations (Mar 4, 2026)
    5. 5Deloitte — State of AI in the Enterprise 2026 (N=3,235, 24 countries)
    6. 6Stanford HAI — 2026 AI Index Report
    7. 7MIT NANDA — The GenAI Divide: State of AI in Business 2025 (v0.1 preliminary findings, PDF mirror)
    8. 8MIT Media Lab — NANDA group overview and publications
    9. 9Arnon Shimoni — MIT's 95% AI Failure Rate Is Wrong (Aug 27, 2025)
    10. 10CIO Dive — S&P Global Market Intelligence on AI project abandonment (Mar 14, 2025)
    11. 11Anthropic — How We Built Our Multi-Agent Research System (token multipliers)
    12. 12Anthropic — Claude API Pricing
    13. 13OpenAI — API Pricing
    14. 14Google — Gemini API Pricing
    15. 15Anthropic — Demystifying Evals for AI Agents (Jan 9, 2026)
    16. 16OpenAI — Why We No Longer Evaluate SWE-bench Verified (Feb 2026)
    17. 17SWE-rebench — fresh, decontaminated software engineering tasks
    18. 18UC Berkeley RDI — Trustworthy Benchmarks: reward hacking agent evaluations
    19. 19Princeton HAL — cost-controlled agent leaderboards
    20. 20METR — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
    21. 21METR — Uplift study update (Feb 24, 2026): the 2025 result is labelled historical
    22. 22LangChain — State of Agent Engineering (n=1,340)
    23. 23Menlo Ventures — 2025: The State of Generative AI in the Enterprise
    24. 24OWASP — Top 10 for Large Language Model Applications
    25. 25NIST — AI Risk Management Framework
    26. 26FinOps Foundation — FinOps for AI
    Chris Machetto - CEO & Founder of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2019 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.