The Question on the Call
"I found three top 10 lists of AI agent development companies. Each one had a different company at number one, and each number one wrote the list. Who do I actually call?"
That is a paraphrase of a question buyers ask, and it's a fair one. I went and looked. Search for a custom AI agent development company and the first page is full of rankings published by Master of Code, RTS Labs, Neurons Lab and others. They are vendor-authored lists. The author tends to do well in its own ranking.
So here is my claim, with the date on it. On September 28, 2026 I scored ten third-party AI agent development companies on six things a buyer can check from a public page in about ten minutes each. 10Clouds came first with 8 of 12. The average was 4.5.
To be clear about the conflict up front: I run Frenchy Digital, which also builds agents. It is not in the scored table. It gets its own disclosed section at the bottom, where I score it on the same rubric (6 of 12, which would be fourth) and explain why I still pick it for one specific kind of buyer, and who should pick someone else.
If you have not decided yet whether you want an agency at all, that decision comes first, and it's covered in freelancer vs agency vs in-house for AI agent work. This article assumes you want a company, and you want a shortlist you can defend to whoever signs the check.
Why Most Top 10 Lists Fail
Most rankings of AI agent development companies fail for one of two reasons: the author is a contestant, or the scores rest on claims nobody outside the company can check.
The first problem is obvious once you notice it. The second is subtler. A list that credits a firm with some accuracy percentage or a faster processing time is repeating the firm's own marketing. There is no neutral benchmark of agent builders. Nobody runs the same workflow through ten agencies and publishes the results.
The wrong model is that a ranking measures who builds the best agents. The real model is that a ranking can only measure what companies disclose. Those are different things, and pretending otherwise is how you end up shortlisting on copywriting.
This matters more than it used to, because regulators have started reading AI claims closely. On January 14, 2025 the SEC issued a settled order against Presto Automation, which sold AI voice ordering for drive-thrus. The SEC found Presto had not disclosed that some deployed units ran on a third party's speech technology, and that while it claimed the product eliminated human order-taking, "the vast majority" of orders needed human intervention. Presto consented without admitting or denying the findings. No civil penalty was imposed.
Presto was not an agency. But the lesson transfers directly to buying from one: ask whose model runs the thing, and how often a human is quietly finishing the job.
Gartner put a name on the wider pattern. In a June 25, 2025 press releaseit described "agent washing", meaning assistants, RPA and chatbots rebranded as agentic without much agentic capability, and estimated that only about 130 of the thousands of vendors claiming agentic AI were real. That is Gartner's estimate, not an audited count, and I could only read the release through secondary coverage because the page blocks automated readers. Still, the direction is clear.
The FTC has been more concrete. It announced Operation AI Comply on September 25, 2024. The DoNotPay order it finalized in February 2025 required $193,000 in relief and barred claims of lawyer-level performance without evidence. The Workado order of August 28, 2025 requires competent and reliable evidence behind accuracy claims, with four years of compliance reporting.
So why would a ranking repeat an accuracy number the FTC would want evidence for? It shouldn't. That's the rule this one follows.
How I Scored Them
Six attributes, 0 to 2 points each, twelve maximum, all checked on September 28, 2026 from pages anyone can open.
- Pricing transparency: 2 if the company publishes figures on its own site; 1 if only a third-party rate band exists; 0 if neither.
- Partner status: 2 if the company appears in a platform vendor's own partner directory; 1 if the partnership is claimed only on the company's site; 0 if none.
- Security attestation: 2 if SOC 2 Type II or ISO 27001 is named on its site; 1 if SOC 2 is named without a type; 0 if none.
- Named AI case studies: 2 for three or more named clients; 1 for one or two; 0 for none.
- IP terms public: 2 for an explicit source code and IP ownership transfer; 1 for IP protection language; 0 for nothing stated.
- Independent review volume: 2 for twenty or more Clutch reviews; 1 for one to nineteen; 0 if not found.
Ties break on the review score, then alphabetically. That is the whole rubric.
What I excluded, and why. Accuracy, ROI and outcome percentages, because every one on these sites is the company's own claim. Team size, because headcount says nothing about who works on your project. Acquired firms: LeewayHertz was left out because The Hackett Group announced an agreement to acquire it on September 16, 2024 (the price was reported at about $7.8 million). Fuselab Creative was left out because I found no agent-build evidence for it.
Where the data came from. Company homepages and about pages. The Claude service-partner directory, which I could fetch directly. Clutch, which blocks automated readers, so review counts come from search snippets and drift: 10Clouds showed 89 reviews on one page and 92 on another the same day. AWS Partner Finder and OpenAI's partner page only via search.
One more caveat on Clutch. Its own help pagedescribes hourly rate and minimum project size as fields vendors are prompted to fill in. The page doesn't say who verifies them. So a Clutch rate band scores 1, not 2.
How to re-check it.Open each company's homepage and about page, search the Claude directory by name, search Clutch for the profile, and look for a pricing page or FAQ. It takes about ten minutes per firm. If a company has published something since September 28, the score should change, and I'd want to know.
The Scoreboard and Evidence
10Clouds leads with 8 of 12. The spread below it is tight, and most of what separates the middle is one or two published facts.
| Rank | Company | Pricing | Partner | Security | Named cases | IP terms | Reviews | Total / 12 |
|---|---|---|---|---|---|---|---|---|
| 1 | 10Clouds | 1 | 2 | 0 | 2 | 1 | 2 | 8 |
| 2 | Azumo | 1 | 0 | 1 | 2 | 0 | 2 | 6 |
| 3 | InData Labs | 2 | 1 | 0 | 1 | 0 | 2 | 6 |
| 4 | Intuz | 0 | 1 | 0 | 2 | 0 | 2 | 5 |
| 5 | Vstorm | 0 | 1 | 0 | 1 | 0 | 2 | 4 |
| 6 | Markovate | 0 | 1 | 2 | 0 | 0 | 1 | 4 |
| 7 | Neurons Lab | 0 | 2 | 0 | 1 | 0 | 1 | 4 |
| 8 | Master of Code Global | 0 | 1 | 2 | 1 | 0 | 0 | 4 |
| 9 | Tribe AI | 0 | 1 | 2 | 0 | 0 | 0 | 3 |
| 10 | RTS Labs | 0 | 0 | 0 | 1 | 0 | 0 | 1 |
Do the arithmetic and a pattern shows up. The ten scores add to 45, so the average is 4.5 of 12, about 38%. The leader, at 8, discloses roughly 1.8x what the average firm does. Four firms sit on exactly 4, which means the middle of this table is decided by tiebreaks, not by any meaningful gap.
Now look at the columns instead of the rows. Pricing: seven of ten score zero. IP terms: nine of ten score zero, and the tenth scores 1 for protection language. Those are the two things a buyer most needs before signing, and they are the two things this market least likes to publish.
That's the finding I'd take away even if you ignore the ranks. You will almost certainly have to ask for price and IP terms in writing, whoever you shortlist.
Every score above traces to one of the entries below. Where a line says the company claims something, I saw the claim on its site, not an audit report or a contract.
| Company | HQ / founded | Pricing | Partner status | Security | Named AI cases | IP terms | Clutch reviews |
|---|---|---|---|---|---|---|---|
| 10Clouds | Warsaw; says 17 years | About $50 to $99/hr band on Clutch (snippet) | Claude directory, Select tier, listed as 10Clouds Financial Institutions | Not publicly disclosed | Four named: Polfund, Trust Stamp, Crescent, Qenta (company's claim) | IP protection and NDA language, not a transfer | About 89 to 92 (snippets) |
| Azumo | San Francisco; says 2016 | Project range about $4,200 to $70,000+ on Clutch (snippet) | Not publicly disclosed | Says SOC 2 Certified, type not stated | Four named: Omnicom, Centegix, Meta, Twitter (company's claim) | Not publicly disclosed | About 21 to 24 (snippets) |
| InData Labs | Miami; says 12+ years | Own FAQ: $15k to $50k PoC; $75k to $200k mid; $400k to $2M+ enterprise | Says AWS and Databricks partner | Not publicly disclosed | Two named: Wargaming, AsstrA (company's claim) | Not publicly disclosed | About 20 (snippet) |
| Intuz | San Ramon, CA and Ahmedabad | Not publicly disclosed | Says AWS Consulting Partner | Not publicly disclosed | Three named: CasePath, Careonix, French Florist (company's claim) | Not publicly disclosed | About 52 (snippet) |
| Vstorm | Wroclaw; says since 2017 | Not publicly disclosed | Says official Pydantic implementation partner | Not publicly disclosed | One named: Mixam (company's claim) | Not publicly disclosed | About 21 (snippet) |
| Markovate | Not publicly disclosed | Not publicly disclosed | Says AWS, Microsoft Solutions Partner, Google Cloud | Says ISO/IEC 27001:2022 and ISO 9001:2015 | None named (industry labels only) | Not publicly disclosed | About 12 (snippet) |
| Neurons Lab | London and Singapore | Not publicly disclosed | AWS Partner Finder listing; AWS Agentic AI competency announced June 2, 2026 | Not publicly disclosed | Two named: HSBC, Visa (company's claim) | Not publicly disclosed | About 5 (snippet) |
| Master of Code Global | Redwood City, CA; says 2004 | Not publicly disclosed (fixed-budget pilots, no figures) | Says Claude Partner Network approval; not in Claude directory on the date checked | Says ISO 27001 | One named: Zipify (company's claim) | Not publicly disclosed | Count not found; own site says 4.7/5 |
| Tribe AI | New York; founded 2019 per PitchBook | Not publicly disclosed | OpenAI partner page exists (not fetched); says Google Cloud | Says SOC 2 Type II | None named on homepage | Not publicly disclosed | Not found |
| RTS Labs | Glen Allen, VA; says 2010 | Not publicly disclosed | Not publicly disclosed | Not publicly disclosed | One named: Landstar (company's claim) | Not publicly disclosed | Not found |
A few cells deserve a note. The 10Clouds directory listing is for an entity called 10Clouds Financial Institutions, at the Select tier of the Claude Partner Network's Services Track. Anthropic's Services Track announcementdescribes Select, Preferred and Global Premier tiers, and Select is listed first, which suggests it is the entry tier. It still counts, because it's the platform vendor saying it, not the agency.
Master of Code Global says in a blog postthat it was approved to move forward in the Claude Partner Network. It did not appear in the Claude directory when I checked. That's not a contradiction, necessarily; onboarding takes time. It is why it scores 1 and 10Clouds scores 2.
Neurons Lab's AWS competency is dated. It announced the AWS AI Competency in the Agentic AI category on June 2, 2026, and AWS created those categories in November 2025. A Partner Finder listing exists, which I saw only through search.
And the security column is the one I'd push hardest on in a sales call. A company writing "ISO 27001" or "SOC 2 Type II" on its homepage is making a claim you can verify in one email: ask for the certificate or the report under NDA. I scored the claim. You should check the document.
The Ten, Profiled
Each profile says who the firm suits, what I could verify, and what I could not.None of this is a quality judgment; I haven't worked with any of them.
1. 10Clouds (8 of 12)
Suits: financial services teams that want a partner the model vendor itself lists, and buyers comfortable with a Central European delivery team.
Verifiable: based in Warsaw, and it says it has been operating for 17 years. It appears in the Claude service-partner directory as 10Clouds Financial Institutions, Select tier, and announced joining the Services Track. Its homepagenames four clients (Polfund, Trust Stamp, Crescent, Qenta) and speaks of a commitment to protecting clients' intellectual property, with NDAs. Clutch snippets show a band of about $50 to $99 an hour and roughly 90 reviews.
Not verifiable:not every named case is agent work; Polfund is the one that reads most like an agent build. The homepage's deployment count and its own Clutch rating are the company's claims, and I don't print them as fact. No SOC 2 or ISO was stated. The IP language is protection, not a transfer, so ask for an assignment clause.
2. Azumo (6 of 12)
Suits: US buyers who want a San Francisco address and a stated SOC 2 posture, with project sizes in the tens of thousands.
Verifiable: Azumo says it was founded in 2016 and gives a San Francisco address. It says it is SOC 2 Certified without naming the type, which scores 1. It names Omnicom, Centegix, Meta and Twitter as clients across generative AI, computer vision, semantic search and chatbot work. A Clutch snippet shows a project range of about $4,200 to $70,000 or more, and around 21 to 24 reviews.
Not verifiable: no partner program I could confirm, no IP terms, and no independent evidence for the named projects. Its site says pricing depends on scope, seniority and engagement model, which is true of everyone and tells you nothing. The SOC 2 type matters: Type II covers controls operating over a period, so ask which one.
3. InData Labs (6 of 12)
Suits: buyers who need a budget number before the first meeting, and data-heavy projects.
Verifiable: InData Labs is the only firm in the ten that publishes cost ranges on its own site. Its FAQ lists about $15,000 to $50,000 for a proof of concept or MVP, $75,000 to $200,000 for mid-complexity work, and $400,000 to $2 million or more for enterprise builds. It gives a Miami address, says it has 12+ years behind it, and names Wargaming and AsstrA. About 20 Clutch reviews.
Not verifiable:its AWS and Databricks partnerships are claimed on its site; I didn't find them in a vendor directory. No security attestation or IP terms stated. And the ranges are wide. The mid band runs $75,000 to $200,000, which is about 2.7x from bottom to top, so treat it as a sanity check rather than a quote.
4. Intuz (5 of 12)
Suits: buyers who want a US office with offshore delivery economics, and who weigh review volume heavily.
Verifiable: Intuz lists San Ramon, California and Ahmedabad. It says it is an AWS Consulting Partner and names CasePath, Careonix and French Florist. Clutch snippets show about 52 reviews, the second-highest count in the ten.
Not verifiable:I've seen 2008 given as its founding year, but I couldn't re-find it on the homepage, so I won't state it. No pricing, security attestation or IP terms were published. The site also carries outcome figures like fraud reduction percentages and a count of systems in production. Those are its own claims.
5. Vstorm (4 of 12)
Suits: Python teams already using Pydantic, and buyers who want a specialist boutique rather than a generalist shop.
Verifiable: Vstorm is based in Wroclaw, says it has operated since 2017 with 40+ people, and calls itself an official Pydantic implementation partner and the first AI consultancy to join the Agentic AI Foundation. It names one client, Mixam, through a quote. About 21 Clutch reviews.
Not verifiable:both partner claims come from Vstorm's own site; I found no independent confirmation. Its site shows accuracy and success-rate figures from projects. I refuse to print them as fact, for the reason given in the refusal section below. No pricing, security or IP statements.
The bottom half isn't worse at building. It publishes less, and in a couple of cases publishes the things enterprise buyers ask for while skipping the things small buyers ask for.
6. Markovate (4 of 12)
Suits: buyers whose procurement checklist starts with ISO certification.
Verifiable: Markovate names ISO 9001:2015 and ISO/IEC 27001:2022 on its site, and lists AWS, Microsoft Solutions Partner, Google Cloud and Stripe relationships. About 12 Clutch reviews.
Not verifiable:founding year and headquarters weren't stated (there's a US phone number). Case studies are labeled by industry, not client, and carry percentage outcomes I won't repeat. No pricing or IP terms. It's the clearest example in the list of a firm scoring well on security and poorly on everything a small buyer would look at first.
7. Neurons Lab (4 of 12)
Suits: AWS-first organizations, particularly in financial services, in the UK or Southeast Asia.
Verifiable: offices in London and Singapore. The AWS Agentic AI competency announced June 2, 2026, and a Partner Finder listing, give it the full 2 on partner status. It names HSBC and Visa. About 5 Clutch reviews.
Not verifiable: founding year, pricing, security attestation and IP terms. The review count is the thinnest of anyone who has one, so the independent signal is small. Ask for references directly.
8. Master of Code Global (4 of 12)
Suits: larger organizations that want a firm that says it has operated since 2004, with multi-cloud and Salesforce relationships.
Verifiable: its About page says it was founded in 2004, is based in Redwood City with offices in Winnipeg and Ukraine, and has 150+ staff. It names ISO 27001 and lists Google Cloud, AWS, Salesforce and Amazon Connect partnerships. It names Zipify as a client. It offers fixed-budget 30-day AI pilots, without figures.
Not verifiable:its own site says 4.7 out of 5 on Clutch but gives no count, and I couldn't find one, so review volume scores 0. The Claude Partner Network status is its claim; the directory didn't show it on the day. Worth noting that Master of Code also publishes one of the top 10 lists I mentioned at the start. That's normal in this market. It's also why this list exists.
9. Tribe AI (3 of 12)
Suits: enterprises that care about a stated SOC 2 Type II attestation and offices in New York, San Francisco or Lisbon.
Verifiable: Tribe AI says it is SOC 2 Type II certified and Microsoft SSPA compliant, lists offices in New York, San Francisco and Lisbon, and announced this month that the HipAI team joined it. PitchBook lists it as founded in 2019 in New York. An OpenAI partner page for it exists, though it blocked my fetch.
Not verifiable: no named case studies on the homepage, no pricing, no IP terms, and no Clutch profile I could find. Tribe is independent; it acquired a team, it was not acquired, which is why it stays in.
10. RTS Labs (1 of 12)
Suits: US mid-market firms that specifically want US-based staff from a firm that says it has operated since 2010.
Verifiable: RTS Labs says it was founded in 2010 in Glen Allen, Virginia, with US-based staff. It names a Landstar agent portal and a Renewal by Andersen testimonial.
Not verifiable:the Landstar agent portal may be a portal for Landstar's human agents rather than an AI agent; the page doesn't say. No partner, security, pricing or IP statements, and no Clutch profile found. A score of 1 is the streetlight effect at its most extreme: this is a firm that tells you almost nothing publicly, which is not the same as having nothing to tell.
Picking by Buyer Type
The rank matters less than the fit. A 4 that matches your constraint beats an 8 that doesn't.
Think of the scoreboard as a filter, not a verdict. Decide your one non-negotiable first (a security attestation, a published budget, a specific cloud, onshore staff) and see who survives it. Then compare the survivors on everything else.
| If you are | Start with | Why | Watch for |
|---|---|---|---|
| A bank or fintech that wants a Claude-certified partner | 10Clouds | Only firm verified in the Claude partner directory, and the listing is its financial institutions entity | Rates are a Clutch band, not a quote; IP terms are protection language |
| A buyer who needs a budget before the first call | InData Labs | The only ranked firm publishing its own cost ranges | Ranges are wide; the mid band spans about 2.7x |
| A security-reviewed enterprise buyer | Markovate, Master of Code Global or Tribe AI | Each names ISO 27001 or SOC 2 Type II on its own site | Ask for the report or certificate itself; we saw none |
| An AWS-first organization | Neurons Lab | AWS Agentic AI competency announced, and a Partner Finder listing | Few public reviews (about 5) |
| A team building on Pydantic in Python | Vstorm | Says it is an official Pydantic implementation partner | Partner claim is its own; outcome numbers are vendor claims |
| A US mid-market firm wanting onshore staff | RTS Labs or Intuz | RTS says US-based staff; Intuz has a San Ramon office and about 52 reviews | Neither publishes pricing |
| A small business that wants one agent live soon | Read the disclosed pick below | A published fixed package price is rare in this list | Package scope is fixed by design |
Two of these rows deserve a longer explanation, because they're where I see buyers make the most expensive mistake.
The security-reviewed buyer.If your procurement team will send a vendor security questionnaire, only three firms in the ten name an attestation that scores 2. That shortlists itself. But don't confuse a certified company with a secure agent. SOC 2 describes the vendor's controls. It says nothing about whether the agent it builds for you can be talked into emailing your customer list to a stranger. You need both conversations.
The budget-first buyer. Suppose you have about $60,000 approved and need to know whether that buys anything. InData Labs' published ranges put $60,000 just above its proof-of-concept band ($15,000 to $50,000) and below its mid-complexity floor of $75,000. That's a useful answer before any call: your budget is roughly 0.8x the mid band's floor, so you're buying a pilot, not a platform. For the full breakdown of what drives that number, including model tokens, hosting and maintenance after launch, see what an AI agent really costs to build and run.
And if the decision is how you'll pay (monthly retainer, fixed project or a packaged product), that structure changes your risk more than the vendor does. I worked through the trade-offs in how AI automation agencies price retainers vs projects.
If geography matters to you, the local rankings apply the same idea to one city. The Los Angeles version is the top AI agent development agencies in LA, and the Los Angeles AI agent developer guide covers independent developers as well as firms.
A Worked Shortlist
A worked example makes the filter concrete. This one is a scenario, not a client.
Consider a regional insurance broker with about 40 staff. It wants an agent that reads inbound policy change requests from email, pulls the policy from its management system, drafts the change and queues it for a licensed agent to approve. Its IT lead says any vendor will get a security questionnaire. Its CFO says the approved budget is roughly $90,000 for the first phase.
Step one is the non-negotiable. The security questionnaire is it, so the first cut keeps the three firms that name an attestation scoring 2 (Markovate, Master of Code Global, Tribe AI) and, if the broker is willing to ask which SOC 2 type it holds, Azumo. Four firms survive. Six drop out, not because they are worse, but because they didn't publish the one thing this buyer needs first.
Step two is the budget sanity check. None of those four publishes a price, so the only public yardstick is InData Labs' mid-complexity band of $75,000 to $200,000. The broker's $90,000 sits near the bottom of it: about 1.2x the floor and 0.45x the ceiling. That tells the broker to expect a scoped first phase, one workflow and one system, not the whole back office.
Step three is the four emails from the list at the end of this article: phased fixed-price proposal, named model providers, a written IP assignment, and the security report under NDA. How fast and how completely a firm answers them is the best free signal you'll get about how it runs a project.
Step four is the boundary. The agent may read the email, look up the policy and draft the change. A licensed person must approve it before anything is written back. Put that sentence in the statement of work. It is also, not by coincidence, the design that limits the damage if a malicious email ever talks the agent into something.
Notice what the scoreboard rank did in all this: almost nothing. The number one firm didn't survive step one, because this buyer's constraint happened to be the one column where it scores zero. That's the point of scoring attributes separately. The total is a summary; the columns are the tool.
The downside of this approach is that it rewards firms that publish, and some very good builders publish very little. If your favorite firm got cut at step one, ask it directly. A single email can turn a zero into a two, and the ranking would be the first thing I'd update.
What the Scores Miss
In my view, three things matter more to an agent project than anything in a public scorecard: who actually does the work, how the scope is phased, and what happens after launch.
Who does the work. A firm with 150 people and a firm with 15 can put exactly the same two engineers on your project. Ask for the names of the people who will write the code, and ask to meet them before you sign. If the senior person on the sales call disappears after kickoff, you bought a brand and received a team you never interviewed.
How the scope is phased.Agent projects are unusually good at revealing surprises late: the system with no API, the permission nobody can grant, the edge case that turns out to be a third of the volume. A phased proposal with a price and an exit at each phase turns those surprises into a decision point instead of an overrun. Think of it like buying a house with an inspection contingency. You still want the house; you just don't want to discover the foundation after closing.
What happens after launch. Models get deprecated, prompts drift, the systems the agent talks to change their fields. Ask what the warranty covers, for how long, and what a retainer costs after it ends. An agent nobody maintains can degrade quietly, and the first person to notice may well be a customer.
I'd weigh these three above anything in the table. The table gets you to a shortlist. These questions get you to a signature.
Numbers I Refuse to Print
Several figures showed up again and again while I built this list, and none of them should travel.
The first group is vendor outcome numbers. Across the ten sites I saw accuracy jumps, success rates, percentage reductions in fraud or processing time, deployment counts, and an NPS score. Every one is published by the company about its own work, with no method, no denominator and nobody outside the firm checking. The Workado order is the FTC saying out loud that accuracy claims need competent and reliable evidence. I don't have that evidence for any of these, so they aren't in the scores, the table or the profiles.
The second group is market-size figures. A search summary of vendor lists offered a global AI agent market size for 2026 to one decimal place. I couldn't trace it to a primary source. Not printed.
The third group is the failure statistics every sales deck opens with: that 95% of generative AI pilots fail, that 85% of AI projects fail, that 87% never reach production. I refuse all three: none traces to a primary study that measured agent builds, so none belongs in a ranking of agent builders.
The Gartner release quoted above also contains a prediction that over 40% of agentic AI projects will be canceled by the end of 2027. It gets repeated as if it were a cancellation rate. It's a forecast, not a count, so I refuse to use it as a statistic; there is no primary source measuring cancellations.
Finally, a vendor list claimed that companies with six or more months of production deployments "consistently outperform". Outperform at what, measured how? No source. Not printed.
If a proposal you receive leans on any of these, that tells you how the rest of it was sourced.
Red Flags in a Proposal
Once you have two or three proposals, these are the lines I'd look for first. The full vetting checklist, with the questions to ask on the call, is in ten red flags when hiring an AI agent developer. This is the short version.
- An accuracy or ROI number with no method: Ask how it was measured, on whose data and by whom. If the answer is our past clients, it's marketing.
- No named model or provider: The Presto order is the reason. You want the model, the speech provider if any, and where the data goes, in writing.
- No clause assigning code and IP to you: Nine of ten ranked firms publish nothing on IP. Get a written assignment in the contract before the first sprint.
- Security claimed, report withheld: A certificate or SOC 2 report can be shared under NDA. A badge on a homepage cannot be audited.
- Prompt injection described as solved: It isn't. A credible builder talks about limiting what the agent can do when an injection succeeds.
- Open-ended time and materials with no phases: You want phases with a price and an exit at each, so a failed pilot costs a pilot, not the whole budget.
- No human in the loop for irreversible actions: Refunds, payments, deletions and outbound messages to customers should hit a confirmation step.
- A statistic from the refusal list on slide one: The failure-rate numbers above are the tell. A deck sourced that way tends to be sourced that way throughout.
Air AI is the cautionary tale for the last category. The FTC sued it in August 2025, alleging deceptive claims including that its conversational AI could replace staff. In March 2026 the FTC announced a proposed settlement that would ban it from marketing business opportunities, with an $18 million judgment largely suspended and a $50,000 payment. The order was proposed, not shown as entered, when I checked. Either way, replaces your staff is a sentence to read slowly.
The Security Question
The most useful question to ask any agent developer is what happens when the agent is tricked, because at some point it will be.
The UK National Cyber Security Centre wrote in December 2025that prompt injection may never be fully mitigated the way SQL injection was, because language models don't separate instructions from data. That's the right frame. You can't patch it out. You can only shrink what an attacker gains when it works.
The OWASP LLM Top 10 for 2026, published in early August 2026, gives you the vocabulary. LLM01 is Prompt Injection. LLM03 is Excessive Agency, meaning an agent with more permissions or autonomy than its job needs. LLM10 is Improper Output Handling, meaning downstream systems trusting whatever the model produced.
Put those together and the design goal is blast-radius reduction. Scoped credentials, so the agent can read the calendar but not the payroll. Deterministic code, not the model, for anything that commits money or data. A human confirmation before irreversible actions. Logs you can replay. That is the whole security conversation in four lines, and it applies to every firm on this list.
In a sales call, this sounds like one question: "If someone hides an instruction in an email your agent reads, what is the worst thing it could do, and what stops it?" A good builder answers with permissions and gates. A weak one answers with the model's safety training. The downside of asking is that it makes calls longer. It's still the cheapest security review you'll ever run.
Limitations
Here is what I chased and could not establish, so you can weigh the ranking accordingly.
- Build quality for any firm: No neutral benchmark exists and I commissioned no test builds. The ranking measures disclosure, not delivery.
- Any certificate or audit report: Every SOC 2 and ISO line is the company's own claim. I saw no reports and can't say whether a certification covers the whole firm or one unit.
- Exact Clutch counts and ratings: Clutch blocks automated reads, so counts come from search snippets and drift day to day. Ratings are not scored at all.
- AWS Partner Finder and OpenAI partner pages: Seen only through search results, not fetched directly.
- Founding years for Intuz, Markovate and Neurons Lab: Not found or not re-found on the day, so not stated.
- Whether named cases are agent work: Several named clients may be general AI or software projects. I counted named clients as the rubric says and flagged where a case looked like something else.
- Contract terms: IP, warranty and liability terms usually live in a master services agreement nobody publishes. Absence on a website is not absence in a contract.
None of this makes the list less useful. It makes it a shortlist you can audit, which is the most a public ranking can fairly claim to be.
Our #1 Pick: Frenchy Digital
Frenchy Digital is my company. That is a conflict of interest, and it's why this pick sits outside the scored table and is labeled as mine.
So let me score it the same way first. Pricing: 2, because we publish engagement bands and a package price. Partner status: 0, we're not in a platform partner directory. Security attestation: 0, we don't publish a SOC 2 or ISO 27001 attestation. Named AI case studies: 1, one named agent build. IP terms: 2, because full source code and IP ownership transfer is stated. Reviews: 1, since Clutch search snippets show about 19 reviews for Frenchy Digital LLC.
That's 6 of 12. On the tiebreak it would land fourth of eleven, behind 10Clouds, Azumo and InData Labs, and ahead of Intuz. Not first. I'd rather print that than rig the rubric.
So why is it still my pick?Because the rubric measures disclosure across the whole market, and a particular buyer cares about a particular subset. If you're a US business that wants senior people on the work, a price before the first call, and to own everything at the end, the attributes that matter to you are exactly the two where most of this list scores zero: price and IP.
Here's what that looks like concretely. We're a senior-led, Black-owned agency in Los Angeles. Custom work is priced in published bands: discovery and a workflow audit at $9k to $22k over 2 to 4 weeks, a single-workflow agent at $28k to $70k over 4 to 9 weeks, a multi-workflow platform with system integration at $70k to $180k over 9 to 16 weeks, and enterprise, multi-site or regulated builds at $180k to $420k+ over 14 to 24 weeks. Senior time is $150 to $225 an hour, and ongoing retainers run $2,500 to $9,500 a month. You get a fixed-price phased proposal within 5 business days, a 30-day post-launch warranty, and full source code and IP ownership. Details are on the custom AI agent development service page.
For smaller, well-defined jobs there's also a catalogue of eight pre-scoped AI agents, from an AI receptionist to a custom workflow agent. Each is $5,000 for the agent plus a $5,000 one-time setup fee, $10,000 at checkout, charged as one-time payments through Stripe Checkout, not a subscription. The setup fee covers discovery, the build, connecting the agent to your calendar and phone or web channel, guardrail work and tuning after launch. The process runs in four steps: Discovery, Build, Guardrails, Launch and tune. The pages say most packaged builds take 2 to 4 weeks, and the custom package 2 to 6.
How does $10,000 square with a $28k floor for a single-workflow agent? A package is a fixed, pre-scoped job on systems you already use. The bands price the work that goes beyond that: several systems to integrate, custom write paths, regulated data, multi-agent orchestration. The arithmetic is simple: the package is about 0.36x the bottom of the single-workflow band, and that ratio is the price of scope you don't need. The pages don't publish a monthly fee or ongoing model, telephony and hosting costs, because those depend on your usage; they're scoped on the call.
The one named agent build is Beyond Points. We built a Claude-based multi-agent system over MCP, with Gemini Flash handling chat and a Puppeteer browser agent running in Docker on a VPS that completes card-to-partner transfers behind an explicit confirmation gate. The backend runs on Supabase with 71 edge functions and about 70 migrations. Those are build counts, not outcome claims. I don't have a verified business result to print, so I won't invent one.
Three Things This Week
You can get from this page to a defensible shortlist in about a week.Here's the order I'd do it in.
- 1.Write down your one non-negotiable (a security report, a budget ceiling, a cloud, onshore staff) and use the buyer-type table to cut the list to two or three firms.
- 2.Email each of them the same four requests: a fixed-price phased proposal, the named model providers, a written IP assignment clause, and any security report under NDA. Note who answers all four.
- 3.On the call, ask what the worst thing the agent could do is if an email it reads contains a hidden instruction, and what stops it. Screenshot every page you relied on, dated, and keep it with the proposal.
Whoever answers those four emails cleanly and that one question with permissions rather than promises is probably your builder. Time to send the emails.
Want a Custom AI Agent You Own?
Book a discovery call with Frenchy Digital, a senior-led Black-owned Los Angeles agency. We map the workflow, the systems and the guardrails, and send a fixed-price phased proposal within 5 business days.
Shortlisting a Custom AI Agent Builder?
Book a discovery call. We map the workflow, the systems it touches and the guardrails, and send a fixed-price phased proposal within 5 business days.
1517 S Bentley Ave Apt 204, Los Angeles CA 90025
Frequently Asked Questions
Sources & References
- 110Clouds, homepage (checked September 28, 2026)↗
- 210Clouds, 10Clouds joins the Claude Partner Network↗
- 3Claude, Service partners directory (checked September 28, 2026)↗
- 4Anthropic, Services Track partner hub announcement↗
- 5Azumo, homepage (checked September 28, 2026)↗
- 6InData Labs, homepage and cost FAQ (checked September 28, 2026)↗
- 7Intuz, homepage (checked September 28, 2026)↗
- 8Vstorm, homepage (checked September 28, 2026)↗
- 9Markovate, homepage (checked September 28, 2026)↗
- 10Neurons Lab, AWS AI Competency in the Agentic AI category (June 2, 2026)↗
- 11AWS, AI Competency adds Agentic AI categories (November 2025)↗
- 12Master of Code Global, About us↗
- 13Master of Code Global, applied for the Claude Partner Network↗
- 14Tribe AI, homepage (checked September 28, 2026)↗
- 15RTS Labs, homepage (checked September 28, 2026)↗
- 16Clutch Help Center, why include hourly rate and minimum project size↗
- 17The Hackett Group, strategic acquisition of LeewayHertz↗
- 18SEC, In the Matter of Presto Automation Inc., Release No. 33-11352 (January 14, 2025)↗
- 19FTC, FTC Announces Crackdown on Deceptive AI Claims and Schemes (September 2024)↗
- 20FTC, FTC finalizes order with DoNotPay (February 2025)↗
- 21FTC, final order against Workado (August 28, 2025)↗
- 22FTC, Air AI proposed settlement (March 24, 2026)↗
- 23Gartner, press release on agentic AI and agent washing (June 25, 2025)↗
- 24OWASP GenAI Security Project, LLM Top 10 for 2026↗
- 25UK NCSC, Prompt injection is not SQL injection (December 8, 2025)↗

