Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    Hiring Checklist
    September 28, 2026
    30 min read

    10 Red Flags When Hiringan AI Agent Developer

    A vetting checklist you can run on a 30-minute call: ten warning signs in AI agent proposals, the question that exposes each one, what a good answer sounds like, and the regulatory and security record behind them.

    A business owner reviewing an AI agent development proposal with a checklist of red flags beside it
    LLM03
    Excessive Agency's position in the OWASP GenAI LLM Top 10 for 2026, the risk behind broad write access
    OWASP GenAI Security Project, LLM Top 10 2026
    60 days
    Minimum notice Anthropic commits to before retiring a publicly released model
    Anthropic, Model deprecations
    $193,000
    Monetary relief in the FTC's final order against DoNotPay over its AI lawyer claims
    FTC press release, February 2025
    30 days
    Default retention of OpenAI API abuse monitoring logs; Anthropic's API default is also 30 days
    OpenAI and Anthropic data retention documentation

    Key Takeaways

    • The ten red flags: guaranteed accuracy or ROI, no evaluation plan, prompt injection called solved, broad write access with no human approval, no monitoring or transcripts, no plan for model retirement, unclear code ownership, no data retention answer, a price that hides running costs, and regulatory hand-waving.
    • Every flag comes with one question to ask and a description of a good answer, so you can run the checklist on a 30-minute call and score each developer out of 20.
    • Performance guarantees have a regulatory record: the SEC's settled order against Presto Automation (January 2025), the FTC's DoNotPay final order ($193,000, February 2025), the Workado final order (August 2025) and a proposed, not yet entered, Air AI settlement (March 2026).
    • Prompt injection is not solved. The UK NCSC says it may never be mitigated the way SQL injection was, so a competent developer designs to limit damage: narrow permissions, human approval, deterministic checks.
    • Contractor-written software is generally not a work made for hire under 17 U.S.C. 101. Without a signed assignment, the developer owns the code you paid for.
    • Models retire on published schedules. Anthropic gives at least 60 days, OpenAI previews can get far less, so a proposal without a migration plan is a proposal with a hidden future invoice.
    • A developer who opens with a banned failure statistic such as the 95% pilot figure is showing you how they treat evidence.

    A Proposal That Promises 95%

    Suppose a founder forwards you a proposal and asks a simple question: is this good? It's for a customer support agent. It's nicely designed. And on page two, before anyone has looked at a single one of her support tickets, it promises the agent will resolve 95% of them.

    Ask her how they arrived at 95% and she won't know. Neither, I suspect, would they. That page is a composite, but every piece of it is the kind of thing buyers are shown.

    That number is the first of ten red flags in this article, and it's the easiest to spot. The others are quieter. They live in what a proposal leaves out: who owns the code, what the agent is allowed to do on its own, what happens when the model underneath it is retired next year, and what it costs to run once the build invoice is paid.

    So here is the checklist I'd hand her. Ten red flags, each with what it looks like in a proposal or on a call, why it matters (with a primary source where one exists), the one question that exposes it, and what a good answer sounds like. At the end there's a scoring sheet so you can compare two or three developers on the same 20 points.

    The wrong model most buyers bring to this is that vetting an AI developer is like vetting any other developer: look at the portfolio, check references, compare day rates. That still matters. But an agent is software that makes decisions and takes actions, so the questions that separate good builders from bad ones are about evidence, permissions and the long run, not about stack or speed. Therefore the checklist below mostly ignores the portfolio.

    If you're earlier than this and still deciding whether you want a freelancer, an agency or an employee, start with how to hire an AI agent developer and come back with a shortlist. If you want names to put on that shortlist, our ranking of the top AI agent development companies scores ten firms only on attributes you can check on a public page.

    Disclosure. Frenchy Digital builds AI agents, so we are one of the developers you might run this checklist on. Near the end I answer the ten questions for our own work, including the places where our answer has limits. Run the same sheet on us that you run on everyone else.

    Why Vetting Feels Impossible

    Hiring an AI agent developer is a lemons market, and red flags are how you inspect the car.

    The idea comes from economics. George Akerlof's "market for lemons" describes what happens when sellers know the quality of what they're selling and buyers don't: good and bad products look identical from the outside, so buyers pay an average price, and the sellers with good products either leave or get undercut.

    Used cars were the original example. The fix in that market was never to get better at reading paint jobs. It was inspections, warranties and history reports: signals that are cheap for an honest seller to give and expensive for a dishonest one to fake.

    AI agent development in 2026 looks a lot like that. Every proposal says "custom", "secure" and "production ready". The demo always works, because demos are rehearsed. You can't read the code, and even if you could, the interesting failures only show up on real traffic.

    So you need questions that work like an inspection. A good question is one where an honest, competent developer can answer in two minutes, and a weak one has to either bluff or change the subject. Every question in this article is built to that spec.

    To be clear, a single red flag doesn't mean a developer is dishonest. Plenty of talented engineers have never had to think about copyright assignment or EU transparency rules. What you're scoring is whether they can take the question seriously once it's asked. The ones who get defensive tell you more than the ones who say "good question, let me come back to you in writing."

    One more piece of framing. The arithmetic of a bad hire is lopsided. A 30-minute call where you ask ten questions costs you half an hour. A single-workflow agent build at a senior shop runs somewhere in the tens of thousands of dollars and a couple of months. If the checklist saves you from one wrong hire in five, it has paid for itself many times over. That asymmetry is the whole case for being a little annoying on the first call.

    Numbers I Refuse to Print

    Before the ten flags, a bonus one: a developer who opens with a scary failure statistic.

    You've seen these in pitch decks. "95% of GenAI pilots fail." "85% of AI projects fail." "87% never reach production." Gartner says 40% of agentic projects will be cancelled. The pitch that follows is always the same: most people get this wrong, so hire us.

    I refuse to print any of them as fact, and here is why. When I tried to trace the 95%, 85% and 87% figures back to a primary source with a sample and a method I could check, I could not verify one, so none of them goes on this page as a finding. And the Gartner line is a prediction in a press release, not a measurement of anything that has happened.

    The same goes for vendor outcome figures: "98% accuracy", "70% faster", "40% cost reduction". When a seller publishes a number about their own product, it's marketing until someone independent reproduces it. There is no independent benchmark for custom business agents that I could find, and several agency sites I checked for the companies ranking publish outcome numbers about their own work that nobody outside has reproduced.

    Why is this a red flag and not just sloppy marketing? Because it shows you how the developer treats evidence. Someone who repeats an untraceable statistic to close a deal will also tell you their agent is "95% accurate" without a test set behind it. The habit is the same habit.

    The question to ask:"Where does that number come from?" A good developer either names a primary source with a sample and a method, or says they can't trace it and drops it. A weak one says "it's well known" or sends you a blog that cites another blog.

    Red Flags in the Promise

    The first two red flags are about what the developer claims the agent will do, and how they intend to prove it. In my view they are the easiest to catch, because they are usually written down.

    Red flag 1: A guaranteed accuracy or ROI number

    What it looks like.A percentage in the proposal before discovery: "resolves 90% of tickets", "books 3x more appointments", "pays for itself in 60 days". Sometimes it's softer: "our agents typically achieve". Either way, the number arrived before your data did.

    Why it matters.An agent's performance depends on your inputs: how messy your tickets are, how many edge cases your booking rules have, how often customers ask things no document answers. Nobody can know that number before testing on your examples. And regulators have a record here.

    The SEC's settled order against Presto Automation (January 14, 2025) concerned claims about its AI drive-thru ordering product. The SEC found that "the vast majority" of orders needed human intervention while the company was describing the product as removing the need for human order-taking. Presto consented without admitting or denying the findings, and no civil penalty was imposed. It's a securities case, not a consumer one, but it's the cleanest public record I know of the gap between an AI automation claim and what was actually happening on the ground.

    The FTC has been busier. Its Operation AI Comply sweep in September 2024 named DoNotPay, Ascend Ecom, Ecommerce Empire Builders, Rytr and FBA Machine. The DoNotPay final order (Commission vote 5-0 on January 16, 2025) required $193,000 in monetary relief, notice to subscribers from 2021 to 2023, and a ban on claiming the service performs like a lawyer without evidence. The Workado final order (August 28, 2025) concerned the accuracy of an AI content detector: accuracy claims now need competent and reliable evidence, and the company files compliance reports for four years.

    Then there's Air AI. The FTC sued in August 2025 over claims about earnings, growth and refunds, including that its conversational AI could replace human staff. On March 24, 2026 the FTC announced a proposed settlement: a stipulated order filed in federal court in Arizona with a ban on marketing business opportunities, an $18 million judgment that is largely suspended, and a $50,000 payment. Status precision matters here. As of September 28, 2026 the FTC case page shows no entered order, so it's proposed, not final. And it is not an "$18M fine", which is how you'll see it described; only $50,000 is being paid.

    None of these cases is about a custom agent developer. But they are the enforcement pattern your developer's promises will end up inside once you repeat them to your own customers.

    • The question to ask: What number will you commit to, how will we measure it, and on what data? If you are going to put a percentage in the contract, show me the test set it will be measured against.
    • What a good answer sounds like: I can't give you a percentage yet. What I can commit to is a test set built from about 100 to 300 of your real examples, a pass threshold we agree before launch, and a rule that the agent does not go live until it clears it. After launch we measure the same thing on live traffic and report it monthly.
    • What a weak answer sounds like: Our agents always hit 90% or better. Or, that is what our other clients see. Neither is evidence about your workflow, and the second one is a vendor outcome claim with no way for you to check it.

    Red flag 2: No evaluation plan or test set

    What it looks like.The proposal has phases called Discovery, Build, Testing and Launch, and "Testing" is one line with no detail. On the call, the developer says they'll "try lots of examples" or "iterate until it feels right".

    Why it matters. An agent built on a language model does not fail the way ordinary software fails. Ordinary code breaks loudly on the same input every time. A model can answer a question correctly nine times and wrongly the tenth, and a prompt change that fixes one case can quietly break three others. The only defence is a fixed set of test cases, rerun on every change, with a score you can compare over time.

    Think of it like a restaurant kitchen tasting every dish before it leaves the pass. You don't taste one plate on opening night and assume the rest are fine. You taste every time the recipe changes.

    NIST's Generative AI Profile (AI 600-1), published in July 2024 as a companion to its AI Risk Management Framework, treats measurement and ongoing monitoring as core risk-management activities rather than optional extras. I'm not saying your developer needs to map their process to NIST. I'm saying the federal government's own framework assumes you measure, and a developer who doesn't plan to is behind it.

    We walk through what a working test set and trace setup look like in AI agent evaluation and observability, including how many cases is enough to start with. The short version: fewer than you fear, but more than zero, and written down before the build starts.

    • The question to ask: Show me the test set you would build for this workflow. Who writes the cases, who decides the right answer for each, and what score does the agent need before it goes live?
    • What a good answer sounds like: We will pull real examples from your inbox or call logs, including the ugly ones, and have someone on your team mark the correct outcome for each. Every prompt or model change reruns the whole set. We agree the launch threshold with you, and the scores go in a report you can read.
    • What a weak answer sounds like: We test thoroughly. Or, the model is very good at this kind of task. Neither tells you what will be measured, or who decides what correct means.

    A third promise problem almost made the list on its own: a developer who can't name anything in your workflow that should not be automated. Every real process has steps where a mistake is expensive, rare or legally loaded (refunds above a threshold, anything medical, anything that commits you to a price). A developer who says the agent can do all of it has either not looked closely or is sizing the project to the budget rather than the risk. I fold it into flag 4, because the fix is the same boundary table.

    Red Flags in the Security Story

    Flags 3, 4 and 5 are about what happens when the agent is wrong, or is being manipulated, and whether anyone will notice. These are the ones that turn a bad month into a bad year.

    Red flag 3: They say prompt injection is solved

    What it looks like.You ask about security and hear "we use a guard model", "our system prompt tells it to ignore malicious instructions", or simply "that's handled". The implication is that an attacker can't make the agent misbehave.

    Why it matters.Prompt injection is when text the agent reads (an email, a web page, a customer message, a PDF) contains instructions that the model follows as if they came from you. It's number one on the OWASP list for a reason, and it is not solved.

    The UK National Cyber Security Centre put it plainly in a December 2025 post, "Prompt injection is not SQL injection". SQL injection was fixed by separating code from data. Language models don't have that separation: everything is text in the same window. The NCSC says prompt injection "may never be totally mitigated in the way that SQL injection attacks can be".

    OpenAI has said something similar about its own browsing agent. In a post on hardening ChatGPT Atlas against prompt injection, it described the problem as a long-term security challenge that is unlikely ever to be fully solved, closer to scams and social engineering than to a bug you patch once. If the company that trains the model says that, your developer's system prompt has not solved it.

    So the right frame is blast radius. You assume the agent will sometimes be fooled, and you design so that being fooled can't do much damage: it can't reach systems it doesn't need, can't take consequential actions without a check, and can't pass its output straight into something that executes it.

    • The question to ask: Suppose a customer email contains instructions telling the agent to forward all recent orders to an outside address. Walk me through what stops that from happening.
    • What a good answer sounds like: We assume some injections will get past the model, so the model is not the control. The agent has no tool that can email arbitrary addresses, its credentials can only read the one mailbox, and anything leaving the business goes through a fixed template or a person. Filters help, but they are not what we rely on.
    • What a weak answer sounds like: The model is trained to refuse that. Or, we have a guard model that catches it. Both treat a probabilistic filter as a lock.

    Red flag 4: Broad write access with no human approval

    What it looks like. The architecture diagram shows the agent connected to your CRM, your inbox, your payment system and your calendar with one arrow each, and nobody has asked what it can changein each one. On the call: "we'll give it an admin API key to keep things simple."

    Why it matters. The OWASP GenAI LLM Top 10 for 2026, published in August 2026, lists Excessive Agency as LLM03 and Improper Output Handling as LLM10. Excessive Agency is exactly this flag: an agent with more functions, permissions or autonomy than the task needs. Improper Output Handling is its sibling: model output passed into another system (a database query, a shell, an email) without validation. Earlier editions used different numbering, so if a developer cites the list, check which year they are reading. The official repository has the current ordering.

    OWASP also published a separate Top 10 for Agentic Applications in December 2025, which is worth handing to any developer who hasn't seen it.

    Combine flag 3 and flag 4 and you get the real risk. An agent that can be tricked (always true) and that can issue refunds, delete records or email anyone (a design choice) is an agent that will eventually do one of those things for someone who isn't you. The only variable you control is the second one.

    The fix is a boundary table, agreed in writing before the build: what the agent may do alone, what it may prepare but a human must approve, and what it must never touch. We cover the mechanics in human-in-the-loop design for AI agents and the credential side in AI agent sandboxing and credential scoping.

    The one agent case study we publish is a good illustration of the shape. In the Beyond Points build, a browser agent running in Docker on a VPS completes card-to-partner transfers, and it does so behind an explicit confirmation gate. The agent does the tedious navigation. A person confirms the move that matters.

    • The question to ask: Which actions can the agent take without a person approving them, and what credentials does it hold for each system it touches?
    • What a good answer sounds like: Here is the table. It can read your calendar and create tentative holds on its own. Refunds, price quotes and anything that deletes data are prepared for a person to approve. It has no access to payroll at all. Each tool gets its own scoped credential, not a shared admin key.
    • What a weak answer sounds like: It will be careful. Or, we will add approvals later if you want them. Approvals added after launch tend never to be added.

    Red flag 5: No monitoring or transcripts

    What it looks like.The proposal ends at "Launch". There's no mention of logs, transcripts, dashboards or alerts, and when you ask how you'll know if something goes wrong, the answer is "customers will tell you" or "we monitor uptime".

    Why it matters.Uptime tells you the agent is answering. It doesn't tell you what it's saying. An agent can be up 100% of the time and quietly giving the wrong refund policy to one customer in twenty, and you won't know until the complaints arrive, if they arrive at all.

    Monitoring for an agent means three things: a full transcript of each conversation, a trace of every tool call the agent made (what it looked up, what it changed), and a person whose job it is to look at a sample of both every week and respond to alerts. Without those, your evaluation plan from flag 2 stops the day the agent launches.

    This doesn't have to be expensive. Hosted tracing tools publish entry tiers in the tens of dollars a month; the cost is mostly the human hour a week to actually read the output. The question is whether it's in the plan at all.

    • The question to ask: After launch, what will I be able to see every morning, where does it live, and who is responsible for looking at it?
    • What a good answer sounds like: You get the full transcript log and a trace of every tool call, in a tool you can log into yourself. We set alerts on a few signals, like repeated handoffs or failed tool calls, and during the warranty period we review a sample every week with you. After that the review is either ours on a retainer or yours, and we write down which.
    • What a weak answer sounds like: We will have logs if you need them. Logs you have to ask for are logs nobody reads.

    Red Flags About Next Year

    Red flag 6 is the one almost nobody asks about, and it's the one that produces the surprise invoice.

    Red flag 6: No plan for model deprecation

    What it looks like.The proposal names a specific model ("built on the latest GPT" or "powered by Claude") and says nothing about what happens when that model is retired. On the call, the developer seems surprised that models retire at all.

    Why it matters. The model under your agent is rented, and the landlord publishes an eviction schedule. Anthropic's model deprecations page commits to at least 60 days' notice for publicly released models. Its recent history shows what that looks like in practice: Claude Sonnet 4 and Opus 4 were notified on April 14, 2026 and retired on June 15, 2026. Claude Opus 4.1 was notified on June 5, 2026 and retired on August 5, 2026. Do the arithmetic on that last one: 61 days from notice to shutdown. Two months.

    OpenAI's deprecations page tells a similar story. gpt-4-0613 and gpt-3.5-turbo-0125 were announced on April 22, 2026 for shutdown on October 23, 2026, about six months. The gpt-4o-realtime-preview models were announced on September 15, 2025 and shut down on May 7, 2026. And OpenAI notes that preview models can get much shorter notice, as little as a couple of weeks. If your voice agent is built on a preview model, your migration window may be measured in weeks, not months.

    Also worth knowing: if you access models through a cloud platform such as Bedrock or Vertex, those platforms set their own schedules, which Anthropic's page says explicitly.

    Here's why this is a red flag rather than a footnote. A new model is not a drop-in replacement. It phrases things differently, follows instructions differently and calls tools differently. Prompts tuned for one model drift on the next. The only safe migration is to switch, rerun the full test set from flag 2, fix what broke and ship. That's real work, and it recurs. If the proposal doesn't mention it, you'll meet it as an unplanned invoice with a two-month deadline.

    Think of it like a lease with a clause that says the landlord can move you to a different flat with 60 days' notice. You'd want to know who pays the movers before you sign.

    • The question to ask: Which exact model version will this run on, what is its retirement policy, and what happens, and who pays, when it is retired?
    • What a good answer sounds like: We pin a specific model version so nothing changes under you without notice. When a retirement is announced, we rerun the test set on the replacement, fix regressions and ship before the shutdown date. That work is covered under the retainer, or quoted as a fixed piece of work if you are not on one. The agent is written so the model can be swapped without a rebuild.
    • What a weak answer sounds like: The models only get better, so it will just improve. Newer models are often better on average and worse on your specific edge cases, which is why you rerun the test set.

    Red Flags in the Contract

    Flags 7, 8 and 9 hide in the paperwork: who owns what you paid for, where your data goes, and what you'll pay after the build.

    Red flag 7: You will not own the code

    What it looks like.The contract grants you "a licence to use the solution", or it says nothing about ownership at all. Sometimes the prompts, test sets or workflow configuration are carved out as the developer's "proprietary framework".

    Why it matters. Most buyers assume that if they paid for it, they own it. Under US copyright law that assumption is often wrong. The definition of a "work made for hire" in 17 U.S.C. 101 has two branches. The first covers work by an employee within the scope of employment. The second covers specially commissioned work, but only in nine listed categories (things like contributions to a collective work, parts of an audiovisual work, translations, compilations, tests and atlases) and only with a signed written agreement. Software is not one of the nine.

    So code written by an independent contractor or agency is generally not a work made for hire, and the author keeps the copyright unless it is transferred by a signed writing. The Copyright Office's Circular 30 walks through the same rules. I'm not a lawyer and this isn't legal advice; have yours read the clause. But the practical point is simple: you need an explicit assignment, not an assumption.

    For an agent, "the code" is more than code. It's the prompts, the tool definitions, the evaluation set, the configuration of the observability tool and the accounts everything runs in. If the developer owns the model provider account and the prompts, you can't move to someone else without starting over. That's the lock-in, and it's usually unintentional, which is why you have to ask.

    • The question to ask: On the last day of the project, who owns the code, the prompts, the test set and the accounts they run in? Can you show me the assignment clause?
    • What a good answer sounds like: Full source code and IP ownership transfer to you, in writing, including prompts, evaluation data and configuration. The model provider, hosting and tracing accounts are in your name from day one, and we are added as users. If we reuse an open-source library, it keeps its own licence and we list which ones.
    • What a weak answer sounds like: You will have full access. Access is not ownership, and a licence you can lose is not an asset you can sell with the business.

    Red flag 8: No answer on data retention

    What it looks like.You ask where your customers' data goes and how long it's kept. The answer is "OpenAI doesn't train on API data", and nothing else.

    Why it matters. Training is one question. Retention is another, and there are more stops than the model provider. At the provider level, the facts are published. OpenAI's API data documentation says API data is not used for training unless you opt in, abuse monitoring logs are kept for 30 days by default, and zero data retention or modified abuse monitoring require OpenAI's approval; some endpoints, including the assistants, threads and conversations endpoints, aren't eligible. Anthropic's privacy center says API inputs and outputs are deleted within 30 days by default, with exceptions: up to two years for policy enforcement, seven years for safety classification scores, five years for feedback, and whatever the law requires. Zero data retention is available by agreement for qualifying customers.

    Then there's everything the developer adds. Transcripts in a tracing tool. Documents in a vector database. Backups. A spreadsheet of test cases pulled from real customer emails. Each is a copy of your data with its own retention period, and unless you ask, there is no reason to assume anyone has written them down.

    If a developer claims SOC 2, ask which report and which period. The AICPA's Trust Services Criteria cover security, availability, processing integrity, confidentiality and privacy, and audit firms describe a Type 1 report as a point-in-time check of control design and a Type 2 as evidence the controls operated over a period. Ask any vendor, including us, for the report itself and read the period it covers. A developer without one can still earn your trust; what isn't fine is not knowing where the data sits.

    • The question to ask: List every place my customers' data will be stored or logged, who can see it, and how long it stays in each one.
    • What a good answer sounds like: A one-page map: the model provider and its retention default, the tracing tool and how long we keep transcripts, the vector store, backups and any test data. Where you need tighter terms, we tell you which provider programme to apply for and which endpoints to avoid.
    • What a weak answer sounds like: The provider does not train on your data, so you are fine. That answers a different question.

    Red flag 9: Pricing that hides running costs

    What it looks like.A build price and nothing else. Or a build price and a line that says "usage costs apply" with no estimate.

    Why it matters.An agent has a meter running. Every conversation spends model tokens; a voice agent adds telephony minutes; there's hosting, a tracing tool and somebody's time for maintenance. None of it is exotic, but it compounds, and the model choice alone moves it a lot.

    Let me do the arithmetic out loud with list prices checked September 28, 2026 on Anthropic's pricing page. Suppose a support conversation uses about 3,000 input tokens and 700 output tokens. On Claude Sonnet 5.5 at $2 per million input and $10 per million output, that's 3,000 × $2 ÷ 1,000,000 = $0.006, plus 700 × $10 ÷ 1,000,000 = $0.007. About $0.013 a conversation. At 10,000 conversations a month, about $130.

    Now the same conversation on Claude Fable 5.1 at $10 and $50: $0.03 plus $0.035, about $0.065 a conversation, or about $650 a month at the same volume. Same agent, same traffic, 5x the bill, purely from the model line in the config. OpenAI's pricing page shows the same spread between tiers: gpt-6-astra at $10 and $50 against gpt-6-sol at $2 and $10.

    Neither number is scary on its own. The red flag is a developer who hasn't done this sum for you, because it means they haven't thought about which model the task actually needs, and you'll find out on your first invoice. Voice agents add per-minute telephony and speech costs on top; our full breakdown of what an AI agent costs works through those lines.

    A related pattern: hourly billing with no fixed scope and no paid discovery. Hourly is fine for exploratory work. But a proposal that is only an hourly rate and an estimate transfers all of the estimation risk to you, and gives the developer no reason to write down the boundary table, the test set or the running costs. I don't count it as its own flag because a good hourly developer can still pass the other nine. It does make the other nine harder to check.

    • The question to ask: At our expected volume, what will this cost per month to run, line by line, and which model are you assuming?
    • What a good answer sounds like: A table: model tokens at an assumed volume and model, telephony or messaging, hosting, tracing, and maintenance hours, with the list prices and the date they were checked. Plus which lever brings it down, such as a smaller model for routing or caching the long system prompt.
    • What a weak answer sounds like: It is pennies per conversation. Maybe, but at which volume and on which model? Pennies times a hundred thousand is a line in the budget.

    Red Flag Ten: Regulatory Shrugs

    The last red flag is a developer who says AI is unregulated, or who quotes a deadline that moved months ago.

    Red flag 10: Regulatory hand-waving

    What it looks like."There's no AI law in the US." "The EU Act doesn't apply until 2027." "That's for your lawyers." Or the opposite: a scary slide about fines with the wrong dates on it.

    Why it matters. The rules are moving, and developers who stopped reading in 2025 are working from an old map. A few dates as of September 28, 2026:

    In the EU, the AI Act's Article 50 transparency obligations, including telling people they are interacting with an AI system, apply from August 2, 2026. That date did not move. What moved was the high-risk regime: the Digital Omnibus on AI, adopted as Regulation (EU) 2026/1744 (published July 24, 2026, in force July 27, 2026), pushed Annex III high-risk obligations to December 2, 2027 and product-embedded Annex I systems to August 2, 2028. Per Gibson Dunn's summary, the Article 50(2) marking duty gets a grace period for existing systems to December 2, 2026. So a developer who says "nothing applies until 2027" is half right, which is the dangerous kind of right if your agent talks to EU customers.

    In the US, Colorado is the example to know, because the earlier law, SB 24-205, has been replaced. SB 26-189, signed on May 14, 2026, repeals and reenacts it with a narrower regime effective January 1, 2027, with attorney general rules due by that date and enforcement by the attorney general only, no private right of action, per Holland & Knight. If you're a Colorado business using an agent in consequential decisions, "there's nothing to plan for" is wrong.

    And the enforcement cases in flag 1 are regulation too. The FTC doesn't need a new AI law to act on a deceptive claim; it used its existing authority in every case above.

    To be clear, I don't expect a developer to be your lawyer. I expect them to know which rules touch the thing they're building, name the dates correctly, and design for the obvious ones (an AI disclosure line, a human handoff, records of what the agent decided) without being asked.

    • The question to ask: Given where we operate and who our customers are, which rules apply to this agent, from what date, and what does the build do about each one?
    • What a good answer sounds like: A short list with dates. For example: EU customers mean an AI disclosure at the start of every conversation, already required since August 2, 2026. Your use is probably not Annex III high risk, but here is why, and counsel should confirm. Colorado's new regime starts January 1, 2027, so we log decisions in a form you can produce if asked.
    • What a weak answer sounds like: AI is basically unregulated. Or a deadline that has already moved, delivered with confidence.

    The Scoring Checklist

    Here are all ten on one sheet, scored 0, 1 or 2, for a maximum of 20. Print it, take it into the call, and score every developer on the same sheet.

    #Red flagThe question to ask0 points1 point2 points
    1Guaranteed accuracy or ROIWhat number will you commit to, and how will we measure it?Quotes a percentage before seeing your dataHedges, but has no measurement planCommits to a test set and a launch threshold, not an outcome
    2No evaluation planShow me the test set you would build for this workflow.Says they will test it manually as they goDescribes testing in general termsNames the cases, who labels them and the pass bar
    3Prompt injection called solvedWhat happens if a customer message tells the agent to ignore its rules?Says their prompt or guard model prevents itMentions filtering onlyExplains how permissions and approvals limit the damage
    4Broad write access, no approvalWhich actions can the agent take without a person approving them?Everything it has credentials forA vague listA written table: acts alone, needs review, never touches
    5No monitoring or transcriptsWhat will I see every morning, and who looks at it?Nothing, or only an uptime checkLogs you could requestTranscripts, tool-call traces and an alert owner
    6No model retirement planWhat happens when the model you chose is retired?Did not know models retireSays they will handle itPinned version, rerun of the test set, named cost
    7You will not own the codeWho owns the code, prompts and test set on the last day?Licence only, or silenceOwnership on final payment with exclusionsWritten assignment of everything built for you
    8No data retention answerWhere does my data go, and how long is it kept at each stop?The model provider handles thatKnows the provider policy onlyMaps provider, logs, vector store and backups
    9Pricing hides running costsWhat will this cost per month to run at our volume?No estimateA single guessed figureA line-by-line monthly estimate at an agreed volume
    10Regulatory hand-wavingWhich rules apply to this agent where we operate?None, AI is unregulatedNames a law without datesNames the rule, the date and what the build does about it

    How I'd read the total. This is my rule of thumb, not a standard. At 16 or above, the developer thinks about agents the way you want someone building yours to think. From 10 to 15, they may be fine, but get the gaps closed in the contract before signing. Below 10, keep looking, however good the demo was.

    And one override. A zero on flag 4 or flag 7 is disqualifying on its own for me, whatever the total. An agent with unscoped write access is a live liability from the day it launches, and code you don't own is money you spent on someone else's asset. Every other gap can be fixed later; those two get more expensive to fix every week.

    Why a score and not just a gut feel? Because by the third call, the most charming developer will have reset your standards. Writing the number down on the same sheet each time is how you compare like with like. It's the same reason pilots use a checklist on the thousandth flight.

    If you've narrowed down to two agencies and want to know how they compare on things you can check without a call, like published pricing, IP terms and review volume, the rubric in the companies ranking linked at the top is built for that. The two methods fit together: the ranking is the paper screen, this checklist is the interview.

    What Good Looks Like

    Every red flag has a green counterpart, and the green ones are mostly visible in the proposal before you ever get on a call.

    Instead of this red flagLook for this green flag
    A promised accuracy percentageA written evaluation plan with a threshold the agent must pass before launch
    Testing by feelA labelled test set built from your own examples, rerun on every change
    Prompt injection is handledA blast-radius design: least privilege, approval gates, output validation
    The agent has admin access to make it easyScoped credentials per tool and a boundary table you signed
    We will keep an eye on itTranscripts and traces you can read yourself, with a named person on alerts
    Built on the best modelA pinned model version, a migration runbook and a budget line for it
    You get a licence to use itA signed assignment of code, prompts, test sets and configuration
    The provider does not train on your dataA retention map covering every system that touches your data
    Build price onlyA monthly running-cost estimate tied to your volume
    Compliance is your lawyer's problemA short list of rules that apply, with dates, and what the build does about each

    Notice what the green column has in common. Nearly every item is a document: a test set, a boundary table, a retention map, a running-cost estimate, an assignment clause. Good developers write things down, because writing it down is how they catch their own mistakes before you pay for them.

    There's a second pattern too. The green answers all admit a limit. "I can't give you a percentage yet." "We assume some injections will get through." "This part should stay with a person." That's the lemons-market signal from earlier: an admission is cheap for a competent builder and costly for a bluffer, because the bluffer's whole pitch depends on there being no limits.

    Two green flags don't fit in the table. First, a developer who asks for your real examples before quoting. If they want to see 20 of your actual tickets or call recordings in the first week, they're sizing the job to your reality. Second, a developer who tells you not to build an agent for part of what you asked for. "This step is a form, not an agent" is one of the most valuable sentences you can hear in a discovery call, because it usually saves you money.

    How We Answer the Ten

    Since we're one of the developers you might vet, here's how Frenchy Digital answers the checklist, including where the answer has limits.

    We sell agents two ways. The first is eight fixed-scope packages, including the custom workflow agent and receptionist, lead generation and support agents, with the custom option on the AI agents hub. Each is a fixed $5,000 agent plus a one-time $5,000 setup fee, $10,000 at checkout, paid through one Stripe Checkout as one-time charges rather than a subscription. The product pages state timelines of 2 to 4 weeks for packaged builds and 2 to 6 weeks for the custom one, depending on how many systems are involved. The setup fee covers discovery, the build, connecting the agent to your calendar and phone or web channel, the guardrail work and tuning after launch.

    What the pages don't publish is a monthly fee or the ongoing model, telephony and hosting costs. Those depend on your volume and model choice, as the arithmetic in flag 9 shows, so we scope them on the call. I'd rather tell you that plainly than print a number that turns out wrong for you.

    The second way is a bespoke engagement through our AI agent creation service, for work that goes beyond a package: several systems to integrate, custom write paths, regulated data, multi-agent orchestration. Those are priced in bands: discovery and workflow audit $9k-$22k over 2-4 weeks; a single-workflow agent $28k-$70k over 4-9 weeks; a multi-workflow platform with system integration $70k-$180k over 9-16 weeks; enterprise, multi-site or regulated builds $180k-$420k+ over 14-24 weeks. Senior-led time is $150-$225/hr and retainers run $2,500-$9,500/month. The packages cost less because they're a fixed, pre-scoped job on systems you already use; the bands price the integration work a package doesn't include.

    Against the flags, then. On guarantees (1) we don't give one; we agree a test set and launch threshold. On evaluation (2), that test set is built from your examples. On injection (3) and write access (4), our process has a Guardrails step whose page wording is "Anything that commits you runs through deterministic code, not model output", and consequential actions get a human gate, like the confirmation gate in the Beyond Points build. On monitoring (5), the launch step promises "You get the full transcript log". On ownership (7), full source code and IP ownership transfer to you. There's a 30-day post-launch warranty, and a fixed-price phased proposal arrives within 5 business days of discovery.

    Where are our limits? We are a senior-led Black-owned Los Angeles agency, not a 500-person firm. If your procurement requires a SOC 2 report, ask us for it the same way you ask everyone else, read the period it covers, and treat the answer as part of the score. Model migration (6) is not priced on our product pages, so ask us to write down in the proposal who does that work and how it is billed. And on regulation (10) we design for the disclosure and logging duties we know about, but we'll tell you to take the final call to counsel.

    Limitations of This Checklist

    A checklist that claims to catch every bad developer would be committing red flag 1 on itself, so here's what this one can't do.

    • It tests answers, not code: A developer can say every right thing on a call and still ship a weak build. The checklist raises the floor; references from clients whose agents are in production, and a paid discovery phase before the full build, are how you check the ceiling.
    • The scoring thresholds are mine: The 16 and 10 cut-offs come from my own judgement, not from a study. No one has published outcome data linking vetting scores to project success, and I would not trust it if they had, for the reasons in the refused-numbers section.
    • Regulatory dates move: The EU and Colorado dates above were correct on September 28, 2026 and both have already moved once. Check them again before relying on them, and have counsel read the enrolled text rather than a summary.
    • Air AI is not final: The Air AI settlement is a proposed stipulated order. If the court enters it, or does not, the description above will need updating.
    • Provider policies change: Retention defaults, deprecation notice periods and list prices are the providers' own published terms on the day checked. They are not contractual commitments to you unless your agreement says so.
    • What I did not verify: I did not read any developer's actual contract for this article, did not test any agent, and did not check the internal security practices of any firm, including ours beyond what our own pages state. The OpenAI Atlas point is a paraphrase because the page blocked our fetcher.
    • It is not legal advice: The copyright and regulatory sections describe published law in general terms. I am not a lawyer. Your contract and your deployment need a lawyer's read.

    Is the checklist still worth running with those limits? Yes. The downside of asking ten questions is half an hour and some mild awkwardness. The downside of not asking is an agent with admin keys you don't own, built on a model that retires in 60 days. That's an asymmetric bet, and it's one you should take.

    Three Things to Do This Week

    You can run this checklist on your shortlist within a week.Here's the order I'd do it in.

    1. 1.Pull 20 real examples of the work you want the agent to do (tickets, emails, call recordings), including at least five of the messy ones. Send them to each developer on your shortlist and see who asks for more.
    2. 2.Book a 30-minute call with each and ask the ten questions from the scoring table, in order. Score every answer 0, 1 or 2 on the same sheet the same day, and ask for the answers to flags 4, 7, 8 and 9 in writing.
    3. 3.Before signing anything, get three documents: the boundary table of what the agent may do alone, needs approval for or must never touch; the running-cost estimate at your volume; and the IP assignment clause. If a developer will not produce all three, you have your answer.

    That's it. Three steps, one week, and you'll know more about each developer than their portfolio could tell you. Time to make some calls.

    Run the Checklist on Us

    Book a discovery call with Frenchy Digital, a senior-led Black-owned Los Angeles agency. Ask us all ten questions; we send a fixed-price phased proposal within 5 business days, and full source code and IP transfer to you.

    Want a Proposal That Passes All Ten?

    Book a discovery call. We map the workflow, write down what the agent may and may not do, and send a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Apt 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    1. 1SEC, In the Matter of Presto Automation Inc., Securities Act Release No. 33-11352 (January 14, 2025)↗
    2. 2FTC, FTC Announces Crackdown on Deceptive AI Claims and Schemes, Operation AI Comply (September 25, 2024)↗
    3. 3FTC, FTC Finalizes Order with DoNotPay (February 2025)↗
    4. 4FTC, FTC Approves Final Order Against Workado (August 28, 2025)↗
    5. 5FTC, Air AI and its owners settle FTC charges (proposed stipulated order, March 24, 2026)↗
    6. 6FTC, Air AI case page (checked September 28, 2026)↗
    7. 7UK NCSC, Prompt injection is not SQL injection (December 8, 2025)↗
    8. 8OpenAI, Continuously hardening ChatGPT Atlas against prompt injection attacks↗
    9. 9OWASP GenAI Security Project, LLM Top 10 2026↗
    10. 10OWASP GenAI LLM Top 10, official repository↗
    11. 11OWASP, Top 10 for Agentic Applications for 2026 (December 9, 2025)↗
    12. 12NIST AI 600-1, Generative AI Profile (July 26, 2024)↗
    13. 13Anthropic, Model deprecations↗
    14. 14OpenAI, API deprecations↗
    15. 1517 U.S.C. 101, Definitions (work made for hire), Cornell LII↗
    16. 16U.S. Copyright Office, Circular 30: Works Made for Hire↗
    17. 17OpenAI, Your data (API data controls)↗
    18. 18Anthropic Privacy Center, How long do you store my organization's data?↗
    19. 19Gibson Dunn, EU AI Act Omnibus agreement: postponed high-risk deadlines↗
    20. 20EUR-Lex, Regulation (EU) 2026/1744 (Digital Omnibus on AI)↗
    21. 21Holland & Knight, Colorado Governor Signs SB 189 (May 2026)↗
    22. 22Anthropic, Claude API pricing (list prices checked September 28, 2026)↗
    23. 23OpenAI, API pricing (list prices checked September 28, 2026)↗
    24. 24AICPA and CIMA, SOC 2 and the Trust Services Criteria↗
    Chris Machetto - CEO & Founder, Frenchy Digital of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2019 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.