Skip to main contentSkip to footer

    Top Rated & Verified

    Top Clutch App Development Company Black Owned United StatesTop Clutch Java Developers France 2026Top Clutch Service Line Blind Company Black Owned 2026Top Clutch App Development Company Minority Owned 2026Top Clutch Web Developers Black Owned 2026Top Clutch App Development Company Black Owned 2026Top Clutch Flutter Developers France 2026Top Clutch Health & Wellness App Developers France 2026Top Clutch Swift Company France 2026Top Clutch Machine Learning Company France 2026Top Clutch Chatbot Company France 2026Top Clutch Artificial Intelligence Company France 2026Top Clutch App Development Company Minority Owned Los Angeles
    Back to Blog
    AI Security
    September 30, 2026
    28 min read

    AI Agent Security:10 Controls Before You Connect Your Data

    Ten concrete controls to have in place before an AI agent gets credentials to your CRM, inbox, database or payment system, each mapped to the current OWASP, NIST and Model Context Protocol guidance, and framed for what they are: ways to shrink the damage, not to make prompt injection go away.

    A security review of an AI agent's connections to business systems, with scoped credentials, approval gates and an audit log
    LLM03
    Excessive Agency's rank in the 2026 OWASP LLM Top 10, up from LLM06 in 2025
    OWASP GenAI LLM Top 10 2026, via CSA research note
    Dec 9, 2025
    Release of the OWASP Top 10 for Agentic Applications
    OWASP GenAI Security Project
    2026-07-28
    Current Model Context Protocol version, with its own security best practices page
    modelcontextprotocol.io versioning page
    12
    Generative AI risks listed in NIST's Generative AI Profile
    NIST AI 600-1, July 2024

    Key Takeaways

    • Prompt injection is not solved and may never be. The UK NCSC says so plainly, so every control here is about shrinking the blast radius, not preventing the attack.
    • Excessive Agency is LLM03 in the 2026 OWASP LLM Top 10, up from LLM06 in 2025. The 2025 numbering is still on OWASP's main Top 10 page, so check which edition a vendor cites.
    • The ten controls split into three groups: limit the agent's authority, distrust what it reads and runs, and keep enough evidence to stop and reconstruct any incident.
    • The MCP spec (current version 2026-07-28) forbids token passthrough and tells clients to show the exact command before launching a local server. Most agent stacks now inherit those rules.
    • Anything that commits the business (a booking, a price, a refund, a record change) should run through deterministic code, not model output.

    The question before the connection

    The question I get most often at the start of an agent project is some version of "can it just log in as me?" The owner wants the agent to read the inbox, update the CRM, check the calendar and maybe issue a refund, and the fastest way to get there is to hand it the same admin login they use every morning.

    It works on day one. It also means that anything that can talk the agent into something can do everything you can do.

    That's the part people underrate. An agent reads text from places you don't control: customer emails, uploaded PDFs, web pages, support tickets, the notes field on a contact. Any of that text can carry instructions. And the model has no reliable way to tell your instructions from the ones hidden in a customer's email.

    So this piece is a checklist of ten controls I'd want in place before an agent gets a single credential to your data. For each one I'll say what it is, what it costs you, and which items in the current OWASP, NIST and Model Context Protocol guidance it answers. Then there's a mapping table you can hand to whoever does your security review.

    If you want the deep version of the attack itself, our guide to prompt injection and the OWASP LLM Top 10 covers it. This one is about what you put around the agent so that when an injection lands, it doesn't land on anything expensive.

    To be clear up front:none of these controls solves prompt injection. Nothing on the market does, and I'd be suspicious of anyone who says otherwise. What they do is shrink the blast radius, which is a different and much more achievable goal.

    Why a smarter model is not the fix

    The belief a competent person usually brings to this is reasonable: models are getting better at spotting manipulation, vendors ship guardrail filters, so the risk is shrinking and will eventually be handled by the model itself. Pick a good model, turn on the filter, connect the data.

    Here's why that doesn't hold. In December 2025 the UK's National Cyber Security Centre published a blog post by its CTO for architecture, Dave Chismon, titled Prompt injection is not SQL injection (it may be worse). The argument is simple. SQL injection got fixed by separating code from data at the database layer. A language model has no such layer. Inside the model, your instructions and the attacker's text are the same kind of thing: tokens to predict from.

    The NCSC's conclusion is that there's a good chance prompt injection will never be mitigated the way SQL injection was, and that defenders should spend their effort reducing impact. Specifically, it points to deterministic safeguards outside the model that constrain what the system can do, plus logging detailed enough to spot misuse.

    Think of it like hiring a very fast, very eager temp who will read every piece of paper put in front of them and believes all of it. You wouldn't fix that by finding a temp who is slightly more skeptical. You'd fix it by deciding which keys they get, which drawers stay locked, and which decisions need your signature.

    Therefore the right question isn't "is this agent safe?" It's "if this agent does the worst thing its permissions allow, what does that cost me, and how fast do I find out?" Every control below makes one of those two numbers smaller.

    One more thing that changed this year. OWASP's 2026 edition of its LLM Top 10 moved Excessive Agency from sixth to third. That's the risk category for agents that have more functions, permissions or autonomy than the job needs. The people who write these lists are telling you where the damage is coming from, and it isn't the model saying something rude.

    The four sources we mapped to

    Every control in this article maps to at least one item in four public documents, each checked on September 30, 2026.Here's what they are and what version is current, because the version confusion alone has tripped up more than one vendor deck I've read.

    OWASP Top 10 for LLM Applications, 2026 edition

    The OWASP GenAI Security Project published the 2026 LLM Top 10 in August 2026 (its resource page is dated August 3). According to the Cloud Security Alliance's research note, the order is: LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM03 Excessive Agency, LLM04 Supply Chain, LLM05 Data and Model Poisoning, LLM06 Unbounded Consumption, LLM07 Misinformation, LLM08 Hidden Context Exposure, LLM09 Vector and Embedding Weaknesses, LLM10 Improper Output Handling.

    The trap: OWASP's own LLM Top 10 landing page still showed the 2025 edition on the day I checked. In 2025, LLM05 was Improper Output Handling and LLM06 was Excessive Agency. In 2026, LLM05 is Data and Model Poisoning and LLM06 is Unbounded Consumption. Same numbers, different risks. When a vendor tells you it "covers LLM06", ask which year.

    OWASP Top 10 for Agentic Applications 2026

    A separate list, released December 9, 2025, for risks that only exist once a model can plan, remember and call tools. As summarized by Giskard, its ten items are ASI01 Agent Goal Hijack, ASI02 Tool Misuse and Exploitation, ASI03 Identity and Privilege Abuse, ASI04 Agentic Supply Chain Vulnerabilities, ASI05 Unexpected Code Execution, ASI06 Memory and Context Poisoning, ASI07 Insecure Inter-Agent Communication, ASI08 Cascading Failures, ASI09 Human-Agent Trust Exploitation and ASI10 Rogue Agents.

    It builds on OWASP's earlier Agentic AI Threats and Mitigations guide (v1.0, February 17, 2025). In September 2026 the project also introduced an Agent Control Standard at version 0.1, a framework for runtime enforcement with policy hooks and an agent bill of materials. It's early; I mention it so you recognize the name, not because you should require it yet.

    NIST AI RMF 1.0 and the Generative AI Profile

    The NIST AI Risk Management Framework 1.0 came out January 26, 2023 and organizes the work into four functions: Govern, Map, Measure and Manage. NIST says it's being revised as part of the White House AI Action Plan. The Generative AI Profile, NIST AI 600-1 (July 26, 2024) adds twelve risks specific to generative AI, among them Confabulation, Data Privacy, Human-AI Configuration, Information Integrity, Information Security and Value Chain and Component Integration.

    Neither document was written for agents. NIST knows: its Center for AI Standards and Innovation announced an AI Agent Standards Initiative in February 2026, alongside a request for information on agent security. The NCCoE followed with a concept paper on agent identity and authorization dated February 5, 2026, which asks, among other things, how to identify, authorize and audit agents and how to control prompt injection. These are inputs to future guidance, not guidance you can be audited against today.

    The Model Context Protocol security best practices

    MCP is the protocol most agent stacks now use to connect a model to tools and data. The current protocol version is 2026-07-28, and it ships a Security Best Practices page covering confused deputy attacks, token passthrough, server-side request forgery, state handle hijacking, local MCP server compromise, OAuth URL validation, mix-up attacks and scope minimization.

    Even if your agent doesn't use MCP, read it. It's the most concrete public document I know on how an agent's credentials go wrong, and it's written in MUST and SHOULD language you can paste into a contract.

    Controls 1 to 3: authority

    The first three controls decide what the agent is allowed to do at all. If you only have budget for three, these are the three, because they cap the damage of every other failure.

    Control 1, Least-privilege scoped credentials

    What it is. The agent gets its own identity and its own credential, scoped to exactly the operations its job requires. Not your login. Not the shared admin API key the office manager set up in 2022. A service account or OAuth grant that can read contacts and create notes, say, and nothing else.

    Why it matters.This is the main lever against Excessive Agency (LLM03) and ASI03 Identity and Privilege Abuse. The MCP spec's scope minimization section is blunt about the failure: a token with broad scopes such as files, db or admin wildcards, once leaked, gives an attacker lateral access and is hard to revoke without breaking everything else. It lists wildcard scopes and bundling unrelated privileges as common mistakes, and recommends starting from a minimal read-oriented scope and elevating only when a privileged operation is actually attempted.

    The same page says MCP servers must not accept tokens that weren't issued for them, a practice it calls token passthrough. In plain terms: the agent's tool server shouldn't just forward your user token downstream. It should hold a credential that's meant for it. The broader OAuth rules behind this are in RFC 9700, the IETF's current OAuth security best practice.

    The cost. Setup takes longer, and some SaaS products don't offer fine-grained scopes at all. When they don't, you put a thin service of your own in front of the API that only exposes the operations you've approved. That's extra code to maintain. It's worth it. We go through the mechanics in our piece on agent sandboxing and credential scoping.

    Control 2, Read-only by default, with approved write paths

    What it is.The agent starts with read access only. Each write is added as a named, narrow operation with a reason: "create a follow-up task on a contact", "add a tag from this list of six". Never "update any field" and never delete.

    Why it matters.Reads can leak data, which is bad. Writes can change your business, which is often worse and harder to notice. ASI02 Tool Misuse and Exploitation covers an agent misusing tools it's legitimately allowed to use. The simplest defense is to not give it tools with more power than the job needs. A tool called update_record with free-form fields is a gift to an attacker. A tool called add_followup_note(contact_id, text) is boring, and boring is what you want.

    The cost.You'll write more tools, each smaller. And you'll have to say no to a few convenient features in the first version. I'd rather add a write path in month two because a real task needed it than remove one in month two after it was abused.

    Control 3, Human approval for consequential actions

    What it is. A list, written before launch, of actions that need a person to approve them: payments and refunds over a threshold, sending anything to more than a handful of recipients, changing a price, deleting data, signing anyone up for anything. The approval happens outside the model, in a UI that shows the exact action and its arguments.

    Why it matters.This is the last line against Excessive Agency, and it answers the NIST 600-1 risk category Human-AI Configuration. But there's a catch that the Agentic Top 10 names directly as ASI09 Human-Agent Trust Exploitation: people learn to click approve. An approval screen that shows "Agent wants to proceed. Approve?" is theatre. It needs to show what will happen, to whom, for how much.

    The cost. Latency and staff time. Every approval is a person interrupted. So keep the list short and make the threshold real: if 95 out of 100 approvals are routine, raise the threshold on that action rather than training your team to rubber-stamp. We wrote about where to draw these lines in designing human-in-the-loop for agents.

    Controls 4 to 6: inputs and tools

    The next three controls deal with what flows into the agent and what the agent is allowed to plug into. This is where most real attacks start.

    Control 4, Treat all retrieved content as untrusted

    What it is.Anything the agent didn't get from you directly is treated as potentially hostile: emails, tickets, PDFs, web pages, search results, documents in your knowledge base that someone else uploaded, even the output of another agent. That content can inform an answer. It can't authorize an action.

    Why it matters. This is LLM01 Prompt Injection in its indirect form, and ASI01 Agent Goal Hijack, where the attacker redirects the agent through text it reads. It also covers ASI06 Memory and Context Poisoning: if the agent stores notes or retrieves from a vector database, a poisoned entry can shape its behavior long after it was planted. That links to LLM05 Data and Model Poisoning and LLM09 Vector and Embedding Weaknesses.

    In practice, the strongest pattern I know is splitting the agent's job: the part that reads untrusted text doesn't get the tools that can do damage, and anything it extracts is passed on as data in a fixed shape (a category, a date, an amount), not as free text the next step obeys. The NCSC also suggests marking untrusted sections clearly in the prompt. That helps at the margin. It doesn't replace the permission boundary.

    The cost. Split designs are more work and sometimes a bit less clever. An agent that can read an email and immediately act on anything in it feels magical in a demo. It also works perfectly for the attacker who sends that email.

    Control 5, Validate output before it touches anything

    What it is. Model output is checked by ordinary code before it goes anywhere that matters. Tool arguments are validated against a schema and a business rule (is this contact ID one this user can see? is this amount under the cap? is this date a slot the calendar actually has?). Anything rendered in a browser is escaped. Anything run as code or a query is parameterized or sandboxed.

    Why it matters. This is LLM10 Improper Output Handling in the 2026 list and ASI05 Unexpected Code Execution in the agentic one. The MCP spec gives a sharp example: a malicious server can hand a client an authorization URL with a javascript: scheme, and a client that opens it blindly has just run attacker code. Its fix is the fix everywhere: allow only expected schemes and values, reject everything else, and never pass model or server output through a shell.

    The cost. Almost nothing, if you do it from the start. Retrofitting validation after the agent has been talking to production for three months is where it gets painful.

    Control 6, Allowlist and pin tools and MCP servers

    What it is.The agent can only connect to tools and MCP servers on an explicit list, each at a pinned version you've reviewed. No installing a new connector from a directory because it looked handy. Updates go through the same review as new tools.

    Why it matters.This covers LLM04 Supply Chain, ASI04 Agentic Supply Chain Vulnerabilities and NIST 600-1's Value Chain and Component Integration risk. The MCP spec's section on local server compromise describes exactly how this goes wrong: a startup command hidden in a client config, or a malicious payload inside the server itself. It says a client offering one-click local server setup must show the exact command without truncation and require explicit approval, and should run servers sandboxed with minimal default privileges. It also covers SSRF, where a malicious server points the client at internal addresses such as a cloud metadata endpoint.

    A tool's description is text the model reads. A tool you don't control can change its own description tomorrow. Pinning is how you notice.

    The cost.You lose the convenience of trying every new connector the week it ships. Your team will grumble. That's fine.

    Controls 7 to 10: evidence and recovery

    The last four controls assume something will go wrong and make sure you can see it, stop it and pay less for it.

    Control 7, Secrets management and vendor data terms

    What it is.API keys and tokens never go in the prompt, the model's context or the logs. They sit in a secrets manager, and deterministic code attaches them to outbound calls after the model has decided what to call. Each agent gets its own credentials so you can rotate or revoke one without touching the rest. Alongside that, you know in writing what each model and tool vendor does with your data: whether it's retained, for how long, and whether it's used for training.

    Why it matters. A secret in the context window is a secret an injected instruction can ask the model to repeat. That's LLM02 Sensitive Information Disclosure, and it sits close to the 2026 list's new LLM08 Hidden Context Exposure. The OWASP Secrets Management Cheat Sheet covers the ordinary mechanics: centralize, rotate, audit access. The vendor side maps to NIST 600-1's Data Privacy risk.

    The cost. Mostly discipline. The vendor terms part is tedious reading, and the terms change. Screenshot or save the version you agreed to.

    Control 8, A complete audit log

    What it is.Every model input and output, every tool call with its arguments and result, the identity the agent acted as, every approval and who gave it, all timestamped and stored somewhere the agent can't edit. Plus a retention period you've chosen on purpose.

    Why it matters.Without it you can't answer the first question after an incident, which is "what did it do?" The NCSC lists monitoring of inputs, outputs, tool use and API calls as one of its four responses to prompt injection. The MCP spec flags audit trail damage as one of the risks of token passthrough, because downstream logs end up showing the wrong identity. And ASI08 Cascading Failures is much easier to contain when you can trace which step started it.

    The cost. Storage, and a privacy question: your log now holds everything the agent saw, which can include personal or health data. Treat the log as sensitive data in its own right, with access control and a deletion schedule.

    Control 9, Rate and spend limits

    What it is. Hard caps outside the model on how many actions the agent can take per minute and per day, how much it can spend on model tokens and third-party APIs, and how much value it can move (refunds, credits, discounts) per day. When a cap is hit, the agent stops and a person is told.

    Why it matters. This is LLM06 Unbounded Consumption in the 2026 list, and it limits ASI02 and ASI08 too, since a looping or hijacked agent does its damage through volume. A cap is the cheapest control on this list and the one I see skipped most often.

    The cost. The occasional false stop on a legitimately busy day. Set caps from real usage after the first few weeks, not from a guess, and make raising one a deliberate act.

    Control 10, Kill switch, incident plan and red-team testing

    What it is.Three things that belong together. A kill switch that stops the agent from acting within minutes, ideally by revoking its credentials at a layer the agent can't reach. A one-page incident plan: who pulls the switch, who checks the log, who talks to customers. And adversarial testing before launch and after every change to the model, prompt, tools or data sources, using the channels the agent actually reads, since indirect injection through a document or email is the realistic attack.

    Why it matters.ASI10 Rogue Agents and ASI08 Cascading Failures both assume an agent will at some point act outside its intended function. NIST's Govern and Manage functions expect a documented response. And testing is how you find out whether controls 1 through 9 work, rather than assuming.

    The cost. A few hours to write the plan, a drill before launch, and ongoing testing time. Testing is never finished, because the model underneath changes.

    The mapping table

    Each control, and the specific items it answers in the four source documents.Hand this to your security reviewer. It's a starting point for their judgment, not a certification.

    ControlOWASP LLM Top 10 (2026)OWASP Agentic Top 10 (2026)NIST AI RMF and AI 600-1MCP Security Best Practices
    1. Least-privilege credentialsLLM03 Excessive Agency; LLM02 Sensitive Information DisclosureASI03 Identity and Privilege AbuseGovern, Manage; 600-1 Information SecurityScope Minimization; Token Passthrough
    2. Read-only by defaultLLM03 Excessive AgencyASI02 Tool Misuse and ExploitationMap, ManageScope Minimization (progressive elevation)
    3. Human approvalLLM03 Excessive Agency; LLM07 MisinformationASI09 Human-Agent Trust Exploitation; ASI10 Rogue AgentsGovern, Manage; 600-1 Human-AI ConfigurationConsent requirements for local servers
    4. Untrusted retrieved contentLLM01 Prompt Injection; LLM05 Data and Model Poisoning; LLM09 Vector and Embedding WeaknessesASI01 Agent Goal Hijack; ASI06 Memory and Context PoisoningMap, Measure; 600-1 Information IntegrityNot addressed directly
    5. Output validationLLM10 Improper Output HandlingASI05 Unexpected Code ExecutionMeasure, ManageOAuth Authorization URL Validation
    6. Tool and MCP allowlist, pinningLLM04 Supply ChainASI04 Agentic Supply Chain Vulnerabilities; ASI07 Insecure Inter-Agent CommunicationGovern; 600-1 Value Chain and Component IntegrationLocal MCP Server Compromise; SSRF; CIMD trust policies
    7. Secrets and vendor data termsLLM02 Sensitive Information Disclosure; LLM08 Hidden Context ExposureASI03 Identity and Privilege AbuseGovern; 600-1 Data PrivacyToken Passthrough; Confused Deputy
    8. Audit logLLM01 Prompt Injection (detection)ASI08 Cascading Failures; ASI10 Rogue AgentsMeasure, ManageToken Passthrough (audit trail risk); scope elevation logging
    9. Rate and spend limitsLLM06 Unbounded ConsumptionASI02 Tool Misuse and Exploitation; ASI08 Cascading FailuresManage; 600-1 Environmental ImpactsNot addressed directly
    10. Kill switch, incident plan, red teamAll ten, through testingASI10 Rogue Agents; ASI08 Cascading FailuresGovern, Measure, ManageScope Minimization (revocation)

    OWASP LLM numbering is the 2026 edition. Agentic item names are as summarized by Giskard from the OWASP list. NIST AI RMF entries name functions only; I haven't mapped to individual subcategories because the framework is under revision. "Not addressed directly" means the MCP security page doesn't have a section on it, not that MCP is unsafe there.

    Two patterns stand out when you read the table down rather than across.

    First, Excessive Agency shows up in controls 1, 2 and 3. That's not a coincidence. The single most effective thing you can do about the number one risk (prompt injection) is manage the number three risk (what the agent can do once injected).

    Second, the MCP column is thickest on credentials and supply chain. That tells you where the protocol's authors have seen real trouble. The spec has almost nothing on prompt injection itself, which is honest: a transport protocol can't fix what the model reads.

    Doing the blast radius arithmetic

    For every system you connect, write down the worst thing the agent's permissions allow, and what it would cost.Then see which control shrinks it. This is the exercise I'd do before any other security work, because it tells you where the effort goes.

    Agent setupWorst plausible actionRough cost of that actionControls that shrink it
    Support agent, full CRM admin tokenExport or delete every customer recordA notification event plus rebuild; unbounded1, 2, 7, 10
    Support agent, read contacts plus create noteWrite a misleading note on some recordsAn hour of cleanup per batch2, 5, 8
    Billing agent that can issue refunds freelyRefund many orders after an injected emailSum of refunds issued before anyone notices3, 9, 10
    Billing agent, refunds under a cap, over cap to a personRefunds up to the cap, capped per dayCap times the daily limit3, 9

    Consider a scenario. A small online store connects a billing agent that can issue refunds, so it can handle "my order arrived broken" emails without waiting for staff. With no cap, one well-crafted email thread could, in principle, trigger refunds until someone looks at the payment dashboard. The worst case is the sum of every refundable order in that window. That number doesn't have a ceiling you can write down, which is the problem.

    Now apply three controls. Refunds under $50 go through automatically, anything above goes to a person (control 3). The agent can issue at most 20 automatic refunds a day (control 9). Every refund is logged with the email that triggered it (control 8). The worst case is now 20 times $50, so $1,000 a day, and you'll see it in the log the same afternoon.

    Going from "unbounded" to "$1,000 and visible" didn't require a better model or a detection product. It took two numbers and a log. That's the whole philosophy in one example, and to be clear, those two numbers are illustrative: set yours from your own order values and volume.

    The same arithmetic works for data. An agent with read access to 40,000 customer records has a worst case of 40,000 records disclosed. An agent that can look up one customer at a time by verified email, capped at a few hundred lookups a day, has a worst case measured in hundreds. Same feature for the customer. Very different incident.

    What we do on our builds

    I'll describe how we build, and I'll be specific about what we don't have, because that part matters as much.

    Our process for AI agent development has four steps: discovery, build, guardrails and launch. Discovery is where most of these controls get decided. We map which systems the agent has to touch and where a human should take over, and that map is what the agent is written against.

    The guardrails step is the one I'd point a security reviewer to. Anything that commits the business runs through deterministic code, not model output. An agent that books appointments can only offer slots the calendar actually has, so it can't invent a time, quote a price you don't charge, or promise something you can't deliver. That's control 5 applied to the actions that matter most, and it's the NCSC's "deterministic safeguards" point in practice.

    After launch, you get the full transcript log, so you can read exactly what the agent says to your customers rather than taking our word for it. You also get the full source code and IP, which means the credentials, the tool list and the caps live in code you own and can have anyone audit.

    What we don't have.Frenchy Digital holds no SOC 2, ISO 27001 or HITRUST attestation. We sign a BAA with healthcare clients. If your procurement requires one of those attestations from the builder, we're not the right fit for that requirement, and I'd rather you know now.

    On cost, since it comes up: a custom agent is $5,000 plus $5,000 setup, so $10,000 to start, and it rises with complexity and the number of integrations. You pay model, telephony and API usage directly at provider rates, which also means spend caps (control 9) are set in accounts you control. After launch it's $2,500 a month. For workflows that touch several internal systems, the custom workflow agent page describes what's included.

    The one real agent build we've published a case study for is Beyond Points AI. The case study describes an agent that searches award and cash inventory, transfers points and books trips through a confirmation-gated browser agent, which is control 3 in its most literal form: the booking doesn't happen until the traveller confirms. I'm not going to claim security outcomes for it beyond what that page says.

    Red flags in a vendor

    When you evaluate an agent vendor or a developer, these answers should make you slow down. None is automatically disqualifying. All of them deserve a follow-up question.

    • "Our guardrails prevent prompt injection.": Nobody can prevent it. Ask what the agent can do if injection succeeds. That answer is the real security posture.
    • They want an admin login or a personal token.: Control 1 fails at the first step. Ask for a dedicated service account and a written list of scopes.
    • They cite "LLM06" without a year.: In 2025 that was Excessive Agency; in 2026 it is Unbounded Consumption. It tells you whether they have read the current list.
    • No list of actions that need human approval.: If they can't name which actions a person signs off on, nobody has decided.
    • Tools installed from a marketplace, unpinned.: Ask for the tool list with versions and who reviews updates.
    • Logs you can't read, or logs they keep and you don't.: You should be able to reconstruct any session without asking permission.
    • No kill switch answer.: Ask them to show, not describe, how the agent is stopped. Time it.
    • A security certificate you can't see.: Ask for the report itself. A badge on a website is a claim, not evidence. "Aligned with SOC 2" is not SOC 2.
    • Vendor-published detection or block rates.: A figure like "blocks 99% of attacks" is the vendor's own test on its own chosen attacks. We do not repeat any such figure, and it should not be the reason you buy.

    One pattern runs through all nine. A vendor who answers in terms of what the agent can't do (scopes, caps, approvals) has thought about security. A vendor who answers in terms of what the model won't do (it's trained to refuse, it detects attacks) is describing a hope.

    Limitations

    Here's what this article can't tell you, stated plainly.

    I took the 2026 OWASP LLM Top 10 ordering from the Cloud Security Alliance's research note, because OWASP's resource page links to a download rather than listing the entries, and its main Top 10 page still showed 2025. The ordering agrees with other published summaries, but if OWASP revises it, trust the PDF over me.

    The Agentic Top 10 item names come from a third-party summary of the OWASP document, not from my own reading of the PDF. The numbering is widely repeated; the exact wording may differ slightly in the original.

    The mapping table is my judgment. OWASP, NIST and the MCP authors don't publish a crosswalk to these ten controls, and a reasonable reviewer could map some rows differently. OWASP has published its own industry framework crosswalk; I didn't use it here.

    NIST's agent work is at the request-for-information and concept-paper stage. There's no NIST agent security standard you can be audited against today, and I haven't tried to predict what it will say.

    I've refused every statistic I found on how often agents fall to prompt injection or how well a given filter blocks it. The ones I saw were vendor tests on vendor-chosen attacks. The same goes for the industry failure-rate numbers that tend to show up in pieces like this ("95% of pilots fail" and the rest); none of them measures security, and none appears here.

    Finally, these controls reduce risk; they don't remove it. An agent with every control on this list can still be manipulated into doing something within its permissions that you didn't want. That residual risk is the reason controls 1 to 3 exist.

    Three things this week

    If you already have an agent connected, or you're about to, do these three things before Friday.

    • 1. List every credential the agent holds.: For each one, write what it can do and whether it's shared with a person. Anything shared or admin-level goes to the top of the fix list.
    • 2. Write the blast radius line for each system.: One sentence: the worst thing the agent's access allows, and a rough cost. If you can't write a number, that system needs a cap or an approval gate.
    • 3. Pull the kill switch once, on purpose.: Revoke the agent's credentials in a test window and time how long it takes to stop acting. If nobody knows how, you've found your first incident-plan gap.

    That's an afternoon of work, and it tells you more about your real exposure than any vendor's security page will.

    Want an Agent With the Guardrails Written In?

    Book a discovery call with Frenchy Digital, a senior-led Black-owned Los Angeles agency. We map what your agent must touch, what it must never touch and where a person signs off, then send a fixed-price phased proposal within 5 business days.

    Want an Agent Built With These Controls?

    Book a discovery call. We map what the agent must touch, what it must never touch, and where a person signs off, then send a fixed-price phased proposal within 5 business days.

    1517 S Bentley Ave Apt 204, Los Angeles CA 90025

    Frequently Asked Questions

    Sources & References

    Chris Machetto - CEO & Founder, Frenchy Digital of Frenchy Digital

    Chris Machetto

    CEO & Founder of Frenchy Digital. Building apps and digital products since 2016 for startups and enterprises across LA, San Francisco, Paris, Geneva, and more globally.