Field note

Sep 28, 2026

AI agent safety: why I stopped trusting a moving cap

AI agent safety means blocking unapproved actions before they run. My pilot exposed a moving spend limit; here is the buying test I use for a real stop.

Damian Moore
Damian MooreSeptember 28, 2026

AI agent safety means keeping an agent inside agreed limits on data, decisions and spending, even when its next proposed action is wrong. I build these systems, so I have a bias toward controls I can actually test. A limit I keep raising myself is not much of a limit. I have supplied that particular demonstration.

A mechanical stop separates an AI agent's requested action from permission to spend

What is AI agent safety in a working business?

I separate answer quality from action safety. A model can misunderstand a request without causing an external change if the surrounding system refuses the action. It can also give a convincing explanation while making a change nobody authorized. I need to test both paths.

For an operator, that means answering plain questions. Which records may this agent see? Which changes can it make without another person? Can it send a message, place a call, or commit money? What happens when the request arrives twice?

OWASP describes excessive agency in terms of excessive functionality, permissions and autonomy. I find that more useful than treating every failure as a model intelligence problem. An agent that only needs to read a customer record should not get permission to delete it just because the connector offers both actions.

I use that distinction when scoping business process automation. The job is not to give the model every tool and hope its instructions are persuasive. It is to decide which business actions belong inside the assignment, then keep the rest outside its reach.

Why did I replace a call-count cap with a spending stop?

On a voice-intake pilot, I had raised a call-count cap three times in one evening. The approved test budget was $5. Counting calls was a convenient proxy, but it was not the same thing as enforcing that budget.

I moved the ceiling into the database and reserved budget for each test case before dialing. The service would refuse to place the call when the money was not available. That made the rule part of the action path instead of another number I could keep nudging to get through the next test.

My opinion is simple: a warning after spending is not a spending control. That $5 was a bounded pilot allowance, not the price of an agent or a claim about normal operating cost. The amount was small. The failure pattern was the important part.

I am not presenting this as proof of production customer outcomes. It was a controlled pilot. What it demonstrated was a useful design decision: check authority before the action that consumes it, rather than asking a person to interpret an alert afterward.

A finite allowance and reserved portions illustrate why in-progress work must count before another paid action starts

When I apply the same idea elsewhere, I start with the thing the business actually wants to limit. A call count is different from money. A daily message allowance is different from permission to contact a particular person. A document limit is different from permission to read that document.

I have covered the broader operator authority test separately. Here the narrower question is whether the system can enforce a refusal when the next action looks useful but is outside the allowance.

How do I judge whether a safety control is real?

I use these questions before I recommend broader access:

  • Does the stop happen before the external action? I want the call, send, payment or record change blocked at the point where it would happen.
  • Does work already in progress count? I want reserved work included, so simultaneous requests cannot each assume the same remaining allowance is theirs.
  • Can ordinary input change the rule? I do not want a customer message or retrieved document to grant new permissions.
  • Can an operator see what was refused and why? I want a useful exception record with an owner, not an unexplained red badge.

These are buying criteria, not a claim that every system needs a custom build. I ask an existing vendor to show its own controls first. If the help desk already has the right permissions and approval path, adding a separate agent service may introduce more operating burden than it removes.

I also distinguish a permission failure from a dependency failure. If an action is forbidden, trying again later should not make it permitted. If a provider is unavailable, the retry may be legitimate, but I still need to know whether the earlier attempt already succeeded.

Which controls belong outside the model?

I keep decisions with clear business limits out of the model's discretion. The model may interpret a request or propose an action. I want the connected system to enforce identity, permission, spending allowance and any required approval.

Business actionWhat I let the agent proposeWhat I want enforced separately
Reply to a customerDraft the responseApproved recipient, send permission and current message
Update an accountSuggest a specific changeAllowed fields and the correct customer record
Place a paid callPrepare the next attemptRemaining allowance before dialing
Approve an exceptionExplain the issueNamed approver and the exact action approved

I do not treat a stronger prompt as a substitute for these checks. OWASP's prompt-injection guidance includes indirect attacks through external sources such as websites and files. The relevant buyer question is not whether the agent will ever read a strange instruction. It is what that instruction could cause the connected system to do.

For example, I would not let text in an inbound invoice authorize a new payment destination. Reading the request and accepting its authority are separate decisions. I would want the payment workflow to keep enforcing its existing approval rules even if the agent proposed the change.

This is also why I ask about the connection identity during legacy-system integration work. A narrow-looking agent interface can still hide broad permissions in the underlying account. I want the permitted actions and the actual account privileges to agree.

What should happen when the agent refuses?

I do not consider the workflow finished just because it avoided the wrong action. The original request still needs a recorded outcome.

For a spent allowance, I want the next action recorded as not attempted, with the reason visible. For missing approval, I want a specific person to own the decision. For unavailable data, I want the request to remain unresolved rather than be silently marked complete.

I also want an explicit way to resume. An operator should be able to identify the pending request, review what has changed, and authorize the next step without blindly replaying everything. Where an earlier action may have succeeded, I want its outcome checked first.

That is the difference between a safety feature and an operating workflow. The first prevents one mistake. The second also tells the team where the work went. My agent workflow approach starts with that business promise rather than the tool sequence.

How do I test a refusal without testing on customers?

I start with controlled records, approved test destinations and a bounded allowance. Then I ask for evidence from the denied path, not just the successful path.

On the same pilot, the urgency rules had 105 tests covering positive cases, ordinary negatives and boundary cases. I kept that separate from live-call acceptance. A test suite can prove the cases it exercises; it does not prove the whole customer experience or replace a real acceptance run.

I had also demonstrated a different failure on an earlier application: after clearing the cache, I was no longer logged in, yet I could still enter the workflow. That is why I do not accept a visible login screen as proof that the protected action is protected. I want the unauthorized attempt tested directly.

Separate accepted and refused request cards show how a test should prove the boundary, not just a successful result

For a buying evaluation, I would ask the supplier to show these attempts using its test environment:

  1. An action outside the agent's assigned permissions.
  2. A paid action after its allowance is exhausted.
  3. The same approved request submitted again.
  4. An action whose approval no longer matches the current request.
  5. An unavailable dependency, followed by recovery.

I would ask to see the resulting records as well as the screen. A refused request should not quietly become a successful action downstream. A failed request should not disappear from the operator's queue. A repeated request should not create another commitment merely because it arrived again.

Do I need a governance platform or a smaller fix?

I start with the workflow that can cause a material problem. If one existing permission setting closes the gap, I would use it. If several systems share the action and none owns the limit, I would consider a small control layer before a broad platform purchase.

The NIST AI Risk Management Framework is voluntary guidance for incorporating trustworthiness into AI design, development, use and evaluation. I use that kind of guidance to structure questions, not as a badge that replaces evidence from the actual workflow.

The current argument about agent behavior can distract from this. In his essay on the term “rogue” agents, Eoin Higgins argues that the language can obscure company responsibility. I do not need to settle every claim in that debate to ask who granted an action, who can withdraw that permission, and who owns the failure.

When I evaluate AI governance platforms, I want them to make those answers easier to prove. A complete inventory is useful. It is not the same as a blocked action when the allowance runs out.

When not to hire us for this

You do not need us for this if the work stays in draft form, an accountable person reviews it before any external action, and your existing software already enforces the relevant access and spending limits. I would keep that arrangement until a specific failure or workload justifies changing it.

I would also pause an autonomous rollout when nobody can agree who owns exceptions. Building more capability does not resolve an unresolved operating decision. I would narrow the assignment first.

My verdict: start with a useful action, a concrete limit and a demonstrated refusal. I trust AI agent safety more when the system can stop me from moving the cap than when it can explain why it intends to be careful.

FAQ

Frequently asked questions

01What is AI agent safety?

I treat it as controlling what an agent can read, change, send and spend, then proving those limits hold when a request is wrong or a dependency fails. A polite answer is not proof that the connected action was safe.

02Is a spending alert enough to control an AI agent?

I use an alert for visibility, not permission. If spending must stop at a ceiling, I want an enforced check before the paid action, including allowance for work already in progress.

03Does human approval make an AI agent safe?

I still check what the person approved, whether it changed, and whether the same approval can be used again. Approval helps only when the actual action remains tied to the reviewed request.

04Can prompt injection come from a document?

Yes. I treat instructions inside retrieved documents as untrusted content, not permission to change the operating rules. I also keep sensitive actions behind controls outside the model.

05What should I ask an AI agent vendor to demonstrate?

I ask for a refused action, an exhausted spending allowance, a repeated request, and an unavailable dependency. I want the resulting business record and the next owner, not only a successful demo.

Related reading

Next step

Want help applying this?

Run the 90-second AI Operations X-Ray and I'll show you where to start.