Pedro Sousa
‹ ← Articles

Article · Blog

An agent that decides too much is just as dangerous as one that decides too little

Most problems with AI agents in production don't come from hallucination. They come from bad boundary design: the agent takes on responsibilities it shouldn't have, or gets stuck waiting for confirmation on trivial things. This article is about how I learned the difference in practice.

An agent that decides too much is just as dangerous as one that decides too little

The first agent I put into production had a problem I couldn't quite name.

It worked. It made correct decisions most of the time. But every now and then it did something no human would have done without asking first. Not because the model was wrong. Because I never defined how far it could go.

It took me a few weeks to understand that the problem wasn't the LLM. It was the autonomy boundary design.

What most people call "hallucination" isn't hallucination

When an agent makes a wrong decision in production, the most common diagnosis is: "the model hallucinated." That does happen. But it's the least frequent cause.

What happens more often is this: the agent was put in an ambiguous situation, without enough context to distinguish between two possible paths, and it picked one. The model did exactly what it was designed to do. The problem is that you shouldn't have let it reach that crossroads alone.

In a system I built for content automation, the agent had access to a publishing tool. The instruction was "publish when the content is approved." What I never defined was what "approved" meant when the intermediate state was ambiguous. The agent interpreted, chose, published. It was technically within the rules I had given it. And it was exactly wrong.

That wasn't hallucination. It was absent boundary design.

The problem of two extremes

There's a real tension in agent design that most articles ignore.

On one side, the agent that decides too much. It has broad autonomy, interprets instructions freely, acts without asking for confirmation. In a test environment, it seems impressive. In production, it eventually does something you didn't authorize, at a moment you didn't expect, with consequences you can't easily reverse.

On the other side, the agent that decides too little. It asks for confirmation on everything. It interrupts flows that should be automatic. It turns automation into semi-automation without making that clear. The user learns to click "yes" without reading, which is worse than letting the agent decide on its own.

Most teams swing between the two without realizing it. They start restrictive out of fear. They loosen up when the product pushes for more autonomy. They don't have a clear mental model for where the boundary should actually be.

Boundary design isn't about trusting the model

Here's the counterintuitive part: the autonomy boundary of an agent shouldn't be calibrated by how much you trust the model.

It should be calibrated by the reversal cost of the action.

A read action has zero reversal cost. An action that writes to a database has low cost if there's a soft delete. An action that sends an email to a thousand users has a very high cost, possibly irreversible. An action that publishes public content has medium to high cost depending on context.

When you think in terms of reversibility, the design becomes clearer. The agent can have full autonomy for actions with low reversal cost. For actions with high cost, you need a human checkpoint or an explicit confirmation layer built into the agent's own flow.

This isn't about not trusting the LLM. It's about recognizing that no system, human or automated, should have unrestricted autonomy over irreversible actions.

What I call the Confident Agent Trap

There's a pattern I've seen repeat itself across different systems.

The agent works well in 95% of cases. The team gains confidence. Autonomy increases incrementally, each adjustment too small to justify an architecture review. Six months later, the agent is making decisions that no one would have explicitly approved if they'd been presented all at once.

I call this the Confident Agent Trap. The autonomy wasn't granted. It was accumulated through inertia.

The symptom is when someone on the team asks "can the agent do X?" and nobody can answer with certainty without going to read the code or the prompts.

What actually works in practice

The approach that makes the most sense to me, after getting it wrong a few times, is to clearly separate three zones:

Autonomous execution zone. Actions the agent can take without any confirmation. Reading, searching, analyzing, formatting, drafting. Anything that doesn't alter external state in an irreversible way.

Execution with logging zone. Actions the agent executes, but that generate an explicit, auditable record in real time. Not for approval, but for visibility. You don't want to stop the flow, but you want to know what happened.

Mandatory approval zone. Actions that require human confirmation before executing. Not because the agent is incompetent, but because the reversal cost justifies the checkpoint.

The hard part is classifying every tool you give the agent into those three zones before going to production... not after the first incident.

Context is what determines the boundary, not the prompt

A mistake I made a few times was trying to solve boundary design in the prompt. "Don't do X without confirming." "Only publish if Y is true."

That works until the model encounters a situation the prompt didn't anticipate. And it always does.

What works better is building the boundary into the architecture. The publishing tool simply doesn't exist in the list of available tools when the context doesn't satisfy the preconditions. It's not the model deciding whether or not to publish. It's the system never offering that option in the first place.

Context here isn't just what you put in the prompt. It's the set of available tools, the active permissions, the session state. The agent decides within the space you built. If the space is poorly defined, the model will explore the edges... and the edges are where the problems show up.

What I would do differently today

I would map tools by reversibility before starting to implement the agent. Not as after-the-fact documentation, but as a design constraint.

I would implement a dry-run mode for any action in the approval zone, where the agent describes what it's about to do before doing it. Not as mandatory human confirmation, but as an observability mechanism and a boundary testing tool.

And I would avoid increasing autonomy under product pressure without reviewing the reversibility map alongside it. Autonomy that grows without boundary review is silent accumulated risk.

The takeaway

There's a common belief that engineering work in agent systems is mainly about the model: choosing the right one, tuning the prompt, improving the context. That matters. But it's not where most production problems show up.

The real work is system design. Where the agent starts and where it stops. What it can touch and what's out of its reach. Not because you don't trust the model, but because reliable systems have clear boundaries regardless of who's doing the executing.

An agent without clear boundary design isn't an autonomous agent. It's a liability with good documentation.