Pedro Sousa
‹ ← Articles

Article · Blog

The context you send to the agent is the product. Not the prompt.

Most teams treat context engineering as an implementation detail. In practice, it's the most important architectural decision in any LLM-based system. When it breaks in production, the model is not the problem.

The context you send to the agent is the product. Not the prompt.

There's a specific moment when you realize the problem isn't the model.

You tested locally. It worked. The agent was making reasonable decisions, following the flow, generating coherent outputs. You pushed to production. Two weeks later, the system starts making decisions that don't make sense. Not always. Not predictably. But enough to start accumulating silent errors that nobody notices right away.

You start debugging the prompt. Adjust the tone. Rewrite the instructions. Add examples. The behavior improves a little. But the problem comes back.

I went through this. And the diagnosis took longer than it should have because I was looking in the wrong place.

The problem wasn't the prompt. It was the context I was injecting alongside it.


What context engineering actually means in production

Context engineering has become a buzzword. But before it was a buzzword, it was a concrete engineering problem: what you send to the model determines what the model will do.

That sounds obvious. And it is. But there's a consequence most teams ignore: if the context is bad, ambiguous, outdated, or excessive, the model will produce bad outputs deterministically. Not because the model is weak. Because you gave it bad information to work with.

A production agent isn't "thinking" in a vacuum. It's responding to what's in the context window. Every decision it makes is a function of what you put there.

When the agent was making wrong decisions in the system I mentioned, here's what I found: the context reaching the model included state events that were already stale. The context ingestion architecture had no expiration mechanism whatsoever. Events from hours ago coexisted with recent events, with no distinction of temporal relevance.

The model tried to reconcile everything. And sometimes it reconciled things in ways that made internal sense but were completely wrong for the actual state of the system.


The problem nobody calls by the right name

I call this Accumulated Context Poisoning.

It's not about the context being wrong from the start. It's about it becoming wrong over time, without any mechanism in the system noticing or correcting it.

It's a particularly tricky problem because:

  1. In tests, the context is always clean and current.
  2. In production, the context accumulates history, transient states, partial information, and events that have already lost relevance.
  3. The model doesn't know how to distinguish what's current from what's noise. It treats everything with the same weight, unless you explicitly instruct it otherwise.

Most teams solve this by adding more instructions to the prompt. "Prioritize the most recent events." "Ignore contradictory information." That softens the problem. It doesn't fix it.

The real solution lives in the ingestion layer, not the instruction layer.


What actually works in practice

Context needs to be treated as a managed resource, not a string assembled on the fly.

Three principles that changed how I build this:

Context has an expiration. Every piece of injected data should have an implicit or explicit TTL. If a state event is older than X minutes, it needs to be discarded or flagged as historical before entering the context window. The model can use historical data. But it needs to know that it's historical.

Context has hierarchy. Not everything that's relevant carries the same weight. Current state information takes priority over past event logs. Explicit user intent takes priority over inferred past behavior. If you don't structure that hierarchy during context assembly, the model will try to infer it. And it will get it wrong in cases you never tested.

Context has an ambiguity cost. Every contradiction inside the context window is a potential failure point. Before injecting any piece of data, the question isn't "is this relevant?" but "does this create ambiguity with anything else that's already here?"


The architectural decision most teams delay for too long

Assembling context inline, inside the same function that calls the model, works in the early cycles. It's fast to implement and easy to trace.

But when the system starts having multiple agents, multiple data sources, and states that change frequently, that pattern creates a problem that shows up gradually: the logic for selecting and filtering context gets scattered. Each model call has its own version of "what's relevant right now." None of them are tested in isolation. None of them are observable as a unit.

The change that made a difference was treating context building as a separate module with a single responsibility: receive the current system state and produce a clean, hierarchical representation with explicit expiration to be injected into the model.

That has a cost. It's more code. It's one more layer to maintain. But it's the layer you'll need to test when the agent starts making wrong decisions in production, and you'll be glad it exists separately.


The trade-off nobody mentions

Richer context isn't always better context.

There's a natural tendency to inject more information because more information feels safer. The model will have everything it needs. Nothing will be missing.

In practice, excessive context degrades response quality in ways that are hard to trace. The model starts paying attention to parts of the context you didn't consider relevant. The signal drowns in the noise. And the output becomes something that's coherent with the injected context but wrong for the actual situation.

The paradox is that reducing context, being more surgical about what you inject, often improves the quality of the agent's decisions. Not because the model got smarter. Because you removed ambiguity.

That goes against engineering instinct. We were trained to think that more data is always better. With LLMs, more data is better only when the right data is in the right place.


What I would do differently

I would have treated the context-building module as a first-class system from the start. With its own unit tests. With its own observability. With logs that show exactly what entered the context window on each call.

Without that, when the agent fails, you're in the dark. You have the wrong output and the prompt. But you have no visibility into what was in the context at that specific moment. And debugging without that is just guessing.

Context observability is not optional in production systems. It's the difference between a bug you can reproduce and a behavior you'll spend weeks trying to explain.


The model isn't the product. The prompt isn't the product. The product is the entire system: how you collect state, how you filter, how you prioritize, how you inject, and what you discard before the model ever sees anything.

When the agent gets it wrong, start with the context. Almost always, that's where the problem is.