Context Engineering

Everyone agrees the agent should see less. The question is who decided, and whether you can read it first.
Explore with us
Context Engineering
The argument about whether to shape an agent's context is over. Progressive disclosure, semantic tool search, runtime synthesis, code mode -- every serious approach now shapes what the agent sees, and several publish good numbers for it. We are not going to pretend otherwise, and we did not invent the practice: Harrison Chase named it in June 2025, and Anthropic gave it rigour that September. Both of their definitions treat tools as first-class, which is what makes this an integration problem rather than a prompting one.
So the live disagreement is not whether the agent should see less. It is whether the shaping is something a human approved, or something a system decided while nobody was looking. Imperative approaches decide at runtime -- in an optimizer, a retrieval index, a generated program, a synthesized connector. They produce a smaller context and no artifact. You can audit what happened; you cannot approve what could happen, because the surface does not exist until the agent asks. Declarative context engineering decides once, in a file: the tools the agent sees, the fields that come back, the upstream operations behind them, and the credential authorizing exactly those and nothing else. The agent may well have written that file. What matters is that a human saw it in between.

Two ways to make the context smaller

Both columns below produce a smaller context. Only one produces a smaller context you can approve in advance. The difference is not how well the shaping is done -- it is when the decision is fixed, and whether anyone could read it beforehand.
DecisionImperativeDeclarative
Where the shaping lives A runtime optimizer, a retrieval index, generated code, model-side disclosure A spec file
When it is decided Per request, or per observed failure Once, at authoring time
Who decides The system, adaptively A human, reviewably
Visible beforehand Configuration and policy The complete surface
How it changes Silently, as the optimizer learns A diff, in version control
Governance mode Post-hoc detection Pre-execution prevention
Failure mode Yesterday's approved behaviour is not today's Staleness — the file has to be maintained

Why the distinction is structural, not rhetorical

You cannot enumerate, by reading general-purpose code, the complete set of upstream operations a credential will ever authorize. Computed URLs, branching and dynamic dispatch defeat static analysis. Extend that to runtime shaping and the limit gets worse: if the surface is synthesized when the agent asks, there is no artifact to read at all.
This is not an argument that observability is worthless. It is an argument that observability answers a different question. A log tells you what the agent did. A reviewer signing off on a production credential is asking what it is able to do, on its worst day, before the credential is attached. Only an artifact answers that one.

Relocation, not compression

Here is the part that gets missed. Every rival technique operates on the same corpus -- disclosure shows fewer of the same tools, search ranks the same catalog, summarisation shortens the same payload. All of them accept the work as the model's and negotiate how much of it to reveal.
A capability changes what the work is. Auth, pagination, retries, correlation and sequencing are not shown in shorter form; they are not shown at all, because the model no longer performs them.
The workWithout a capabilityWith one
Choosing among N operations The model, every turn The author, once — the surface is the choice
Sequencing three calls The model, three turns plus re-planning Declared steps, one turn
Correlating two responses The model, with both payloads in context A lookup, server-side, never in context
Auth, pagination, retries The model, or generated glue code The engine, declared
Selecting from a 60-field record The model, after receiving all 60 Output mapping, before projection

Why this matters strategically

Compression has a floor. You cannot compress below what the model must still reason over, and the whole field is approaching that floor together, which is why the published figures are converging. Relocation has no floor, because it changes the denominator rather than the ratio.
It is also the honest explanation for why our numbers are not comparable to a disclosure benchmark. The cheapest token is the one you never put in the context window. An 85% catalog reduction and a two-tool surface that never had a catalog are simply different claims about different quantities.

What it pays, in the order the buyer weighs it

One declaration pays on three axes at once, which is what removes the usual trade-off between governance and speed -- the same file that shrinks the context bounds the blast radius. But the axes are not equally weighted, and we have historically led with the wrong one.
AxisWeightWhat it does
Risk Decisive Not a benefit so much as the unblocker. Undeclared fields cannot leak because they never project; the agent never holds the upstream credential; the engine cannot improvise a call the spec does not contain. Without this there is no deployment to optimise.
Velocity Strong N round-trips collapse into one. The opaque vendor hop disappears. Authoring is paid once and amortised across every consumer, rather than re-paid on every run — and one spec projects to REST, MCP and Agent Skills.
Cost Supporting Tool catalogs drop from thousands of tokens per turn to hundreds, and that cost is per turn for the life of the agent. Responses carry a handful of domain fields instead of sixty. Caching removes upstream calls outright.

The honest boundary

Declarative context engineering requires the shape of the task to be known. It is a warm path, and we would rather say so than have it extracted from us. When an agent meets a genuinely novel task there is no approved capability yet, and something has to produce one -- an agent drafts it, a human reviews it, and the cold path becomes warm.
There are other limits worth stating plainly. A capability engine is a hop, so on a single-call passthrough it adds latency; the wins are on multi-call paths, cacheable reads, and wherever a vendor runtime sits in as an opaque hop. Caching trades freshness for speed. And novel reasoning stays with the model, where it belongs.
Nor is any of this an argument against the approaches above. Compaction, note-taking and sub-agents operate inside the agent; a declared capability determines what is ever eligible to enter. Anthropic's own note that runtime exploration is slower than retrieving pre-computed data points the same way: a declared capability is the pre-computed half of that hybrid, with governance attached.

MCP is necessary, and radically insufficient

The Model Context Protocol won. It is the de-facto wire format for exposing tools to agents, it is named in Anthropic's own definition of context engineering as part of the context state to manage, and if you are putting tools in front of a model in 2026 you are speaking it. None of what follows disputes any of that.
The problem is what the sentence “we have an MCP server” is being asked to carry. It is routinely offered as evidence that an integration story is finished, and it is not -- in the same way that “we have HTTP” was never an API strategy. MCP answers how an agent discovers and calls a tool. It is deliberately, correctly silent on which tools that agent should have been given. A transport that also specified which tools belonged in your enterprise would be a worse transport. But the silence is a governance vacuum, and something has to fill it before a security reviewer will let an agent near a system of record.

Six questions, and where MCP stands on each

The first question is what MCP is for, and it answers it well. The other five decide whether the thing ever reaches production.
QuestionAnswered by MCP?Who has to answer it
How does an agent discover and call a tool? Yes — this is precisely what the protocol is for, and it does it well. MCP. Settled.
Which tools should this agent have? No. Whoever authored the surface. If nobody authored it, the answer defaults to ‘everything the upstream API can do.’
What shape should the response be? No. The tool implementation. A vendor's sixty-field object arrives whole unless something shaped it.
Which upstream operations may this credential reach? No. The credential's scope, which is usually far broader than the tools on offer — and invisible from the protocol.
Who approved this surface, and when? No. Nobody, unless an artifact exists for them to have approved.
How is any of it reviewed before it runs? No. This is the question the whole buying decision turns on, and the protocol has nothing to say about it.

What the silence gets mistaken for

Each of these is a real conversation, and in each one a protocol-level fact is doing work it cannot do.
The claimWhy it does not follow
“We exposed our API as MCP, so agents can use it.” A 1:1 export turns every endpoint into a tool. The agent now chooses among two hundred options every turn, and the catalog is re-sent on each one. Faithful, and the wrong size for a task.
“The MCP server is authenticated, so it is governed.” Authentication establishes who is calling. It says nothing about what that caller may reach once admitted — and scopes declared per adapter admit a token to every tool the adapter exposes.
“We can see every tool call, so we have oversight.” Observability is a record of what happened. It cannot tell you what could happen, which is the question a reviewer is actually asking before signing.
“The vendor operates the server, so the vendor is responsible.” The credential is yours, the blast radius is yours, and the tool descriptions your model reads as instructions are fetched from someone else's server at connect time.
“A smarter model will navigate two hundred tools fine.” Possibly, and it would not help. Trust is not a model-capability problem: a smarter model does not make an over-broad credential acceptable, and no amount of intelligence tells a reviewer what the agent is able to reach.
“The protocol keeps revising, so this is all unstable.” The opposite, if the surface is declared. The 2026-07-28 revision was the largest break in MCP's history and required no recertification of existing capabilities, because a capability is declared against a stable schema and projected to whatever the protocol currently is.

What fills the gap

The five unanswered questions have one thing in common: they are all answerable from an artifact, and unanswerable without one.
The question MCP leaves openWhat a declared capability contributes
Which tools should this agent have? The capability exposes the operations its author wrote. The surface is the selection, made once and reviewable, rather than inherited from whatever the upstream API happens to offer.
What shape should the response be? Typed output parameters select the fields the task declared. An undeclared field is not shortened — it is absent, because it is never projected.
Which upstream operations may this credential reach? The agent calls the capability; the capability calls upstream. Two blast radii, no shared secret, and the reachable set is written down.
Who approved this surface, and when? It is a file in version control. The approval is a commit, and a change to what the agent can reach arrives as a diff someone reviews.
How is it reviewed before it runs? The spec is data before it is behaviour, so it can be linted in the IDE, gated in CI, and read end-to-end before a credential is attached.

Someone has to build the trusted server

None of this is a reading of the protocol against its authors' intent. The specification says plainly that it cannot enforce its own security principles at the protocol level, and that tool descriptions should be considered untrusted unless they come from a trusted server. That is an instruction, and it is addressed to us rather than to the protocol: someone has to build the trusted server.
We are not claiming a perimeter. A capability engine closes some risk classes by construction and leaves others open, and we would rather publish that list than improvise coverage in a sales call. What it does close, it closes structurally — because there is a file to read.
MCP is how the agent reaches the tool. What the tool is, what it returns, what it may touch, and who agreed to that — those are still yours to declare.

Start small, scale smart

Everyone else makes the context smaller. We make most of it unnecessary.
Chart your next course
API Reusability