Prompt Engineering Was the Easy Part. Context Engineering Is the Job.
The term that replaced "better prompts"
For about two years, "prompt engineering" was the whole conversation. Pick the right words, give a few examples, specify the output format, and the model does what you want. That still matters for single-shot tasks — classify this ticket, summarize this email — where the entire problem fits in one well-crafted instruction.
It stops being the relevant skill the moment you're building an agent instead of calling a prompt. An agent that reads files, calls tools, retrieves documents, and runs for dozens of turns isn't operating on a prompt you wrote once. It's operating on a context window that gets assembled and re-assembled at every single step: system instructions, conversation history, tool definitions, retrieved chunks, the results of the last five tool calls, and whatever scratch notes the agent left itself. Getting that assembly right — deciding what's in the window, what's been summarized, what's been dropped, and what's been moved to external storage — is a different discipline. Anthropic's applied AI team formalized this as "context engineering" in September 2025, defining it as the set of strategies for curating and maintaining the optimal set of tokens during inference, and framed it explicitly as the evolution of prompt engineering rather than a replacement vocabulary for the same thing. The term itself was popularized a few months earlier, in June 2025, by Shopify's Tobi Lütke and AI researcher Andrej Karpathy, both making the same observation from different angles: the hard part of building with LLMs had quietly shifted from phrasing to architecture.
If you're a business evaluating a vendor who's building you an AI agent, or an engineering lead deciding whether to build one in-house, this distinction is not academic. It's the difference between a demo that works and a system that holds up in week three of production use.
Why more context doesn't mean better context
The intuitive fix for an agent that's missing information is to hand it more: paste in the whole document, keep the entire conversation history, give it every tool you can think of. This is where most naive agent builds go wrong, because LLM performance does not scale cleanly with context length.
Chroma Research tested this directly in mid-2025, running 18 models through roughly 194,000 LLM calls designed to isolate the effect of input length on retrieval and reasoning accuracy, independent of task difficulty. Every model they tested degraded as input length grew — and the degradation showed up well before the context window was anywhere near full. The effect wasn't just "the model ran out of room." It was driven by how long the input was, how many distractor items sat near the relevant fact, how closely the target information matched the surrounding text, and how the content was structured. Anthropic's own engineering writeup names the same phenomenon "context rot": as token count grows, accuracy and instruction-following degrade, which is why curating what's in the context window matters as much as how much room the window has.
This is counterintuitive to anyone who's been tracking the context-window arms race — 1M-token windows get marketed as a capability upgrade, and in a narrow sense they are. But a bigger window doesn't fix a model's tendency to lose track of information buried in the middle of a long input, sometimes called "lost in the middle" behavior. It just means you can bury more things. For an agent doing real work — reading code, searching documents, executing a multi-step task — dumping everything into context because you can is a reliability bug, not a feature.
The three techniques that actually hold up
The teams building agents that stay reliable past turn ten aren't writing cleverer system prompts. They're managing the context window as a finite, actively-curated resource, using a small number of recurring patterns:
Compaction. When a conversation or task history gets long, the agent (or the harness around it) summarizes the older portion into a compact, high-fidelity representation and drops the verbose original. This is why tools like Claude Code periodically compact session history rather than letting it grow unbounded — the agent keeps working with the gist of turns 1 through 40 instead of carrying every raw tool output forward.
Structured external memory. Instead of keeping everything in the live context window, the agent writes progress, decisions, and intermediate findings to external storage — a scratch file, a task list, a notes document — and reads back only what's relevant to the current step. This is structured note-taking: it lets an agent work across a task that's far longer than any single context window could hold, because the working memory lives outside the window, not inside it.
Sub-agent isolation. Rather than one agent carrying the full history of a complex task, a coordinating agent delegates focused pieces of work to sub-agents that start with a clean, narrow context — just what they need for their slice of the problem — and return a condensed summary rather than their full working history. The parent agent's context stays small even as the total amount of work done gets large, because the exploration and dead ends happened in a context that got thrown away afterward.
These three compose. A production coding or research agent typically uses just-in-time retrieval to avoid stuffing context up front, compaction to keep long-running sessions bounded, and sub-agents to isolate deep or exploratory work from the main thread. None of them are about finding better words for the system prompt. They're about information architecture.
What this means if you're buying, not building
If a vendor pitches you an "AI agent" for your business — triaging support tickets, processing applications, managing a workflow — the question worth asking isn't "what model does it use." It's "what happens to this agent's context after fifty tickets, or on a conversation that runs long." A vendor who can describe how they bound context growth, what they summarize versus retain, and how they isolate sub-tasks has actually built an agent. A vendor who shrugs and says "the model has a big context window now" has built a demo that will degrade quietly in production, in ways that are hard to detect until a customer notices the agent forgot something it was told three steps ago.
This also changes what to budget for. A system prompt is a one-time cost. Context engineering is an ongoing architectural decision that shows up in latency, token cost, and reliability simultaneously — and it's the part of an agent build that actually requires engineering judgment, not prompt iteration.
The takeaway
A better-written prompt still helps for simple, single-shot tasks. For anything that runs as an agent — multi-turn, multi-tool, operating over real data — the reliability ceiling is set by what's in the context window at each step, not by how well the instructions are phrased. Ask any vendor building you an agent how they manage that window as work accumulates; if they don't have a concrete answer, assume the demo won't survive contact with real usage.
Have a system like this in mind?
Get a scoped plan ↗