An agent with access to your data can find things. An agent with context knows what those things mean inside your operation: what the task is for, where the workflow currently stands, which definitions apply, which source is authoritative and current, what it is permitted to see and do, which actions need approval, and how the organization decided last time. Retrieval supplies material; context is the designed layer of state, provenance, permissions and decision rules around it. Without that layer, more access produces more confident mistakes.
Consider an illustrative bid team. Their new agent is connected to the shared drive, the CRM, the policy library and three years of past proposals. Asked to draft a response, it produces a fluent, well-referenced document in minutes. It quotes a delivery commitment from a 2023 bid the firm no longer offers, cites a security policy that was replaced in the spring, and promises a discount only a director may approve. Nothing failed technically. Every retrieval succeeded.
The agent had access. It did not have context. This is the gap that decides whether an enterprise agent can be trusted with a step in a real workflow, and it is not closed by connecting more sources. It is closed by designing what the agent knows about the operation before it acts.
What Is Enterprise Context?
Enterprise context is the operational understanding an agent needs to act correctly inside a specific organization: the purpose of the task, the current state of the workflow and its dependencies, the business definitions in use, which sources are authoritative and current, what the agent is permitted to see and do, which decisions require approval, and how comparable situations were decided before. It is distinct from the raw material in documents and systems; it is what makes that material usable for a consequential action.
Anthropic describes context engineering as the successor to prompt engineering: the question is no longer how to phrase an instruction but "what configuration of context is most likely to generate our model's desired behavior". In an enterprise, that configuration is mostly not text the agent reads. It is the state, permissions and rules the surrounding system enforces.
Access, Retrieval, Memory and Context Are Different Things
Data access is a connection. Retrieval finds relevant passages for a request. Memory carries information across steps and sessions. Context is the designed layer that tells the agent what a passage means, whether it is current, who may use it, and what may be done with it. Each layer is necessary; none of the first three substitutes for the fourth.
| Layer | What it gives the agent | What it cannot tell the agent on its own |
|---|---|---|
| Data access | A connection to a drive, database, API or application | Which of the thousands of reachable items matter for this task |
| Retrieval (RAG) | Passages ranked as relevant to the request | Whether a passage is current, superseded, in scope or permitted for this requester |
| Memory | Facts and history carried across steps and sessions | Which remembered facts are still true, and which were decisions rather than facts |
| Operational context | Workflow state, definitions, provenance, permissions, approval rules, precedent | Nothing it was not designed to hold: it is only as good as the sources and rules behind it |
Retrieval-augmented generation, introduced by Lewis et al. (2020), remains the right way to bring large, changing bodies of knowledge to a model at the moment of use. The design mistake is treating the retrieved passage as the answer rather than as evidence that still has to pass a check: is it the current version, is the requester entitled to it, and does acting on it need a human decision.
Seven Things an Agent Must Know Before It Acts
- Purpose and success criteria. What the step is for and what an acceptable outcome looks like, so the agent can tell a complete result from a plausible one.
- Workflow state and dependencies. Where the case stands, what has already been decided upstream, what is waiting downstream, and what must not be changed once a later step has consumed it.
- Business definitions. What "active customer", "approved supplier" or "delivery date" means here, not in general. Most enterprise terms have a local definition and an owner.
- Provenance, version and freshness. Where each input came from, which version it is, when it was last confirmed, and which source wins when two disagree.
- Roles, permissions and boundaries. What the agent may read and write, on whose behalf, and which systems are out of scope for this task even if technically reachable.
- Policies and approval requirements. Which actions are the agent's to take, which need a named approver, and which are never automated.
- Decisions, exceptions and outcomes. How the organization handled comparable cases, which exceptions were granted and why, and what happened afterwards.
And one rule for when the list is not satisfied: if information is missing, stale or contradictory, the agent should stop and escalate with what it found, not fill the gap. A reliable agent is one whose refusals are as well designed as its answers.
Why Storing More Information Is Not Automatically Better
Because attention is a finite budget, not a warehouse. Anthropic notes that a transformer forms nΒ² pairwise relationships across n tokens, so every additional token dilutes the model's ability to track the ones that matter. Liu et al. found that model performance follows a U-shaped curve by position: strongest for information at the start or end of a long input, weakest in the middle, even for models built for long context. Bulk context also widens the attack and leakage surface. Curated, permissioned, labelled context beats more context.
The security consequences are documented. The OWASP Top 10 for LLM Applications (2025) lists Excessive Agency as an agent that has been given more functionality, permissions or autonomy than its task needs, so that a manipulated or simply mistaken output becomes a damaging action; and Vector and Embedding Weaknesses covering poisoned retrieval stores and cross-tenant leakage in RAG pipelines. Indirect prompt injection, where instructions hidden in a retrieved document redirect the agent, is exactly a context problem: the agent could not tell evidence from instruction.
The practical consequence for a CIO or operations leader: the question to ask a vendor or a team is not "how much can the agent see?" but "how does the agent know which of what it sees is current, permitted and authoritative for this step?"
Example: An RFP-Response Agent That Has to Tell Five Things Apart
Illustrative scenario. It reflects the structure of a proposal workflow, not a specific client's data or results. World AI X's own AI Proposal Intelligence case describes a deployed version of this workflow.
A tender arrives. The agent's job is to build the compliance matrix and draft a response from approved material. To do that reliably it has to distinguish five kinds of input that look alike on a shared drive:
| Input | How the agent must treat it |
|---|---|
| Current tender requirements | The authoritative scope for this response. Every requirement becomes a row in the matrix; nothing else defines what "compliant" means. |
| Approved company capabilities | The only statements the draft may assert as fact, drawn from a governed library with an owner and a review date. |
| An older proposal | Evidence of how the firm has answered before, useful for structure and tone, never a source of current claims without re-approval. |
| A superseded policy | Must be recognized as retired by version and effective date, and excluded even if it is the most semantically similar passage. |
| A commercial commitment | Pricing, delivery dates and liabilities are drafted as placeholders and routed to the named approver; the agent never asserts them. |
Notice that none of these distinctions is made by retrieval quality. A better embedding model will find the superseded policy faster. The distinctions come from metadata the system maintains (status, version, owner, effective date), from permissions that scope what the agent may assert, and from an approval rule that turns certain outputs into requests rather than statements.
The Context Checklist: Before an Agent Takes a Consequential Action
Four questions, asked of every input the agent will rely on and every action it may take. If any answer is "no" or "unknown", the step is not ready for autonomy.
βThe task purpose and success criteria are stated, not inferred
βThe current workflow state and upstream decisions are readable by the agent
βLocal business definitions exist for the terms the step depends on
βEach source carries a version, owner and effective or review date
βSuperseded and draft material is marked and excluded by default
βA precedence rule says which source wins when two disagree
βThe agent acts under a named identity with least-privilege access
βRead, write and out-of-scope systems are defined per step, not per agent
βConsequential actions have a named approver; some are never automated
βEvery output can be traced to the sources and versions it used
βDecisions, exceptions and approvals are logged with a reason
βMissing or contradictory information triggers escalation, and that is logged too
How Do You Know the Agent Used the Right Context?
Test the context, not just the answer. Build evaluation cases that include stale, superseded and out-of-scope material alongside the correct sources, and score whether the agent chose, cited and respected the right ones. Measure escalation behaviour when information is missing or contradictory. And review accepted business outcomes, not successful tool calls; an agent that completed every step and produced an unusable draft did not succeed.
- Provenance coverage. Share of consequential statements in an output that trace to an approved, current source.
- Distractor tests. Representative tasks seeded with retired policies and old proposals; the pass condition is that they are recognized, not merely avoided by luck.
- Escalation rate and quality. How often the agent stops when it should, and whether what it hands over is enough for a person to decide quickly.
- Permission conformance. Zero reads or writes outside the step's defined scope, verified from logs rather than assumed from configuration.
These measures belong in the business case before the build, not in a post-mortem after it. They are the difference between an agent you can audit and one you can only observe.
Where This Sits in World AI OS
In World AI OS, this layer is Brain, the operational context layer. It does three things that map directly onto the checklist above. Context integration connects structured and unstructured information across documents, databases, applications, APIs and operational systems, leaving the information in its existing systems while giving the OS the context it needs to work across them. The operational graph structures that environment around relationships rather than isolated records: workflows to tasks, people to decisions, systems to data, assets to events, actions to outcomes. Operational memory preserves the evidence, assumptions, decisions, changes and outcomes created as work moves through Studio, Factory and production, so each new transformation starts with more context than the last.
The point of a shared context layer is that Studio, which redesigns the workflow, and Factory, which builds and runs it, work from the same evolving model of the operation instead of each rebuilding its own. How much of an operation is modelled, and how quickly, is scoped in each engagement; the architecture is designed so that context accumulates rather than being reconstructed per project.
Frequently Asked Questions
What is enterprise context for an AI agent?
The operational understanding an agent needs to act correctly inside a specific organization: the purpose of the task, the current state of the workflow, the business definitions in use, which sources are authoritative and current, what the agent is permitted to see and do, which decisions need approval, and how similar situations were handled before. Access to documents and systems supplies raw material; context tells the agent what that material means and what it may do with it.
Is retrieval-augmented generation (RAG) enough to give an agent context?
No, but it is a necessary part of the answer. Retrieval finds relevant passages at the moment of the request. It does not, by itself, know whether a passage is current or superseded, whether the requester is allowed to see it, whether acting on it needs approval, or what happened the last time the organization faced the same situation. Those have to be designed around the retrieval layer.
Why is giving an agent more information not automatically better?
Language models have a finite attention budget: the more tokens in the context, the thinner attention is spread, and research shows recall degrades for information buried in the middle of long inputs. Unfiltered context also widens the surface for indirect prompt injection and for retrieving data the requester should not see. Curated, permissioned, well-labelled context outperforms bulk context.
How do you evaluate whether an agent used the right context?
Trace every consequential output back to the sources, versions and permissions it relied on; run representative tasks that include stale, superseded and out-of-scope material to see whether the agent recognizes them; measure how often it escalates when information is missing or contradictory rather than guessing; and review accepted outcomes, not just successful tool calls.