Discovery World AI OS Case Studies Blog Company
Book a Call β†’
Appearance
πŸŒ™ β˜€οΈ
Language
EN FR AR
Share
Platform & Architecture Β· Series 10 of 12

Why Enterprise AI Agents Need Context, Not Just Access to Data

What an agent has to understand about an operation before it can act inside it

Most enterprise agents are wired to everything and understand very little. They can read the documents and call the systems, and they still cannot tell a current requirement from a superseded one, or a suggestion from a commitment that needs sign-off. This article defines enterprise context, separates it from data access, retrieval and memory, and gives you a checklist for what must be in place before an agent takes a consequential action.

WX
AuthorWorld AI X EditorialWorld AI X
Published
Updated
Reading time9 min
Download PDF
The short answer

An agent with access to your data can find things. An agent with context knows what those things mean inside your operation: what the task is for, where the workflow currently stands, which definitions apply, which source is authoritative and current, what it is permitted to see and do, which actions need approval, and how the organization decided last time. Retrieval supplies material; context is the designed layer of state, provenance, permissions and decision rules around it. Without that layer, more access produces more confident mistakes.

U-curveModel recall drops for information in the middle of long inputs β€” Liu et al., TACL 2024
nΒ²Attention relationships grow with the square of tokens in context: a finite budget β€” Anthropic, 2025
3Root causes of Excessive Agency: excess functionality, permissions, autonomy β€” OWASP LLM06:2025
LLM08Vector & embedding weaknesses, incl. cross-tenant leakage in RAG β€” OWASP Top 10 for LLMs, 2025

Consider an illustrative bid team. Their new agent is connected to the shared drive, the CRM, the policy library and three years of past proposals. Asked to draft a response, it produces a fluent, well-referenced document in minutes. It quotes a delivery commitment from a 2023 bid the firm no longer offers, cites a security policy that was replaced in the spring, and promises a discount only a director may approve. Nothing failed technically. Every retrieval succeeded.

The agent had access. It did not have context. This is the gap that decides whether an enterprise agent can be trusted with a step in a real workflow, and it is not closed by connecting more sources. It is closed by designing what the agent knows about the operation before it acts.

01 β€” Definition

What Is Enterprise Context?

Enterprise context is the operational understanding an agent needs to act correctly inside a specific organization: the purpose of the task, the current state of the workflow and its dependencies, the business definitions in use, which sources are authoritative and current, what the agent is permitted to see and do, which decisions require approval, and how comparable situations were decided before. It is distinct from the raw material in documents and systems; it is what makes that material usable for a consequential action.

Anthropic describes context engineering as the successor to prompt engineering: the question is no longer how to phrase an instruction but "what configuration of context is most likely to generate our model's desired behavior". In an enterprise, that configuration is mostly not text the agent reads. It is the state, permissions and rules the surrounding system enforces.

02 β€” Four layers

Access, Retrieval, Memory and Context Are Different Things

Data access is a connection. Retrieval finds relevant passages for a request. Memory carries information across steps and sessions. Context is the designed layer that tells the agent what a passage means, whether it is current, who may use it, and what may be done with it. Each layer is necessary; none of the first three substitutes for the fourth.

LayerWhat it gives the agentWhat it cannot tell the agent on its own
Data accessA connection to a drive, database, API or applicationWhich of the thousands of reachable items matter for this task
Retrieval (RAG)Passages ranked as relevant to the requestWhether a passage is current, superseded, in scope or permitted for this requester
MemoryFacts and history carried across steps and sessionsWhich remembered facts are still true, and which were decisions rather than facts
Operational contextWorkflow state, definitions, provenance, permissions, approval rules, precedentNothing it was not designed to hold: it is only as good as the sources and rules behind it

Retrieval-augmented generation, introduced by Lewis et al. (2020), remains the right way to bring large, changing bodies of knowledge to a model at the moment of use. The design mistake is treating the retrieved passage as the answer rather than as evidence that still has to pass a check: is it the current version, is the requester entitled to it, and does acting on it need a human decision.

03 β€” The content of context

Seven Things an Agent Must Know Before It Acts

  • Purpose and success criteria. What the step is for and what an acceptable outcome looks like, so the agent can tell a complete result from a plausible one.
  • Workflow state and dependencies. Where the case stands, what has already been decided upstream, what is waiting downstream, and what must not be changed once a later step has consumed it.
  • Business definitions. What "active customer", "approved supplier" or "delivery date" means here, not in general. Most enterprise terms have a local definition and an owner.
  • Provenance, version and freshness. Where each input came from, which version it is, when it was last confirmed, and which source wins when two disagree.
  • Roles, permissions and boundaries. What the agent may read and write, on whose behalf, and which systems are out of scope for this task even if technically reachable.
  • Policies and approval requirements. Which actions are the agent's to take, which need a named approver, and which are never automated.
  • Decisions, exceptions and outcomes. How the organization handled comparable cases, which exceptions were granted and why, and what happened afterwards.

And one rule for when the list is not satisfied: if information is missing, stale or contradictory, the agent should stop and escalate with what it found, not fill the gap. A reliable agent is one whose refusals are as well designed as its answers.

04 β€” The limit

Why Storing More Information Is Not Automatically Better

Because attention is a finite budget, not a warehouse. Anthropic notes that a transformer forms nΒ² pairwise relationships across n tokens, so every additional token dilutes the model's ability to track the ones that matter. Liu et al. found that model performance follows a U-shaped curve by position: strongest for information at the start or end of a long input, weakest in the middle, even for models built for long context. Bulk context also widens the attack and leakage surface. Curated, permissioned, labelled context beats more context.

The security consequences are documented. The OWASP Top 10 for LLM Applications (2025) lists Excessive Agency as an agent that has been given more functionality, permissions or autonomy than its task needs, so that a manipulated or simply mistaken output becomes a damaging action; and Vector and Embedding Weaknesses covering poisoned retrieval stores and cross-tenant leakage in RAG pipelines. Indirect prompt injection, where instructions hidden in a retrieved document redirect the agent, is exactly a context problem: the agent could not tell evidence from instruction.

The practical consequence for a CIO or operations leader: the question to ask a vendor or a team is not "how much can the agent see?" but "how does the agent know which of what it sees is current, permitted and authoritative for this step?"

05 β€” Worked example

Example: An RFP-Response Agent That Has to Tell Five Things Apart

Illustrative scenario. It reflects the structure of a proposal workflow, not a specific client's data or results. World AI X's own AI Proposal Intelligence case describes a deployed version of this workflow.

A tender arrives. The agent's job is to build the compliance matrix and draft a response from approved material. To do that reliably it has to distinguish five kinds of input that look alike on a shared drive:

InputHow the agent must treat it
Current tender requirementsThe authoritative scope for this response. Every requirement becomes a row in the matrix; nothing else defines what "compliant" means.
Approved company capabilitiesThe only statements the draft may assert as fact, drawn from a governed library with an owner and a review date.
An older proposalEvidence of how the firm has answered before, useful for structure and tone, never a source of current claims without re-approval.
A superseded policyMust be recognized as retired by version and effective date, and excluded even if it is the most semantically similar passage.
A commercial commitmentPricing, delivery dates and liabilities are drafted as placeholders and routed to the named approver; the agent never asserts them.

Notice that none of these distinctions is made by retrieval quality. A better embedding model will find the superseded policy faster. The distinctions come from metadata the system maintains (status, version, owner, effective date), from permissions that scope what the agent may assert, and from an approval rule that turns certain outputs into requests rather than statements.

06 β€” The checklist

The Context Checklist: Before an Agent Takes a Consequential Action

Four questions, asked of every input the agent will rely on and every action it may take. If any answer is "no" or "unknown", the step is not ready for autonomy.

Available

βœ“The task purpose and success criteria are stated, not inferred

βœ“The current workflow state and upstream decisions are readable by the agent

βœ“Local business definitions exist for the terms the step depends on

Current

βœ“Each source carries a version, owner and effective or review date

βœ“Superseded and draft material is marked and excluded by default

βœ“A precedence rule says which source wins when two disagree

Permissioned

βœ“The agent acts under a named identity with least-privilege access

βœ“Read, write and out-of-scope systems are defined per step, not per agent

βœ“Consequential actions have a named approver; some are never automated

Traceable

βœ“Every output can be traced to the sources and versions it used

βœ“Decisions, exceptions and approvals are logged with a reason

βœ“Missing or contradictory information triggers escalation, and that is logged too

07 β€” Evaluation

How Do You Know the Agent Used the Right Context?

Test the context, not just the answer. Build evaluation cases that include stale, superseded and out-of-scope material alongside the correct sources, and score whether the agent chose, cited and respected the right ones. Measure escalation behaviour when information is missing or contradictory. And review accepted business outcomes, not successful tool calls; an agent that completed every step and produced an unusable draft did not succeed.

  • Provenance coverage. Share of consequential statements in an output that trace to an approved, current source.
  • Distractor tests. Representative tasks seeded with retired policies and old proposals; the pass condition is that they are recognized, not merely avoided by luck.
  • Escalation rate and quality. How often the agent stops when it should, and whether what it hands over is enough for a person to decide quickly.
  • Permission conformance. Zero reads or writes outside the step's defined scope, verified from logs rather than assumed from configuration.

These measures belong in the business case before the build, not in a post-mortem after it. They are the difference between an agent you can audit and one you can only observe.

08 β€” In World AI OS

Where This Sits in World AI OS

In World AI OS, this layer is Brain, the operational context layer. It does three things that map directly onto the checklist above. Context integration connects structured and unstructured information across documents, databases, applications, APIs and operational systems, leaving the information in its existing systems while giving the OS the context it needs to work across them. The operational graph structures that environment around relationships rather than isolated records: workflows to tasks, people to decisions, systems to data, assets to events, actions to outcomes. Operational memory preserves the evidence, assumptions, decisions, changes and outcomes created as work moves through Studio, Factory and production, so each new transformation starts with more context than the last.

The point of a shared context layer is that Studio, which redesigns the workflow, and Factory, which builds and runs it, work from the same evolving model of the operation instead of each rebuilding its own. How much of an operation is modelled, and how quickly, is scoped in each engagement; the architecture is designed so that context accumulates rather than being reconstructed per project.

09 β€” FAQ

Frequently Asked Questions

What is enterprise context for an AI agent?

The operational understanding an agent needs to act correctly inside a specific organization: the purpose of the task, the current state of the workflow, the business definitions in use, which sources are authoritative and current, what the agent is permitted to see and do, which decisions need approval, and how similar situations were handled before. Access to documents and systems supplies raw material; context tells the agent what that material means and what it may do with it.

Is retrieval-augmented generation (RAG) enough to give an agent context?

No, but it is a necessary part of the answer. Retrieval finds relevant passages at the moment of the request. It does not, by itself, know whether a passage is current or superseded, whether the requester is allowed to see it, whether acting on it needs approval, or what happened the last time the organization faced the same situation. Those have to be designed around the retrieval layer.

Why is giving an agent more information not automatically better?

Language models have a finite attention budget: the more tokens in the context, the thinner attention is spread, and research shows recall degrades for information buried in the middle of long inputs. Unfiltered context also widens the surface for indirect prompt injection and for retrieving data the requester should not see. Curated, permissioned, well-labelled context outperforms bulk context.

How do you evaluate whether an agent used the right context?

Trace every consequential output back to the sources, versions and permissions it relied on; run representative tasks that include stale, superseded and out-of-scope material to see whether the agent recognizes them; measure how often it escalates when information is missing or contradictory rather than guessing; and review accepted outcomes, not just successful tool calls.

Brain Β· The operational context layer

Give your agents the operation, not just the files.

See how Brain connects, structures and remembers the context every AI-native workflow depends on.

β–ΆSources6 references
Anthropic (2025). Effective context engineering for AI agents: context as a finite resource; the nΒ² attention budget; curating the minimal high-signal token set.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F. & Liang, P. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the ACL, 12, 157–173: U-shaped recall by position, including for long-context models.
OWASP GenAI Security Project (2025). LLM06:2025 Excessive Agency: excessive functionality, permissions and autonomy as root causes.
OWASP GenAI Security Project (2025). LLM08:2025 Vector and Embedding Weaknesses: poisoned retrieval stores, cross-tenant leakage and injection through RAG.
Lewis, P. et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020: the retrieval-augmented pattern this article builds on rather than replaces.
World AI X. Brain β€” the operational context layer (context integration, operational graph, operational memory) and AI Proposal Intelligence case study.
Found this useful?
Share it with whoever is wiring an agent to the shared drive.
LinkedIn
X
Copy link