SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Explainer 20 mins

Agent Memory on AWS: Retention, Retrieval and Tenant Boundaries

A practical AWS reference design for separating agent conversation history, operational state and long-term memory—with clear retention and tenant boundaries.

The PADISO Team ·

An agent that remembers earlier conversations can feel more useful, but “memory” is not one storage problem. A transcript, a paused task and a durable preference serve different purposes. Combining them in one undifferentiated record makes it harder to answer basic operational questions: what should be retained, what can be retrieved for this request, and which tenant is allowed to see it?

A safer design begins by separating three concepts: conversation history, the record of interaction; operational state, the current facts needed to run a task; and agent memory, selected information intended to improve later interactions. These categories may be implemented with different stores, or with distinct records and access paths in a shared service. The important boundary is their purpose and handling, not the number of databases.

This explainer develops an AWS-oriented reference design from those distinctions. It shows how request identity, tenant scope, retrieval, retention and task execution fit together, then applies the design to a clearly hypothetical support agent. The choices are starting points for architecture review, not a claim that one layout fits every workload.

1. Begin with three definitions

Conversation history is the sequence of messages and relevant interaction events associated with a session or conversation. It helps the agent interpret references such as “that order” or “use the same format as before.” History is contextual evidence, not automatically a verified record of business truth. A user may correct themselves, provide inaccurate details or ask the agent to ignore earlier instructions.

Operational state is the structured, current information required to complete work. It might include a task identifier, the stage of a process, a selected item, a pending action or a result returned by an authoritative business system. State answers “where is this operation now?” rather than “what did the user say?” It should be updated according to the application’s business rules, not silently inferred from a conversation transcript.

Agent memory is a deliberately selected set of information that may be useful beyond the immediate session. Examples include a user’s stated preference for concise summaries or a team’s approved terminology. Memory is a candidate source of context, not an authority that overrides current instructions or a system of record. A remembered preference can be stale, ambiguous or incorrectly attributed.

The terms describe roles, not necessarily separate products. One implementation could store conversation events in an archive, task state in an application database and selected memory in a retrieval service. Another could use a common storage platform while enforcing separate schemas, retention policies and access checks. Either can be sound if the distinctions remain explicit and testable.

AWS documentation describes AgentCore Memory as supporting short-term and long-term memory. For the application design considered here, make retention and tenant filtering explicit decisions. That makes the key architectural work the boundaries around what is stored and retrieved, rather than assuming that a memory feature determines those policies for an application. (Amazon Bedrock AgentCore Memory)

2. Why the distinction changes the architecture

Suppose a user tells an assistant, “I’m working on the north region account,” then asks it to draft a customer update. That sentence belongs in conversation history because it explains the current exchange. If a multi-step workflow is underway, a task record might separately say that the draft is awaiting review. A durable memory might capture a stable preference such as “include a short summary before detailed notes,” if the product has a clear reason and policy to retain it.

Those records have different lifetimes. The immediate reference to a region may stop being useful when the session ends. A task may need to remain available until it is completed, cancelled or expired. A preference may be useful across many sessions but should be editable and removable under the application’s retention rules. Applying one expiration period to all three is convenient to implement and often wrong for the product.

They also have different consequences when wrong. A missing conversation turn can make a response less coherent. A stale task status can cause repeated or skipped work. A memory retrieved for the wrong person can expose information across a tenant boundary. The impact depends on what the agent is permitted to do with the retrieved item, so retrieval and action design must be considered together.

Treating all content as memory also encourages a common failure: every message is stored, then a similarity search is expected to sort out relevance later. Search relevance does not establish authorization, freshness or correctness. It cannot turn a statement into a verified fact, and a high similarity score does not mean that a record belongs in the current request’s context.

A useful first classification asks four questions of each field: Is it needed to understand this exchange? Is it needed to resume a specific operation? Is it intended to influence future exchanges? Which application or person is authoritative for correcting it? A field with multiple answers should be split where possible. For example, retain a user’s original request in the conversation record, while storing a separately validated status in the task record.

3. A reference design: identity, state and service boundaries

The design below keeps the request path understandable without prescribing a particular AWS storage product. It assumes an application receives a user request, establishes trusted actor and tenant context, and then calls separate application services for history, memory and task state. Those services can be deployed using components appropriate to the workload, but each should expose a narrow purpose-specific interface.

  • Request boundary: Accepts the user request and trusted identity context from the application’s authentication path. It must not treat a tenant name or user identifier typed into the prompt as proof of identity.
  • Context service: Loads only the conversation history needed to understand the active exchange and asks the memory service for eligible records using the resolved scope.
  • Memory service: Applies tenant and subject scope, retention eligibility and retrieval rules before returning a bounded set of records.
  • Task service: Reads and updates structured operational state under the application’s business rules. It provides the current task status rather than asking the model to reconstruct it from old messages.
  • Action boundary: Sends any proposed business operation to the system responsible for that operation. It receives a result that can be recorded as an outcome; a model’s statement that an action succeeded is not itself proof of success.

In a practical deployment, these are logical boundaries even if some share a process or storage platform. Make the caller and purpose visible in each interface. A history lookup should not become an unrestricted memory query simply because both return text. A task update should not be smuggled through a general-purpose “save context” operation.

The request flow is intentionally small. Tenant and actor validation comes before any retrieval, and the later services receive the resolved scope rather than deriving it from generated text. If the context is missing or cannot be trusted, the system rejects or routes the request for a controlled recovery path instead of searching broadly.

flowchart TD
  accTitle: Scoped agent context flow
  accDescr: A request is checked for trusted tenant and actor context before conversation history, eligible memory, and task state are assembled for a response. Invalid context is rejected before retrieval.
  A["Request with actor context"] --> B{"Tenant and actor valid?"}
  B -- "No" --> C["Reject before retrieval"]
  B -- "Yes" --> D["Load scoped conversation"]
  D --> E["Retrieve eligible memory"]
  E --> F["Read or update task state"]
  F --> G["Compose response and record outcome"]

The diagram shows the order of checks, not a requirement to retrieve all three kinds of data for every request. A stateless informational question may need no task lookup. A request to resume a workflow may need the task record and a small portion of history, but no long-term preference. Retrieval should be driven by the task’s needs, not by a default to attach everything available.

The boundary also clarifies where to look when a response is wrong. If it used an incorrect preference, inspect memory selection and its origin. If it claimed a workflow was still pending after completion, inspect task-state updates. If it misunderstood a pronoun, inspect conversation context. Without distinct records and traces, all three problems can look like “the model forgot.”

4. Tenant boundaries are retrieval rules, not labels

A tenant identifier is useful only when it is carried from a trusted source into each relevant access decision. Storing a tenant_id field beside a memory record is not enough if a query can omit the filter, use a caller-supplied value or return results before checking their scope. The retrieval path must constrain candidates to the tenant and, where appropriate, to a person, team or application-defined audience.

A helpful design distinction is between tenant scope and subject scope. Tenant scope separates one customer or organizational boundary from another. Subject scope determines which people or roles within that tenant may use an item. A shared team preference and an individual user preference may have the same tenant identifier but should not automatically have the same audience.

Define scope when an item is created, not only when it is retrieved. A memory record can carry fields such as tenant_key, subject_kind, subject_key, record_type, created_at, expires_at, source_ref and review_state. These names are illustrative, not prescribed schemas. The design goal is to make the dimensions required for authorization and lifecycle decisions explicit enough to validate.

At retrieval time, construct the allowed scope from trusted request context and policy. Then require every candidate record to satisfy that scope before it reaches the model. If an item’s scope is absent, malformed or incompatible with the request, fail closed for that item. Do not ask the model to inspect a mixed-tenant result set and decide what it is permitted to see.

There are tradeoffs between physically separating tenant data and using logically partitioned records. Physical separation can make boundaries more legible and reduce the consequences of a faulty filter, but can increase operational complexity as tenant count and data volume grow. Logical separation can simplify shared operations, but depends more heavily on consistent scope enforcement across every write, read, export and deletion path. The appropriate choice depends on the application’s threat model, scale and operational capacity; neither pattern removes the need to test boundary behavior.

For an agent acting on behalf of a user, do not substitute a shared application identity for a clear actor context. The detailed question of delegated AWS identity is separate from memory design; the related discussion of delegation without shared credentials is useful when designing that part of the request path. Here, the essential requirement is that the memory service receives trustworthy scope and does not let the model invent or broaden it.

5. Retention and retrieval are linked decisions

Retention asks how long a record may remain available and what event ends that period. Retrieval asks whether a record is eligible and useful for a particular request. They are connected, but not interchangeable. A record that has not expired may still be irrelevant or unauthorized. A record that appears relevant should still be excluded if its retention has ended or its source is no longer acceptable.

Start by setting a purpose for each record type. Conversation history might exist to support continuity within an active session. Operational state might be kept until a task reaches a terminal outcome and a defined follow-up period passes. A user preference might persist until the user changes or removes it, or until an application-defined expiration. Those are examples of policy choices, not universal retention periods.

For each type, define the lifecycle events that change eligibility: creation, correction, completion, expiration, deletion request, source update and tenant closure. A retention rule that says “keep it for a while” leaves implementation teams without a reliable way to decide when to delete, whether to refresh a record or what to do when the source disappears. State the event, the owner of the decision and the expected effect on retrieval.

Retrieval should apply several tests before assembling context:

  • Scope: Does the item belong to the resolved tenant and permitted subject?
  • Lifecycle: Is it still within its allowed retention period and in an eligible state?
  • Purpose: Is this record type useful for the current task, or should it be excluded?
  • Freshness: Is the item recent or confirmed enough for the decision at hand?
  • Provenance: Can the application identify where the item came from and how it was created or corrected?

These tests are complementary. A fresh item can still be out of scope. A correctly scoped memory may be stale. A record with a credible source may still not be appropriate to include in a response. The application should be able to explain which rule admitted or rejected a candidate without relying on the model’s generated rationale.

Bound the amount of retrieved context as well as its scope. A small set of records selected for a stated purpose is easier to review, debug and correct than an open-ended dump of past interactions. If retrieval returns no eligible memory, that is a valid result. The agent can proceed using the current request or ask the user for missing information rather than silently expanding the search.

6. Worked example: a hypothetical service assistant

Consider a hypothetical B2B support assistant used by several customer organizations. A customer employee asks, “Can you continue the access review we started yesterday and put the results in our usual format?” The assistant needs to resolve which organization and person made the request before retrieving anything. The words “our usual format” are not sufficient to select another customer’s preference, and “the review” is not enough to establish task status.

The request boundary supplies trusted actor and tenant context. The application then looks for an active access-review task within that tenant and checks whether the actor is allowed to resume it. If there is one eligible task, the task service returns structured status such as task_ref, stage, last_updated, owner_ref and pending_input. If there are multiple matching tasks or none, the assistant should ask for clarification or present a scoped choice rather than guessing from a broad history search.

The context service may load the most recent relevant exchange for the selected task, such as the user’s explanation of a requested review window. That conversation text helps resolve “started yesterday,” but it does not decide whether the review has actually begun or completed. The task record does. This distinction prevents a statement like “I finished the review” from being treated as authoritative if the application has not recorded a result.

For “our usual format,” the memory service can retrieve an eligible preference whose subject scope is the customer organization or the individual, depending on how the preference was created. The service checks tenant, audience, lifecycle and purpose before returning it. If the preference is missing or ambiguous, the assistant should use a neutral format or ask. It should not broaden retrieval to other tenants in an attempt to find a likely match.

Suppose the task record says the review is awaiting a list of accounts, while the recent transcript suggests the user supplied that list. The system should reconcile the discrepancy through a defined application path: inspect the source event, validate whether the input was accepted, and update operational state if appropriate. The model can help identify the mismatch, but it should not silently rewrite structured state based on a plausible narrative.

When the task is completed, the application records the result and its source. The conversation remains subject to its own history policy; task state can move to a completed or archived status; and the format preference can remain eligible if its purpose and retention rules allow it. The same interaction therefore produces three different lifecycle outcomes rather than one generic “memory saved” event.

This design costs some simplicity at the interface level: the assistant must call purpose-specific paths and handle missing or conflicting information. In exchange, an operator can correct a stale preference without altering a task, resume work without searching every transcript, and answer which tenant scope was used for a retrieval. For background on broader platform selection, see how enterprise agent platforms differ; the memory boundaries here remain relevant regardless of platform choice.

7. Counterexample: one shared “memory” record

A tempting shortcut is to append every user message, tool result and task update to one tenant-labelled text record. At request time, the application searches that record and sends the most similar passages to the model. This can work in a small demonstration where one user, one task and one retention period are assumed. Those assumptions break as soon as users share an organization, tasks outlive sessions or records need different correction and deletion behavior.

Imagine the shared record contains a transcript where an employee says that a review is complete, a later application result showing it is still pending, and a previous user’s formatting preference. A similarity search for “continue the review” might return all three. Even if the tenant label is correct, the system still has to distinguish verified state from user text, determine which person’s preference is in scope and decide whether each item remains eligible.

The shortcut also obscures failure ownership. If a task repeats an action because the agent read an old transcript, is the fix to tune retrieval, add a timestamp, update the task service or change the action flow? A single blob makes those responsibilities difficult to isolate. Adding labels later helps only if every write and query honors them consistently, and older records can be classified reliably.

A shared store is not inherently a bad choice. The counterexample is sharing purpose, lifecycle and access semantics—not simply using one storage product. A common platform can support separate record types and service interfaces. The application still needs different policies for history, state and memory, and it must apply those policies before context reaches the model.

8. Failure analysis: trace the timeline, not just the answer

When an agent gives an answer based on the wrong context, preserve a trace of the decisions that led to it. A useful operational record identifies the request correlation reference, resolved tenant and subject scope, record categories queried, candidate references admitted or excluded, task-state version read and final application outcome. Avoid turning the trace into a second unbounded transcript; it should help reconstruct decisions while following the application’s own retention policy.

A cross-tenant retrieval failure often starts earlier than the model call. The request may have accepted a tenant value from untrusted text, a service may have defaulted a missing filter, or a cache key may have omitted scope. Investigate the entire chain from identity resolution through cache and storage query. A model instruction to “ignore unrelated tenants” does not repair a retrieval boundary that already exposed their records.

A stale memory failure can arise when the source changes but the memory remains eligible. Record the source and the event that should trigger correction or expiration. If the source cannot be checked at retrieval time, define a conservative freshness rule for that record type, or treat the memory as a hint that must be confirmed before consequential use. Do not present a remembered value as current merely because it was retrieved successfully.

A lost or duplicated task update is a state-management problem, not a prompt-writing problem. A worker may stop after receiving an external result but before recording it, or retry after an uncertain response. Design task transitions so an operator can distinguish “not started,” “in progress,” “outcome unknown” and “completed.” Record external outcomes separately from model-generated summaries. Do not promise exactly-once external effects; use the receiving system’s supported reconciliation or idempotency mechanisms where available, and make uncertain outcomes visible for review.

A deletion mismatch occurs when one representation is removed while a derived or cached copy remains retrievable. Inventory where each category can appear: primary record, indexes, caches, logs, exports and any generated summaries. Define how a deletion or correction event propagates and how completion is verified. This is an architecture and operations requirement; it should not be reduced to a button in the user interface.

A memory poisoning or misattribution failure can begin with content that was never intended to become durable memory. Separate the source event from the selected memory item, record who or what initiated its creation, and define correction paths. For preferences that materially change behavior, consider asking for confirmation at creation or when the preference is applied in a consequential context. A model’s confident restatement is not evidence that a user endorsed the stored interpretation.

Operational ownership should follow these distinctions. The team responsible for the application’s task model should define valid state transitions; the team responsible for context retrieval should implement scope and eligibility checks; product owners should define which information is useful enough to retain. These responsibilities can sit with a small team, but they should not collapse into an undocumented prompt that nobody can reliably audit or change.

9. A decision worksheet for a first implementation

Use this worksheet in a design review for each proposed record type. Fill in concrete answers before selecting a storage pattern. If two categories have different answers, split them into separate record types even if they share infrastructure.

Design questionConversation historyOperational stateAgent memory
Primary purposeInterpret the active exchangeResume or control a taskSupply selected context across interactions
Authoritative correction pathUser or application policyBusiness system or task ownerUser, authorized owner or defined source
Typical retrieval triggerResolve references in the current sessionResume, check or advance a taskApply a relevant, eligible preference or fact
Scope dimensionsSession, actor and tenant as neededTenant, task and permitted actorsTenant plus person, team or approved audience
Lifecycle end eventSession or history-policy eventTerminal outcome plus defined lifecycleExpiration, correction, removal or policy event
Failure to preventIrrelevant or excessive transcript contextIncorrect stage or repeated workStale, misattributed or cross-scope influence

Before implementation, write one sample record for each category using fictional values. Include the tenant and subject scope, source reference, creation time, eligibility status and a realistic expiry or review event. Then ask an engineer who did not design the schema to explain which record a request can retrieve and why. If the answer depends on interpreting a free-form note, the access or lifecycle rule is not yet explicit enough.

Next, walk through three requests: a normal request within one tenant, a request with missing identity context, and a request where the relevant record is expired or belongs to another subject. For each, define the expected retrieval result and the user-visible response. “No record returned” should have a deliberate behavior, such as asking for clarification, using a neutral default or declining to continue—not an automatic fallback to a wider search.

A compact printable review can use these checkboxes:

  • Purpose: Every stored field belongs to history, task state or memory, with an identified reason.
  • Scope: Tenant and subject boundaries are resolved from trusted context and enforced before retrieval.
  • Authority: The design identifies who or what can correct each value and distinguishes claims from verified state.
  • Lifecycle: Creation, refresh, expiry, correction and removal events have defined effects on eligibility.
  • Fallback: Missing, stale, conflicting and unauthorized records produce a predictable response.
  • Traceability: Operators can identify which record category and scope influenced a response without retaining unnecessary content.
  • Test cases: The design includes cross-tenant, wrong-subject, expired-record and conflicting-state cases.

This artifact is useful before choosing a specific managed service or database. It exposes policy and interface decisions that remain necessary across storage implementations. If the design spans multiple AWS accounts, teams or deployment environments, cloud platform engineering can help establish reusable boundaries and operational patterns without making those product-policy decisions on behalf of the application team.

10. Evolve retrieval without blurring the categories

A first version can be intentionally narrow. Start with a clearly defined conversation window, a task record with explicit transitions and a small set of memory types whose source and audience are understandable. Add a memory type only when a real task needs context beyond the current exchange and the product can explain its lifecycle. Avoid collecting broad conversational material on the assumption that future retrieval will make it useful.

As usage grows, measure operational quality in terms that map to the design: whether correct task state was loaded, whether eligible memory was available, whether out-of-scope records were excluded, and whether operators could resolve stale or conflicting context. These are proposed evaluation dimensions, not benchmark claims. A high rate of retrieval is not success if the returned item is wrong, while an empty result can be correct when no eligible record exists.

Retrieval may also interact with concurrency, latency and downstream actions. Those are important workload concerns, but they do not justify weakening tenant or lifecycle checks. The separate discussion of constraints in scaling AWS agent workloads can inform capacity planning; memory scope remains a correctness condition at every scale.

Keep tool execution distinct from memory retrieval. A retrieved preference may shape a response, but it should not itself authorize an operation. When a task proposes an external change, the application should validate the precise action and record the result from the system responsible for that action. The design of governed tool access is a separate topic covered in turning existing APIs into governed agent tools.

The practical principle is straightforward: store each kind of information for a stated purpose, retrieve it only within a verified scope, and let the system responsible for a fact remain its authority. Conversation history explains what was said; operational state explains what the application believes is happening; memory supplies selected context that may help later. When those roles stay distinct, teams can change retention, retrieval and task behavior independently—and investigate failures without treating every problem as the agent’s memory.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call