Purpose and experiment boundaries
An agent’s context is the information available while it decides what to do next: the current request, conversation history, retrieved records, tool results and any compacted notes. Context engineering is the deliberate choice of what to retain, condense or fetch. The practical question is not whether one policy is best in the abstract. It is which policy preserves the information a particular task needs, at an acceptable cost and with evidence you can inspect.
This walkthrough builds a small experiment around three policies. Keep passes a fixed, complete task packet to the agent. Summarise replaces earlier material with a compact note, then adds the current task. Retrieve stores source records separately and selects a small set for each task. All three receive the same initial evidence and must complete the same tasks. The experiment changes context policy—not task wording, available tools, model settings or scoring rules.
Treat the result as a decision aid, not a model leaderboard. A short benchmark can expose a dangerous omission or retrieval miss; it cannot establish that a policy will behave identically across every project, prompt, model or workload. Run the design against representative cases, preserve the inputs and outputs, and report the limits alongside any aggregate score.
The experiment also needs a clear boundary between measured and illustrative information. The example records and worked calculations below are hypothetical. They demonstrate how to structure and interpret a run; they are not results from executed model tests. When you run the harness with your own agent, label actual observations as measured and retain the run artifacts that support them.
What you need before starting
You need a working agent or an adapter that can accept a task plus a context packet and return a response. The agent may call tools, but for the first comparison either disable them or expose precisely the same tools and tool results to all three policies. If one arm can fetch information from a live system while another cannot, the experiment no longer isolates context policy.
Prepare a small, versioned task set. Include the source records required to answer each task, at least one distractor, and an answer key or scoring rubric. Use synthetic records for the first pass so that the experiment does not expose customer or employee information. Keep each record stable across runs; changing source text midway makes results hard to compare.
Choose a fixed agent configuration and record its name, version or deployment identifier, system instructions, sampling settings if available, tool definitions, and any output schema. Do not infer a specific model capability from this exercise. If the agent configuration changes, start a new experiment version rather than blending outputs from unlike runs.
Install Python 3.10 or later if you want to use the standard-library harness below. The code does not call a model or provider. It prepares policy contexts and records outputs supplied by an adapter. That separation lets you connect your existing agent without embedding invented API names, credentials, or provider-specific behavior in the experiment.
Finally, set a maximum context budget for each policy in the same unit. Token counts are useful if your agent exposes them consistently; character counts are a rough fallback for an offline prototype, not a substitute for provider token accounting. Record the method. Do not claim that equal character counts mean equal token counts.
flowchart TD
accTitle: Compare three context policies on the same cases
accDescr: Replay the same held-out cases through keep, summarize and retrieve policies, with model and tools fixed. Evaluate outputs under one outcome rubric.
A["Fixed held-out cases"] --> B["Keep full permitted context"]
A --> C["Summarize permitted context"]
A --> D["Retrieve permitted context"]
B --> E["Same model and tools"]
C --> E
D --> E
E --> F["Blind outcome grading"]
F --> G["Compare errors and total cost"]
Each arm receives the same task cases and authorization boundary. The context policy changes; the model, tools and grading rule stay fixed. This is the experiment design, not a report of measured results.
Experiment design: hold the task constant
Write tasks so that success depends on information dispersed across records. For example: identify the current delivery date for a project, name the one prerequisite that remains open, and cite the source supporting each answer. A task that can be answered from one obvious sentence will not reveal much about summarisation or retrieval. A task requiring secret or unavailable information tests something else entirely.
Set the expected answer before generating contexts. For each task, define required facts, acceptable variants, required source identifiers and disallowed inferences. Keep the rubric separate from the agent prompt. Otherwise, you risk changing the question to fit an output that looks convincing.
Use three policies with explicit rules. In Keep, include the same fixed source packet each time, even if parts are irrelevant. In Summarise, create a short note from the packet under a fixed length limit, retain the note’s source references, and give the agent that note rather than the underlying records. In Retrieve, index the same records and apply a fixed query and selection rule to choose the context for each task. Do not manually improve a weak retrieval result after seeing the answer.
A fair comparison gives each policy the same task instructions, output format, source corpus and maximum context allowance. It does not require identical context contents: those contents are the treatment being tested. Keep the agent’s available tools fixed. If you compare a no-tool agent against an agent that can inspect raw sources, the result measures both context strategy and tool access.
Define an evaluation sheet before the run. Useful measures include task success against the rubric, required-fact coverage, unsupported claims, source-attribution accuracy, context size, elapsed time, and any available input and output usage counts. Record failures individually, not just as a mean. A high average can hide one serious class of omission.
Step 1: create a small source corpus
Start with a deliberately compact example. The hypothetical project packet below contains five records. Each record has an identifier, a timestamp, a topic and a statement. In a real experiment, include stable source identifiers and the original text or a durable reference to it. The identifiers must survive summarisation and retrieval so that an answer can be traced back to its origin.
| ID | Topic | Hypothetical source statement |
|---|---|---|
| R01 | Schedule | The Northstar pilot is planned to start on 14 October. |
| R02 | Schedule update | On 22 September, the pilot start moved to 21 October, pending security review. |
| R03 | Readiness | The security review is complete as of 25 September. |
| R04 | Readiness | The data mapping review remains open; its owner expects an update by 3 October. |
| R05 | Scope | The pilot covers two internal teams; external customers are out of scope. |
The task is: “State the current planned start date, identify the remaining open readiness item and its expected update date, and state whether external customers are included. Give a source ID for each claim. If a fact is absent, say so rather than infer it.” The answer key is 21 October, data mapping review, 3 October, and external customers are out of scope. The evidence should point respectively to R02, R04 and R05. R01 is intentionally stale; R03 is a completed item; both test whether the policy preserves status and recency rather than merely matching keywords.
Add tasks that vary the failure pressure without changing the underlying corpus. One can ask for a single current fact, another can ask for two facts with different source records, and a third can ask for a fact that is genuinely absent. The absent-fact task checks whether an agent invents an answer when no context policy can supply one. Keep task variants and answer keys in the run manifest.
Avoid making every question depend on the newest timestamp. If all tasks reward recency, the benchmark may measure date handling more than context selection. Add a case in which an older source remains authoritative because it describes a scope boundary that was never changed. Explain that distinction in the rubric.
Step 2: prepare the three contexts
For Keep, assemble the complete packet in a deterministic order. The order should not change across tasks or runs. Include all five records, including the old date and completed review, because filtering them out would turn “keep” into a separate selection policy. If the packet exceeds your chosen budget, record that outcome; do not silently truncate it.
For Summarise, create one compact note from the same packet. Specify a maximum length and a required format. For example, ask for current status, unresolved items, scope and source IDs. A summary that omits citations has already failed a provenance requirement, even if its factual prose happens to be correct. Keep the summary text as an artifact, not just an intermediate hidden in a log.
The summary policy needs a defined creation procedure. If a model creates the note, use the same summariser configuration and instructions throughout the experiment, and record its response. If a human writes it, record that as a human-prepared summary and do not compare it as if it were an autonomous agent policy. The provenance of a summary matters because an omission can occur before the task agent ever sees the material.
For Retrieve, store each source record separately. For a tiny corpus, a transparent lexical selector is sufficient to test the mechanics: score records by terms shared with the task, then select the highest-ranked records up to a fixed budget. That is not a claim that lexical retrieval is suitable for production. It provides a reproducible baseline whose missed records can be inspected. If your real system uses a different retrieval method, test that method and capture its query, ranking and selected IDs.
Do not change the retrieval query after reading a bad answer. If you need to tune it, mark the changed configuration as a new experiment version and rerun all policies against the same task set. Otherwise, one arm benefits from hindsight while the others do not.
Step 3: build a context-preparation harness
The following Python example prepares the three contexts, stores their source references and records one output per policy and task. It uses only the standard library. The agent_call function is deliberately a placeholder: connect it to your own agent interface, or populate the output field from a separate run. It does not execute a model, calculate provider token usage or verify factual correctness.
from dataclasses import dataclass, asdict
from datetime import datetime, timezone
import json
import re
@dataclass(frozen=True)
class Record:
id: str
topic: str
text: str
RECORDS = [
Record("R01", "schedule", "The Northstar pilot is planned to start on 14 October."),
Record("R02", "schedule", "On 22 September, the pilot start moved to 21 October, pending security review."),
Record("R03", "readiness", "The security review is complete as of 25 September."),
Record("R04", "readiness", "The data mapping review remains open; its owner expects an update by 3 October."),
Record("R05", "scope", "The pilot covers two internal teams; external customers are out of scope."),
]
TASKS = {
"T01": "State the current planned start date, open readiness item and expected update date, and customer scope. Cite source IDs; do not infer absent facts."
}
STOP = {"the", "and", "for", "with", "from", "that", "this", "are", "its"}
def terms(text):
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in STOP}
def keep(records):
return list(records)
def summarise(records):
# Fixed, human-authored note for a reproducible illustrative run.
return [{"id": "SUMMARY-01", "text":
"Current start: 21 October [R02]. Security review complete [R03]. "
"Data mapping review open; update expected 3 October [R04]. "
"Two internal teams only; external customers out of scope [R05]."}]
def retrieve(task, records, limit=3):
q = terms(task)
ranked = sorted(
records,
key=lambda r: (len(q & terms(r.topic + " " + r.text)), r.id),
reverse=True,
)
return ranked[:limit]
def serialize(items):
return "\n".join(
f"[{item['id']}] {item['text']}" if isinstance(item, dict)
else f"[{item.id}] ({item.topic}) {item.text}"
for item in items
)
def agent_call(task, context):
raise NotImplementedError("Connect your agent adapter or enter a captured response")
def make_run(policy, task_id, task, items, output):
context = serialize(items)
return {
"policy": policy,
"task_id": task_id,
"task": task,
"selected_ids": [x["id"] if isinstance(x, dict) else x.id for x in items],
"context": context,
"context_chars": len(context),
"output": output,
"captured_at": datetime.now(timezone.utc).isoformat(),
}
for task_id, task in TASKS.items():
inputs = {
"keep": keep(RECORDS),
"summarise": summarise(RECORDS),
"retrieve": retrieve(task, RECORDS),
}
for policy, items in inputs.items():
context = serialize(items)
try:
output = agent_call(task, context)
except NotImplementedError:
output = "UNRUN: connect adapter or add captured output"
print(json.dumps(make_run(policy, task_id, task, items, output)))
The fixed summary in this example is human-authored to make the walkthrough deterministic. It is not a result from a summarisation model. Replace it with the procedure you actually intend to evaluate and retain that procedure’s output. The simple retrieval selector ranks by overlapping terms and breaks ties by ID; it may select a stale or irrelevant record. That weakness is useful in a prototype because it is visible, but it should not be mistaken for a tested production retriever.
Step 4: run identical tasks and preserve provenance
Connect agent_call to the agent you intend to evaluate. Pass the task and context through the same prompt template for all policies. Keep role instructions and output schema fixed. Instruct the agent to cite only supplied source IDs and to mark unsupported details as unknown. If your agent uses tools, capture the tool inputs and outputs and ensure each policy has the same tool access and stopping conditions.
Run each policy against each task. Randomize the order of policies or use a balanced order across repeated runs if generation variation could affect the outcome. If you repeat a task, preserve every attempt; do not replace an inconvenient answer with a later response. Record the run identifier, experiment version, task ID, policy, source corpus version, selected source IDs, exact context text, exact task text, output, timestamp and available usage measurements.
A context record without provenance is difficult to debug. For a summary, preserve both the summary and the original source IDs it claims to condense. For retrieval, preserve the query, ranking, selected records and any filtering rule. For keep, preserve the complete packet and order. These artifacts let a reviewer distinguish an agent reasoning error from an upstream omission.
If your existing agent already has durable, file-backed state or a multi-session design, keep that behavior constant while testing context selection; the related implementation questions are covered in Memory Patterns for Multi-Session Agents: File-Backed Context. Similarly, if the orchestration layer itself is under consideration rather than held fixed, separate that decision from this experiment; Claude Agent SDK: When It Beats Your Own Orchestration addresses that distinct choice.
Step 5: score outcomes, not appearances
Use a rubric with one row per required fact. For T01, score whether the output gives 21 October, identifies the data mapping review as open, gives 3 October as the expected update, excludes external customers, and cites R02, R04 and R05 correctly. Mark a fact wrong if the answer cites a source that does not support it. A fluent answer that states the right date but cites R01 should not receive full provenance credit.
Score unsupported claims separately. An unsupported claim includes a detail that appears in neither the task nor the cited sources, or a conclusion that exceeds what those sources say. The phrase “the review will be approved by 3 October,” for instance, would overstate R04, which gives an expected update date rather than a promised approval. Record the exact claim and the source comparison that led to the score.
Use two passes where practical. In the first pass, a reviewer scores against the answer key without seeing the policy label. In the second, a reviewer resolves ambiguous cases and logs the reason. This does not eliminate evaluator judgment, but it makes disagreement visible. If one person performs both passes, record that limitation rather than suggesting independent review.
Keep measures distinct. Task success can be a pass/fail decision based on required fields. Fact coverage can be the fraction of required facts stated correctly. Citation accuracy can be the fraction of cited claims supported by their cited source. Context size should use a consistent measurement method. Latency and usage should come from the same instrumentation for all arms. Do not combine these into one score until you have a reasoned weighting scheme; separate measures reveal tradeoffs more honestly.
For a small task set, show counts and the underlying failures. “Four of five tasks passed” is more interpretable than a percentage with false precision. Report the number of tasks, number of repeats and any missing observations. A policy that looks strong on three easy tasks has not demonstrated reliable performance across a broad workload.
Step 6: compare and interpret the outputs
The following table shows a hypothetical illustration, not measured results. It demonstrates the kind of differences to look for. Assume a retrieval selector returns R01, R04 and R05, missing R02 because its query did not distinguish “current start” from the older date. Assume the summary contains the note shown in the code. The full keep packet contains all five records.
| Policy | Context contents in this illustration | Likely advantage to inspect | Failure to inspect |
|---|---|---|---|
| Keep | R01–R05 | Both old and current dates remain available for comparison | More irrelevant material competes for attention; stale facts can be selected |
| Summarise | Summary-01 with references | Compact statement carries the current date and open item | Omitted nuance or incorrect summary becomes the only task evidence |
| Retrieve | R01, R04, R05 | Small packet focuses on scope and readiness | Missing R02 makes the current date unanswerable from retrieved sources |
In this illustration, retrieval should not be rewarded for producing a confident date from R01. A correct behavior is to report that the current date is not supported by the selected context, or to ask for the missing source if the system permits another retrieval attempt. If it claims 14 October as current, record a factual error and a retrieval miss. If it claims 21 October without R02, record an unsupported answer even though the date happens to match the answer key.
A summary can appear superior because it already states the answer. That result is meaningful only if the summary was created from the same corpus using a declared procedure and its construction cost is counted where relevant. If the note was manually written with the answer key in view, it is a prepared-context baseline, not evidence that an autonomous summarisation step will reliably preserve the right facts.
Keep may appear wasteful on a small packet while being safer when relationships, exceptions or source history matter. But more context is not automatically better: an agent can confuse an old statement with a current one. The appropriate observation is not “keep always wins” or “shorter is better.” Inspect which facts each policy exposed, which ones the agent used, and whether the evidence cited actually supports the claim.
Step 7: diagnose failure paths
A failure can originate in at least three places: context preparation, agent response, or evaluation. A retrieval miss means the required source was not selected. A summary omission means the required fact or its qualification was lost before the task response. A response error occurs when the needed evidence was present but the agent misread or ignored it. An evaluation error occurs when the rubric or source interpretation is itself wrong. Label the earliest observable failure and keep later contributing failures as notes.
The most important operational distinction is between not present and not found. If a source exists but retrieval omitted it, the agent should not be treated as having seen the fact. If the corpus genuinely lacks a source, the expected behavior is to state that the answer is unknown. A run manifest should record the corpus version and available record IDs so reviewers can tell those cases apart.
Watch for false provenance. An agent may attach a valid ID to a claim that the record does not support, or cite a summary ID without linking it to underlying records. Check claim-to-source alignment, not merely whether an ID appears in the output. For decisions with material consequences, make source inspection part of review rather than treating a citation-shaped string as proof.
Watch for silent truncation as well. If a context exceeds a limit, record the actual delivered text and the truncation method. A run that logs the intended packet but not the packet the agent received cannot explain an omission. Fail the run as invalid if the delivered context cannot be reconstructed; do not score it as a normal policy outcome.
When a call times out or an adapter fails, record an execution failure rather than substituting a blank answer. Set a retry rule before running, including the maximum attempts and whether retries count as separate observations. Preserve the failed attempt and the retry relationship. This prevents operational instability from disappearing inside a cleaned-up results table.
Step 8: summarize evidence without overstating it
Create a result sheet that reports the experiment version, date, agent configuration, task set, context policy definitions, measurement method and number of runs. Present outcomes by task and policy. Include a brief failure narrative for each material error: which source was required, what context was delivered, what the agent returned, and how the rubric scored it.
An illustrative calculation shows why per-task detail matters. Suppose, purely for demonstration, that a team runs six tasks once per policy, with five required facts per task. If Keep gets 27 of 30 facts correct, Summarise gets 28 of 30, and Retrieve gets 24 of 30, those totals do not show whether one policy failed an important “unknown” case or repeatedly made a low-impact omission. These figures are assumed for the calculation and are not observed results. Report the individual task outcomes before considering the aggregate difference.
If you calculate context reduction, define the baseline and units. For example, compare the character count of the delivered summary context against the character count of the complete packet for the same task. Label that as a character-count reduction, not a token or cost saving. If your agent supplies consistent usage measurements, report those separately and retain the raw values. Do not convert them into financial savings without verified rates and a clearly stated calculation basis.
For repeated runs, report spread as well as the mean or pass rate. A policy that succeeds inconsistently may be less suitable than one with a slightly lower average but fewer severe failures. The right treatment depends on the task: an internal draft and a customer-facing action do not necessarily deserve the same failure tolerance. Do not hide a critical unsupported claim inside a blended average.
The result should lead to a bounded decision, such as “use retrieval for tasks whose required records are consistently selected, with a fallback when evidence is missing,” rather than a universal declaration. State the contexts tested, the limitations and the next experiment that would challenge the conclusion. Keep the raw run artifacts available to the people responsible for adopting the policy.
Troubleshooting common problems
The three arms receive different instructions. Compare the rendered prompts, not just the template files. Confirm that the task, output format, tool availability and stopping rules match. Treat prompt changes as a new experiment version.
The summary wins every task. Check who or what generated it, whether its creator saw the answer key, whether it includes source IDs, and whether its preparation effort was recorded. Add a task requiring a preserved qualification or an absent fact. A prepared note can be a useful benchmark, but it is not automatically a fair measure of automated summarisation.
Retrieval selects the wrong records. Save the query, ranking and selected IDs. Determine whether the query wording, ranking method, tie-breaking or context budget caused the miss. Change one component at a time and rerun every policy under the revised experiment version. Never quietly hand-add the missing source to a single failed run.
The agent answers correctly but cites the wrong source. Score factual accuracy and provenance separately. Inspect whether the cited record supports the specific claim, whether the context included the correct record, and whether the output instructions distinguish current from superseded facts. Correctness without traceable support is a different result from evidence-grounded correctness.
Context counts do not match expectations. Inspect the exact serialized context and the measurement unit. Whitespace, metadata and wrappers can change length. If truncation is involved, capture the final delivered content. Treat character counts as a local comparison measure only unless you have a consistent token-counting mechanism.
A run cannot be reproduced. Check that source text, task wording, policy code, configuration identifiers, selected records and outputs were saved. A timestamp alone is not enough. If any material input is missing, mark the run non-reproducible and avoid using it as decisive evidence.
Use this decision worksheet
Copy this worksheet into the experiment record before running. Complete it once per experiment version, then attach the per-task rows and raw outputs. The worksheet is intentionally compact; its purpose is to make assumptions and provenance inspectable, not to replace the source artifacts.
- Decision under test: What concrete context-policy choice could this run inform?
- Task set: Which task IDs, answer keys and absent-fact cases are included?
- Source corpus: What version and record IDs were available to every policy?
- Policy definitions: What exact rules produce Keep, Summarise and Retrieve contexts?
- Fixed configuration: Which agent instructions, settings, tools and output rules remain unchanged?
- Budget and measurement: What context limit applies, and how are size, latency and usage measured?
- Provenance: Can every delivered context, selected source, response and score be reconstructed?
- Scoring: How are required facts, unsupported claims and citations judged?
- Failure handling: How are timeouts, retries, truncation and missing evidence recorded?
- Interpretation: What result would change the decision, and what limitations will be reported?
For each task-policy pair, record: run ID; task ID; policy version; delivered source IDs; delivered context; output; required facts passed; unsupported claims; citation accuracy; size and latency measurements; execution status; reviewer note. Keep the original output even when it is malformed or incomplete. Do not replace it with a corrected answer.
Choose the next experiment
A useful first run is small enough to inspect by hand and varied enough to expose distinct errors. Begin with a few tasks that test recency, cross-record synthesis, source attribution and absence. If one policy misses a source, make the next experiment target that mechanism: adjust the retrieval query or selection rule, revise the summary constraints, or test whether a complete packet creates confusion. Change one factor at a time where possible so the next result is interpretable.
Once the mechanics are sound, expand the corpus and tasks using representative, appropriately handled data. Retain a held-back task set if you plan to tune a policy; otherwise, repeated adjustment against the same questions can make the experiment look stronger than it is. Revisit the rubric when source meaning or acceptable answers are genuinely ambiguous, and document the change rather than rescoring silently.
If the experiment informs an agent used in an operational workflow, define what happens when required evidence is missing: stop, ask for clarification, retrieve again under a bounded rule, or route the task for human review. The chosen behavior should be explicit and tested. Context selection alone does not verify that a business outcome occurred, and an agent’s statement that it completed a task is not evidence that an external system changed as intended.
The broader context-engineering principle is that context is limited, so selection, retrieval and compaction are design choices—not neutral plumbing (Anthropic’s engineering discussion). This experiment makes those choices inspectable by keeping the task constant, preserving source provenance and recording the actual delivered context.
For work that needs a more complete implementation plan, AI agent engineering is the relevant next step. Keep the engagement focused on a defined workflow, reproducible evaluation and observable failure handling rather than treating a single comparison as proof of general reliability.
The useful conclusion from a keep, summarise or retrieve experiment is a policy tied to a task class and a known failure response. Keep the task set, sources and scoring rules stable enough to reproduce the comparison; preserve the evidence behind each score; and make the next revision answer a specific failure you observed. That produces a practical engineering decision without turning a small benchmark into a universal claim.