Purpose and result
An agent trace can show that a model was called, a tool returned data, or a workflow took several steps. Those events are useful for reconstructing execution, but they do not establish that the requested business task succeeded. A response may sound complete while the downstream record remains unchanged; a service may accept a request that applies to the wrong account; or an agent may stop after encountering an error without making that failure clear to its caller.
This tutorial builds an AWS-oriented reference design that connects execution telemetry to a separately verified business outcome. Its deliverable is a compact task record joining a task identifier, trace references, state transitions, and a result-verification status. The design treats identity, durable task state, telemetry collection, and business-system effects as distinct boundaries. You can adapt those boundaries to your own AWS account layout and runtime rather than assuming a particular deployment or vendor integration.
The distinction matters operationally. Tracing helps answer what the system attempted and where execution diverged. Outcome verification answers whether the intended result is visible in the authoritative business system. A useful service view needs both, without treating model-generated text as proof of an external effect.
The AWS observability guidance describes traces and metrics. In this proposed architecture, business-outcome verification belongs in an application-owned record. This design follows that separation: collect execution signals, then record and verify task results in an application-owned outcome ledger. AWS AgentCore observability.
Prerequisites and setup requirements
Before implementation, choose one bounded workflow with a clear completion condition. A suitable first candidate has a stable task identifier, an authoritative system where its result can be checked, and a limited set of actions whose consequences are understood. Avoid beginning with a broad assistant that can take arbitrary actions across multiple business systems. Its many possible paths make it harder to decide what counts as complete.
You will need an AWS environment where your team can deploy the application components it selects; an agent runtime or orchestration layer; an identity mechanism for the workload; a durable store for task state; a telemetry destination; and a way to query or otherwise verify the business system that owns the desired result. The exact services and deployment choices depend on your existing architecture. The design below defines responsibilities and data contracts rather than prescribing undocumented service APIs.
Assign clear identifiers before instrumenting calls. At minimum, create a task_id for the business request, a run_id for one execution attempt, and an event_id for each recorded state or outcome event. If the runtime exposes a trace identifier, store its value as a correlation reference, not as a substitute for the task identifier. One task can involve more than one attempt or trace, and a trace can contain many model and tool events.
Agree on a completion predicate with the business owner. For example, “the case exists with the requested account, category, and status” is testable; “the agent says the case was created” is not. Identify the authoritative source and the fields that must match. Also decide what the application should do when verification is delayed, unavailable, or inconclusive.
Before production use, review what the telemetry may contain. Tool inputs and outputs can include customer data, credentials accidentally included in payloads, or confidential business details. Define field-level capture rules, redaction or omission behavior, access boundaries, and retention expectations in line with your organization’s policies. These are design decisions to make for your environment, not guarantees supplied by a tracing tool.
If your implementation uses OpenTelemetry generative-AI semantic conventions, pin the instrumentation and convention version used by the application, and review what content is captured. The conventions can change, so teams should not treat a moving convention set as a stable data contract. OpenTelemetry GenAI semantic conventions.
Reference design and boundaries
The design uses four logical areas. The task entry boundary authenticates the caller and creates a task record. The agent execution boundary runs a single attempt using an explicitly identified workload identity. The telemetry boundary receives execution events and associates them with task and run identifiers. The business-system boundary performs or observes the consequential action and supplies evidence for a separate verification step.
These areas may be implemented with different AWS services, or some may share an execution environment. The important constraint is that their responsibilities remain distinguishable in logs, permissions, and failure handling. For example, the ability to emit telemetry should not silently imply permission to modify a customer record, and an agent’s ability to propose an action should not itself count as confirmation that the action succeeded.
A task state record should be small enough to inspect and stable enough to query. One proposed shape is: task_id, run_id, status, created_at, updated_at, trace_ref, action_ref, verification_status, verification_checked_at, and a short failure_class. Add a business key only when it is needed for controlled verification, and protect it according to its sensitivity. Do not copy full prompts or full tool payloads into this record merely to make tracing easier.
The task state and telemetry serve different purposes. State answers “what is the current business workflow status?” Telemetry answers “what happened during execution?” The outcome record answers “what evidence supports the completion claim?” Combining those concepts into one mutable status field makes dashboards ambiguous. For instance, agent_finished is not synonymous with business_effect_verified.
A proposed sequence is: accept a task, persist its initial state, execute the agent, record each material transition, submit an allowed action, then verify the resulting business state. If the action response is missing or verification cannot establish the result, record an unresolved status and preserve enough references for investigation. Do not infer failure solely from a timeout: the external system may have accepted a request even if the caller did not receive its response.
flowchart TD
accDescr: Workflow stages and decisions: Accept task and persist ID, Run agent and emit correlated events, Submit permitted business action, Verify authoritative business state, Record verified outcome, Record unresolved and investigate. The adjacent text explains the conditions and exceptions.
accTitle: Tracing AWS Agents Across Models, Tools and Business Systems workflow
A["Accept task and persist ID"] --> B["Run agent and emit correlated events"]
B --> C["Submit permitted business action"]
C --> D["Verify authoritative business state"]
D -->|"Matches task predicate"| E["Record verified outcome"]
D -->|"Missing or uncertain"| F["Record unresolved and investigate"]
The diagram shows a deliberate decision after the action, not an assumption that a successful tool response equals success. The verification branch should evaluate the task’s agreed predicate against the authoritative system. A match permits the application to record a verified outcome. A missing or uncertain result remains unresolved, even if the model presents a confident summary. Keep the action boundary and verification boundary separate so that a tool response cannot mark its own effect as verified.
For background on tool-call trace structure, see MCP Observability: Tracing Tool Calls Across an Agent Loop. This tutorial focuses on joining that execution view to business task completion, not on reproducing a tool-tracing implementation. For state that persists across turns or attempts, AI Agents in Production: Memory Architecture treats the distinct memory design problem.
Step 1: Define the task and its completion predicate
Write down one representative request and its completion condition before adding instrumentation. Suppose an internal operations workflow is asked to create a case for account A-104 with category delivery_issue and status open. The task is complete only if the authoritative case system contains a case linked to that account with the requested category and status. The agent’s generated sentence “Case created” is a claim to compare with that condition, not the condition itself.
This is an illustrative scenario, not a statement about any PADISO deployment or a tested integration. Its value is that it gives each layer a precise job. The task record carries the request identifier and expected attributes; execution events record how the agent reached an action; the external system remains authoritative for whether the case exists; and a verifier compares the observed record with the expected attributes.
Create an explicit state vocabulary. A starting set might include accepted, running, action_submitted, verification_pending, verified, unresolved, and failed_before_action. Use states that describe observable application facts. Avoid labels such as successful unless the team has defined precisely which evidence permits that transition.
Define transitions and their evidence. For example, action_submitted means the application recorded a request to perform an action, not that the external system applied it. verified means a verification procedure observed the required fields. unresolved means there is insufficient evidence to establish either success or failure. This vocabulary makes dashboards and incident notes more reliable than a single Boolean success flag.
Record the expected result in a form the verifier can evaluate, but do not turn unrestricted model output into executable verification logic. For the illustrative case, the predicate is a small set of field comparisons. In a real system, the application should own and validate the predicate, and the business owner should confirm it describes the outcome that matters.
Step 2: Establish identity and service boundaries
Give each workload component an identity appropriate to its role. The task-entry component needs only the authority required to accept and persist requests. The agent execution component needs the permissions required for its defined model and tool interactions. The verifier needs read access sufficient to inspect the outcome fields. A telemetry writer needs authority to emit the allowed event data. These roles are a proposed separation of responsibilities; map them to the identity mechanisms and services available in your environment.
Do not pass broad business-system credentials through model context, prompts, or traces. Keep credentials outside the data used to explain a task, and ensure that a tool call receives only the authority required for its specific operation. Where an action changes business state, make the action service boundary explicit so that the application can validate the target and payload before submission.
Use a trace reference as correlation data, not as authorization. Knowing a trace identifier should not let a caller read a task record or inspect a customer record. Similarly, a task identifier is not a secret or a bearer credential. Apply access checks at the data and action boundaries independently of whether a request is associated with a valid trace.
Record which identity performed each consequential operation in the application’s event model. A useful event includes the actor or component identity, task and run identifiers, action type, timestamp, and an opaque reference to the request or result. Keep sensitive content out unless the operational case justifies it and the handling rules permit it. The purpose is to reconstruct responsibility and sequence without turning telemetry into a second, less-controlled copy of the business system.
For a broader discussion of AWS foundation choices that shape identity and service boundaries, consult An AWS Agent Landing Zone: The First Architecture Decisions. The boundary decisions here are limited to what this task’s telemetry and outcome verification require.
Step 3: Persist task state before execution
Create the task record before starting the agent run. This gives retries and later investigation a stable business identifier even if execution ends unexpectedly. The initial record should include the task ID, current state, creation time, and a link to a controlled representation of the requested operation. Persist the record before emitting “running” or attempting a consequential action so that an interrupted execution does not disappear from the operational view.
For each state transition, append an event or otherwise retain a history that lets an operator distinguish current state from earlier state. A minimal event might contain event_id, task_id, run_id, event_type, occurred_at, component, and a bounded set of non-sensitive attributes. Update the current-state projection separately if that helps queries, but preserve enough history to reconstruct ordering and repeated attempts.
Decide how retries relate to tasks. A retry of the same business request should normally retain the task identifier while receiving a distinct run identifier. This makes it possible to ask whether one task had several attempts and which attempt produced the eventual verified result. If the system treats a repeated request as a new task, document that behavior and how duplicate external effects are prevented or detected.
Do not promise exactly-once effects at an external boundary. A request may be applied while its response is lost, and a retry may encounter the already-created result. Design the action and reconciliation path with that ambiguity in mind. If the business system supports an idempotency mechanism, validate its semantics in the chosen integration before relying on it; do not assume one exists or that it covers every failure mode.
A compact state update can be expressed in pseudocode without tying it to an undocumented AWS API:
function record_transition(task_id, run_id, next_state, details):
assert allowed_transition(current_state(task_id), next_state)
append_event({
event_id: new_id(),
task_id: task_id,
run_id: run_id,
event_type: next_state,
occurred_at: now(),
details: allowlisted(details)
})
update_task_projection(task_id, next_state, run_id)
This is illustrative application logic, not executable code for a particular persistence service. allowed_transition should reject illegal changes; allowlisted should exclude data that the telemetry policy does not permit. In a concurrent system, state updates also need a strategy for detecting conflicting writes, such as a version check or an equivalent conditional update supported by the chosen store.
Step 4: Emit correlated execution events
Instrument events at meaningful boundaries instead of treating every internal detail as equally useful. Record when a run starts and ends; when a model or tool operation begins and returns; when an action is submitted; when a timeout or exception occurs; and when verification concludes. Associate each event with task_id and run_id, and attach the available trace reference consistently. If an event cannot be linked to a task, make that gap measurable rather than silently placing it in the same dashboard as complete records.
Choose an event schema that supports the operational questions the team expects to ask. For a tool call, the useful shape often includes the tool category, start and end times, outcome class, and a reference to the action record. It does not automatically require the full input and output. For model activity, capture only the attributes that are useful and approved for diagnosis. Token, prompt, response, or other detailed content should not be collected by default merely because an instrumentation library can represent it.
Test instrumentation changes against saved, sanitized trace fixtures before rollout. Compare attribute presence, redaction and task correlation with the prior release; a dashboard should not silently change the meaning of success when an instrumentation package changes.
Define the correlation contract in code and dashboards. For example, every agent-run event must carry a task_id; every action event must carry both the task_id and run_id; and every verification event must refer to the action or observed business record it evaluates. Enforce those expectations at the application boundary. A field that is usually present but not checked becomes a recurring investigation gap.
Do not collapse separate attempts into one trace-shaped story. An agent may produce multiple runs for one task, and a business action may be retried or reconciled after an ambiguous response. Keep attempt-specific data on the run while retaining task-level status and result references. This allows an operator to see, for example, that the first attempt timed out, the second found an existing case, and the verifier confirmed the expected state.
Step 5: Verify the business result independently
Implement verification as an application-owned operation with a clear input and output. Its input can be the task identifier, expected predicate, and action reference. Its result should distinguish at least “matches,” “does not match,” and “unable to determine.” A query failure, timeout, or incomplete record is not the same as a confirmed mismatch, and neither should be silently converted to verified success.
For the illustrative case, verification reads the authoritative case record and compares account, category, and status with the task’s expected values. If all required fields match, the verifier records a verified outcome with a timestamp and a reference to the observed record. If no matching record is visible, it records unresolved or a carefully defined mismatch state according to the system’s consistency behavior. It should not simply ask the model whether the task succeeded.
Some systems expose changes only after a delay or return partial data. Define an explicit verification schedule suited to the system’s behavior, such as an initial check followed by a bounded retry interval. The exact interval is a workload decision, not a universal default. Record each check and its result so operators can distinguish “not yet visible” from “never checked.” Stop automated checking at a defined limit and route unresolved tasks to the operational path your team has designed.
Be careful with compensating actions. If verification detects a mismatch, automatically issuing a second mutation may create duplicates or make the original state harder to understand. First determine whether the initial action could have taken effect, whether a safe reconciliation rule exists, and whether the action is reversible. Where that cannot be established, preserve the evidence and use an appropriate human or application-controlled resolution path rather than guessing.
The task status should reflect the strongest evidence available. A successful transport response may justify action_submitted; a matching authoritative record may justify verified; a lost response followed by no conclusive query may justify unresolved. Keeping those states separate gives operators a more honest account of what the system knows.
Step 6: Build task-completion views and useful alerts
Create a task-level view that answers operational questions rather than presenting a wall of raw events. For each task, show current status, creation and last-update times, number of runs, the latest run outcome, a trace reference, action reference, verification status, and a short failure class. Provide a way for an authorized operator to move from the task to relevant execution telemetry without exposing unrelated task data.
At service level, track counts and rates that distinguish execution from outcome. Examples include tasks accepted, runs ending with errors, actions submitted, outcomes verified, unresolved outcomes, and the age of tasks still awaiting verification. If you calculate a completion rate, state its denominator and time window. “Verified tasks divided by tasks accepted during the same period” is interpretable only if the reporting rules address tasks still in progress and late-arriving results.
An alert should point to a condition that needs action. A rise in unresolved outcomes may indicate a downstream system problem, a verification query issue, or a correlation defect. A rise in run errors may indicate execution trouble but does not, by itself, tell you whether earlier actions took effect. An alert that includes sample task identifiers and the relevant time window is more useful than one that reports only a percentage.
Treat missing telemetry as a quality signal. If a task reaches a terminal application state without a corresponding verification event, surface it as an instrumentation or workflow defect. Likewise, if many telemetry events lack task_id, do not count them as ordinary successful traces. A system that records only well-behaved runs can make its apparent completion picture look better precisely when its weakest paths go unobserved.
Keep dashboards consistent about timestamps and grouping. Distinguish event time from ingestion time where delayed delivery is possible, and distinguish task-level counts from run-level counts. When one task has multiple runs, counting each run as a separate task can inflate workload and distort completion measures. Put the unit of analysis in the panel title or accompanying description.
Step 7: Exercise failure paths before rollout
Test a normal completion, a tool error before any action, an action timeout with an uncertain external result, a successful action followed by a failed verification query, and a retry that encounters an already-present result. These are scenario-based acceptance checks, not a claim that a particular implementation has been tested. For each, confirm the task state, event history, visible dashboard result, and operator next step.
A useful acceptance condition is that every accepted task can be found by its task ID; every execution attempt has a distinct run reference; an action response does not itself mark the task verified; and every verified outcome points to the evidence used by the verifier. Add a separate check that sensitive fields excluded by policy do not appear in captured events or diagnostic views.
Consider a failure timeline. At 10:02, a task is accepted and persisted. At 10:03, a run submits a case-creation request. At 10:04, the caller times out before receiving a response. At 10:05, the verification query is unavailable. The correct operational conclusion at 10:05 is not “failed” and not “verified”; it is “unresolved.” When the query becomes available, the verifier can determine whether the case exists and matches. Preserving that sequence prevents an eager retry from creating an unintended duplicate.
A counterexample helps expose an inadequate design: the application logs the model’s final sentence as task_status=success, then retries whenever the client connection drops. The sentence may be inaccurate, the request may already have applied, and the retry may repeat the effect. Such a design confuses model output, transport state, and business state. Separate records and a verification boundary make the ambiguity visible instead of disguising it.
Troubleshoot common symptoms by following the identifiers. If a task has no trace reference, check whether the run-start event was emitted and whether correlation fields were passed across the application boundary. If a trace exists but no task appears, check task persistence timing and the task-level query. If an action is recorded without verification, inspect the verifier’s availability, predicate, and permissions. If verified status appears despite a missing business record, inspect the transition rules and identify any code path that trusts an agent claim or tool response.
If dashboards show inconsistent counts, first check whether panels count tasks, runs, events, or actions. Then examine late-arriving events, duplicate delivery, and retries. Do not “fix” a discrepancy by changing the aggregation until the unit and source of each count are understood. If sensitive content appears unexpectedly, stop the capture path that exposes it, follow the organization’s incident procedures, and revise the allowlist before restoring that instrumentation.
Implementation worksheet and rollout decision
Use the following worksheet in a design review. Fill it in for one workflow before expanding the pattern. The answers should be concrete enough that an engineer can implement the data contract and an operator can interpret an unresolved task without asking the model what happened.
| Decision | Record before rollout |
|---|---|
| Task unit | What business request receives one stable task_id? |
| Attempt unit | What creates a new run_id, and how is a retry represented? |
| Completion predicate | Which authoritative fields must match for the task to be verified? |
| Action boundary | Which component may submit the consequential request, and what does its response prove? |
| Verification boundary | Which component reads the authoritative result, and what are its three outcomes? |
| Correlation contract | Which events must contain task, run, and trace references? |
| Capture policy | Which fields are allowed, omitted, or redacted in telemetry? |
| Unresolved path | What happens when a response or verification result is ambiguous? |
| Operational view | Can an authorized operator find the task, its runs, action, and evidence? |
| Acceptance checks | Which normal and failure scenarios must pass before rollout? |
For a printable summary, copy this short checklist into your implementation ticket:
- Define one bounded workflow and an observable business completion predicate.
- Assign stable task IDs and distinct run IDs; persist the task before execution.
- Separate workload identity, action authority, verification access, and telemetry writing.
- Record state transitions with bounded, policy-approved fields.
- Correlate execution events to tasks without treating trace IDs as authorization.
- Verify effects against the authoritative business system; keep uncertain results unresolved.
- Test retries, timeouts, missing telemetry, and verification outages before rollout.
- Pin instrumentation versions and review event content when conventions change.
- Show task counts separately from run and event counts in operational views.
Start with a small set of tasks and inspect whether operators can answer three questions from the records: what was requested, what actions were attempted, and what evidence supports the final status? If one answer depends on interpreting a model-generated summary, improve the application record or verification path before broadening the workflow.
The design also needs to fit the operational constraints of the runtime and downstream services. Concurrency and capacity planning are separate from task-outcome semantics; see Scaling Agent Workloads on AWS: Concurrency Is Only One Constraint when scaling behavior is the decision at hand. Browser and code-execution boundaries require their own treatment, covered in AgentCore Browser and Code Execution: Designing Safe Boundaries.
When the implementation crosses several AWS services, identity systems, and business applications, a platform review can help keep these contracts consistent. Teams planning that work can explore cloud platform engineering as a relevant next step.
Rollout sequence
Roll out the instrumentation before making the agent workflow broadly consequential. First run the task and event model against representative requests, including the failure scenarios above. Check that the verifier can distinguish matching, mismatching, and unavailable results, and that the task view stays useful when a task has several attempts. Then enable a limited operational workflow with a defined path for unresolved outcomes, while watching for missing identifiers, unexpected payload capture, and confusing state transitions.
Expand only when the evidence chain remains understandable at the task level. A reliable record is not the one with the most captured content; it is the one that lets an authorized operator connect a request to its execution and determine whether the intended business result was actually observed. Keep the task ledger, execution telemetry, and authoritative business record distinct, and make their links explicit. That gives engineering and operations a shared basis for investigating failures without asking a model’s final answer to stand in for the state of the business.