SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
Tutorial 19 mins

Tracing a Foundry Agent from User Request to Business Outcome

Tracing a Foundry Agent from User Request to Business Outcome. Practical examples, tradeoffs and implementation guidance for technology leaders.

The PADISO Team ·

What a trace can—and cannot—tell you

A trace is a connected record of work performed while handling a request. For an agent, that work may include the incoming request, model interaction, tool invocation, returned tool result, and application response. A trace helps an operator reconstruct the execution path and locate delays or failures. It does not, by itself, establish that the business task was completed correctly.

That distinction matters because an agent response can sound complete while a downstream operation has failed, or a tool can return success while the requested business change remains unapplied. To connect execution evidence to a business outcome, you need both technical spans and an independently recorded transaction result. Keep those records linked with a stable identifier, but do not treat them as interchangeable.

For Azure AI Foundry, tracing requires explicit instrumentation and correlation; tracing alone does not guarantee business success. Follow the current Foundry tracing setup guidance for the supported instrumentation path in your chosen application. Keep the application’s own transaction checks alongside that trace. The walkthrough below shows how to organize the correlation and verification around those records without assuming that a model response is proof of completion.

This is a tutorial for engineering teams who already have an agent workload or a representative test environment. If the platform choice itself is still unsettled, first compare the operational and control requirements in the enterprise agent platform selection discussion. For tracing patterns that focus specifically on MCP tool calls and agent loops, see the MCP tracing implementation discussion. This article stays narrower: connect a request to tool activity, failure evidence, and a separately verified business transaction.

Prerequisites and setup requirements

Before changing instrumentation, identify one business action that can be tested safely and whose completion can be verified independently. Good candidates have a durable record or a clear state transition, such as a work item moving from “submitted” to “assigned.” Avoid beginning with an action that has ambiguous completion semantics, multiple undocumented side effects, or no way to distinguish a real change from an agent’s confirmation message.

You need an application environment where you can instrument the agent execution path, a test identity and test data, and a way to inspect the downstream system’s resulting state. You also need an agreed retention and access approach for telemetry. Prompts, tool arguments, and tool results can contain sensitive content; decide what may be recorded, redacted, or omitted before enabling content capture. Trace identifiers should support investigation without becoming a second copy of sensitive business records.

Pin the version of any instrumentation and semantic-convention packages used by your application. Generative AI semantic conventions can change, so a pinned version makes fields and dashboards more predictable across deployments. Protect captured content according to its sensitivity and minimize it where possible. Review the current Generative AI semantic conventions before relying on a particular attribute name or interpretation; do not assume that a field observed in one version will remain unchanged in another.

Set up three distinct records before you start:

  1. Execution trace: spans for the request path and relevant tool operations, carrying a shared trace context where propagation is supported.
  2. Business transaction record: an application-owned record of the intended action, its stable transaction identifier, and its verified final state.
  3. Diagnostic event: a concise account of a failure or exceptional transition, with references to the trace and transaction rather than copied sensitive payloads.

These records answer different questions. The trace explains what the agent and application attempted. The transaction record says what the business system ultimately reflects. The diagnostic event helps an operator find and classify cases that need intervention. A dashboard that merges them into one green “success” signal hides the very boundary this design is intended to expose.

Step 1: Define the transaction boundary

Write down what the user asked for, what durable change would satisfy the request, and what evidence will prove that change occurred. Keep this definition independent of the agent’s wording. For example, “the agent said it submitted the request” is not a completion criterion. “The target system contains a request with the expected external reference and an accepted state” is a verifiable criterion, assuming that system is the authoritative record for the action.

Give each business attempt a transaction identifier generated or assigned by your application. It should identify one intended operation, not a person, conversation, or broad batch. If one user request can create several independent business actions, create a distinct transaction identifier for each action and retain a parent request identifier to group them. That structure makes partial completion visible instead of forcing the whole request into one ambiguous status.

Record a minimal transaction envelope before invoking tools. Useful fields include the transaction identifier, request identifier, action type, target-system reference if known, creation time, current status, and a redacted or classified input reference. Use a controlled status vocabulary such as requested, in_progress, pending_verification, succeeded, failed, and needs_review. Define allowed transitions in application logic; otherwise, different code paths may report contradictory states.

A request identifier is for grouping work. A transaction identifier is for a particular intended business effect. A trace identifier groups spans for an execution path. Keep all three concepts distinct even if an implementation happens to reuse a value in a narrow case. Their lifetimes and cardinalities can differ: a request may create more than one transaction, and a retry may create a new trace while continuing investigation of the same transaction.

Step 2: Instrument the execution path

Use the Foundry tracing setup appropriate to the agent framework and application, following the current product instructions rather than copying instrumentation code from a different framework or version. Instrument the application boundary where a request enters and the points where it calls tools. The objective is to make the path observable with meaningful parent-child relationships, not to capture every internal detail or maximize span volume.

For each tool operation, record a start and finish, an outcome category, and a reference to the business transaction when the operation serves one. Capture duration and a stable tool name. Add a small number of operational fields that let responders distinguish action type and environment. Do not put a raw access token, full customer record, or unrestricted prompt into a span attribute. If content must be retained for a specific diagnostic purpose, apply an explicit access and retention decision rather than making it the default.

The following pseudocode illustrates the application-level relationship. It is not a Foundry SDK sample and does not prescribe a specific tracing API. Adapt the concepts to the instrumentation supported by your application and test propagation end to end.

request_id = new_request_id()
transaction_id = create_transaction(action_type, status="requested")

with request_trace(request_id=request_id):
    try:
        result = run_agent(user_request)
        record_agent_completion_claim(transaction_id, result.status)
        verification = read_authoritative_business_state(transaction_id)
        set_transaction_status(
            transaction_id,
            status=classify(verification),
            verification_reference=verification.reference
        )
    except error:
        record_failure(
            request_id=request_id,
            transaction_id=transaction_id,
            category=classify_error(error)
        )
        raise

The important property is not the syntax. It is that the transaction exists before the tool action, the trace context covers the execution being investigated, and a verification step updates the business record based on an authoritative observation. The classify operation must be your application’s explicit decision logic; do not infer success merely because the agent returned normally.

A useful expected result after this step is one request-level trace that can be found using the request identifier, with tool spans that show their operation names and outcomes. The transaction record should separately show the same request reference and a transaction identifier. If your tools or downstream services cannot propagate trace context, record the trace reference at the application boundary and preserve the transaction identifier in the application’s own records. Do not label uncorrelated downstream activity as a child span unless your instrumentation actually establishes that relationship.

Step 3: Carry correlation through each tool call

For every tool invocation, decide which identifiers cross the boundary and which remain internal. Pass a transaction identifier to a downstream operation only when the receiving system can store or use it safely and the identifier has a defined purpose. If the tool accepts a business reference but not tracing context, store the relationship in your application’s transaction record. Avoid adding arbitrary fields to a vendor API request: use only fields supported by that system and its contract.

A tool span should let an operator answer: which operation ran, for which transaction, when it began, how it ended, and where to inspect the result. It should not require an operator to reconstruct the relationship from a full prompt transcript. Where operation names or error categories are user-controlled, normalize them to a bounded vocabulary so dashboards do not fragment into thousands of near-duplicate values.

For a multi-tool path, retain a clear sequence. If a tool reads information and another writes a change, their spans should be distinguishable. A read that succeeds does not establish that the later write succeeded. Likewise, if one user request triggers two business transactions, correlate each write tool with its own transaction and retain the shared request reference. This makes partial outcomes queryable without treating the whole agent turn as a single atomic action.

The boundary between agent tracing and tool-level tracing can be a separate implementation concern. If your tool protocol is MCP, use the MCP observability discussion for that specific tracing problem. Here, verify the resulting identifiers and transaction outcome in your own application rather than assuming a protocol-level span proves the downstream business effect.

Step 4: Verify a business outcome independently

Implement a verification function that reads from the system of record or another agreed authoritative source after the action. Its inputs should include the transaction identifier and the expected action details. Its output should be a structured result: verified success, verified failure, pending, or indeterminate. “Indeterminate” is important. A timeout while reading the record does not mean the write failed, and it does not mean the write succeeded.

Keep the verification evidence concise. Record the system reference, observed state, observation time, and the rule that mapped the observation to the status. When the evidence contains sensitive details, keep those details in the business system and store a reference or redacted summary in telemetry. A responder should be able to follow the reference through authorized tools without exposing the underlying record in a broadly accessible trace viewer.

Consider timing explicitly. A downstream write may be asynchronous: the tool returns an accepted response, while the durable state appears later. In that case, move the transaction to pending_verification and schedule a bounded follow-up check under your application’s operational policy. If the check reaches its deadline without authoritative evidence, transition to needs_review or another non-success state. Do not quietly convert a timeout into success to improve completion metrics.

Where an operation needs human approval, require approval before execution, bind approval to the exact payload being authorized, and expire it according to your process. A changed payload is a different action and should not inherit the previous approval. This tutorial does not prescribe an approval product or claim that a particular agent platform enforces these conditions; they are application-level controls for actions where authorization is required.

Step 5: Connect failures to transactions

Use failure categories that lead to different operational responses. At minimum, separate a tool invocation that returned an explicit error, a timeout with unknown downstream state, a verification failure, an application or instrumentation failure, and a policy or validation rejection. Record the stage at which the problem occurred. A single agent_failed label is not enough to determine whether a tool retry is safe or whether the business effect may already exist.

For each category, define the next state of the business transaction. A validation rejection can move to failed before any write occurs. A timeout after a write request should normally remain uncertain until a read-back or reconciliation determines whether the effect happened. A verification mismatch may require needs_review, especially if the application cannot distinguish a delayed update from a conflicting one. Preserve the initial error and the verification outcome as separate events so that later evidence does not erase how the issue began.

Retries need particular care. A retry can create duplicate external effects if the first attempt completed but its response was lost. Do not promise exactly-once behavior across an external system unless that system and your implementation provide a verified mechanism for it. Before retrying, check whether the operation is safe to repeat, whether a stable idempotency mechanism is supported by the receiving system, and whether the transaction can be queried for an existing result. If those conditions are not established, route the transaction for reconciliation rather than replaying blindly.

If an approval is required, do not treat a retry as automatically covered by an earlier approval. Confirm that the retry uses the exact approved payload and that the approval remains valid. Record the approval reference and payload version in the transaction record, not in a public trace attribute that may expose the payload itself.

Step 6: Validate the design with a hypothetical transaction

Consider a hypothetical internal service agent asked to create a work item from a request. The application assigns request R-2048 and transaction T-771, with action type create_work_item. The agent invokes a creation tool. The tool call times out after the downstream system may have accepted the request. The agent then returns a message saying it could not confirm completion. This scenario is illustrative; it does not describe a PADISO deployment or a tested product behavior.

A weak implementation would mark the transaction failed because the tool timed out, then let the user try again. If the first write actually completed, that second attempt could create a duplicate. Another weak implementation would mark it successful because the agent produced a reassuring final response. Neither decision is supported by the available evidence.

A stronger path preserves the timeout as the execution outcome and places T-771 in pending_verification. A read-back searches for the work item using a stable business reference that the application supplied to the target system, if that system supports it. If exactly one matching item is observed with the expected state, the application records succeeded and stores its reference. If none is observed, the transaction remains pending until the defined verification window ends or another authoritative check resolves it. If multiple matches appear, the status becomes needs_review; the application does not choose one arbitrarily.

The trace then helps explain the sequence: request handling, tool invocation, timeout, and later verification. The transaction record explains the business result: one matching item found, none found yet, or ambiguous duplicates. A dashboard can show both records together, but should display their separate states. An agent’s final message is a third observation: it communicates to the user, but does not replace either execution evidence or business verification.

Step 7: Build a focused operations view

Start with a small set of questions rather than a broad dashboard. For an individual request, an operator should be able to find the trace from the request or transaction reference, see which tool operation failed, and open the transaction record to inspect its last verified state. For a group of requests, the operator should be able to count transactions by outcome and failure category over a chosen period, with a path to the underlying records.

Separate technical success rates from business completion. For example, a tool span may have a successful response while business verification is pending. Track these as different measures. If you combine them, specify the rule and expose the underlying categories so leaders can see whether an apparent improvement reflects faster tool responses, better verification coverage, or genuinely more completed transactions.

Useful operational signals include the number of transactions awaiting verification beyond their expected interval, the age of the oldest unresolved transaction, repeated failures by tool and category, and the share of attempted actions with a verified outcome. Define the denominator carefully. If requests that never reach an action are included in a transaction-completion percentage, the metric answers a different question than the success rate of attempted transactions.

Keep content capture out of the default dashboard. Prefer bounded attributes such as environment, action type, tool name, outcome category, and transaction status. Restrict access to detailed traces and business records according to their sensitivity. These choices reduce unnecessary exposure and keep the primary operational view focused on routing and diagnosis rather than reading entire user conversations.

Azure architecture diagram and responsibility matrix

The flow below distinguishes execution tracing from business verification. A tool result is not automatically a completed transaction; the application checks an authoritative state and preserves an uncertain result for follow-up. The diagram is conceptual and does not assert that every component is a built-in Foundry feature.

flowchart TD
  accDescr: Workflow stages and decisions: User request, Application creates IDs, Foundry agent runs, Correlated tool call, Application verifies state, Verified transaction, Pending or review. The adjacent text explains the conditions and exceptions.
  accTitle: Tracing a Foundry Agent from User Request to Business Outcome workflow
    A["User request"] --> B["Application creates IDs"]
    B --> C["Foundry agent runs"]
    C --> D["Correlated tool call"]
    D --> E["Application verifies state"]
    E --> F["Verified transaction"]
    E --> G["Pending or review"]

Use the following matrix to assign implementation responsibility in an Azure deployment. It is an architecture worksheet, not a claim about automatic ownership by a cloud service. Confirm each boundary against the services and identity configuration in your actual environment.

ConcernApplication / agent teamAzure platform teamDownstream system ownerEvidence to retain
Request and transaction IDsCreate, validate, and propagate identifiersProvide approved conventions and runtime configurationAccept a business reference only where supportedRequest-to-transaction mapping
Trace instrumentationAdd supported instrumentation and span boundariesMaintain environment configuration and observability accessPropagate trace context if the interface supports itTrace reference and instrumentation version
Workload identityUse the configured identity without embedding credentialsProvision and maintain identity configuration and access pathGrant the narrow access needed for the operationIdentity reference, operation, and result category
Network pathSurface connection errors with the affected operationDefine and operate the network path to required servicesDocument reachable endpoint and response contractConnection outcome and target reference, not secrets
Business verificationDefine status transitions and perform read-backKeep observability dependencies availableIdentify authoritative state and query behaviorVerification time, state, and record reference
Sensitive contentMinimize attributes and redact application dataSet access and retention controls for telemetryProtect source business recordsRedaction policy and authorized reference

The matrix is useful because many apparent “agent tracing” incidents cross ownership boundaries. An application team may see a timeout, while a platform team can establish whether the configured network path was available and the downstream owner can determine whether the write arrived. Preserve enough identifiers for those teams to collaborate without copying the full business payload into every diagnostic system.

Identity and network diagnosis are not substitutes for transaction verification. A successful connection only says something about the path at a particular time; it does not establish that a requested state change persisted. Similarly, a successful identity check does not prove that the action was authorized for the exact payload or that the downstream system completed it. Keep these as distinct findings in the incident record.

If your requirements include recovery across a regional failure, treat that as a separate architecture exercise; the Azure agent disaster-recovery discussion addresses what must survive that event. For data separation between customers, see the tenant isolation treatment. Neither topic is a substitute for the request-to-transaction correlation designed here.

Step 8: Troubleshoot common gaps

The tool span exists, but the downstream work is not visible. Check whether the downstream service accepts trace context and whether your implementation propagates it. If it does not, use the application’s transaction record to connect the tool attempt with the downstream business reference. Do not manufacture a span relationship based only on similar timestamps.

The transaction identifier appears on the request but not on the tool operation. Find the point where the application constructs the tool call and verify that the identifier is available there. If the receiving system does not support a transaction field, retain the association at the application boundary and use a supported business reference for read-back. Avoid silently changing an external request contract to carry telemetry data.

The trace looks successful while the business status is pending. This may be correct: the tool returned normally, but verification has not established the durable outcome. Inspect the verification event, its observation time, and the authoritative record reference. Keep the status pending or indeterminate until the defined rule resolves it.

A trace is missing or split into unrelated pieces. Confirm that instrumentation is enabled in the environment handling the request and that correlation context survives each application boundary. Compare identifiers recorded at the request entry and tool invocation. If the execution crosses a queue or another asynchronous boundary, define how the application carries its own request and transaction references across that handoff rather than assuming the original in-process context is still present.

Dashboard fields change after an instrumentation update. Check the pinned instrumentation and semantic-convention versions, then review the current convention definitions before changing queries. Update dashboards deliberately and compare the old and new field mappings in a non-production environment. A field rename or semantic shift can look like an abrupt operational failure even when the underlying workload has not changed.

The trace contains too much content. Remove unneeded prompt and result capture, redact fields at the point they enter telemetry, and reduce access to any remaining detailed content. Retain identifiers and outcome categories needed for diagnosis. If a responder needs the original business record, direct them to its authorized system of record rather than duplicating it into broadly accessible traces.

Step 9: Test the correlation and failure path

Use a controlled test identity and a non-production target or a harmless operation. Before the test, write down the expected request identifier, transaction identifier, tool name, final business state, and verification evidence. This makes the expected result precise enough to inspect. Do not use a production business action merely to prove that spans appear.

Exercise at least three paths. First, run a normal action and confirm that the request, tool operation, and transaction can be connected and that the business state is independently verified. Second, cause a validation failure before the write and confirm that the transaction is recorded as failed without a false success state. Third, simulate an uncertain outcome using a safe test mechanism, such as a controlled interruption before verification, and confirm that the transaction remains pending or requires review instead of being declared complete.

For each path, inspect the trace and transaction record separately. Confirm that a trace can be found without searching through sensitive content, that the tool outcome is distinguishable from verification, and that the transaction status follows the documented state transition. Check that an operator can identify the next action: wait for a defined check, inspect the downstream record, or route the case for reconciliation.

Record the outcome of your own validation in your deployment notes. The steps above describe what to test; they do not claim that a test was run or that a particular environment has passed it. Repeat the checks after changing instrumentation, tool contracts, transaction state logic, or downstream verification behavior.

Troubleshooting decision artifact

Use this compact worksheet during design review or incident triage. Fill it in for one action type at a time; a generic answer such as “the agent did it” is not sufficient evidence.

DecisionRecord your answer
What is the intended business action?Name one durable state change.
What identifies this attempt?Record the request ID and transaction ID separately.
Which tool operation performs the action?Use a stable, bounded operation name.
What is the system of record?Name the source used to verify the result.
What evidence means success?Specify the expected state and matching reference.
What happens on timeout?Keep the outcome uncertain until a defined check resolves it.
When is human review required?State the ambiguity or deadline that triggers review.
What content must not enter telemetry?List sensitive fields and the redaction or omission rule.

Printable summary: Create a transaction before execution; connect the request, trace, and tool attempt with explicit references; classify failures by stage; verify the durable business state independently; preserve uncertainty rather than guessing; and route ambiguous outcomes to reconciliation. A green agent response is not a transaction receipt.

Implementation handoff

Assign an owner for each part of the path: application instrumentation, Azure identity and network configuration, downstream business verification, and operational dashboards. Agree on the identifier format, status vocabulary, verification deadline, sensitive-field handling, and retry decision before enabling real actions. A short design review is usually more valuable than adding a large collection of unbounded trace attributes that no responder knows how to interpret.

For teams that need help connecting application instrumentation, identity and network boundaries, and operational evidence into a maintainable Azure design, enterprise platform engineering is a relevant next step. Bring one representative transaction path and the worksheet above; the useful starting point is a concrete action, its authoritative outcome, and the failure state that currently leaves operators uncertain.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call