SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 20 mins

Scaling Foundry Agents: Plan for Quotas, Queues and Backpressure

A practical method for sizing Foundry agent throughput, setting concurrency limits, and using queues and backpressure to protect users when demand spikes.

The PADISO Team ·

Prerequisites

Before changing a deployment, gather four things: a defined user-facing workload, recent or defensible estimates of request volume, the model and deployment choices under consideration, and an owner who can pause or reduce traffic. You do not need a perfect forecast. You do need to distinguish measured facts from assumptions so that a sizing decision can be revised when production behavior differs.

This guide focuses on model throughput and concurrency budgets: how much work reaches a model, how much may be in flight, and what the system does when demand exceeds either limit. It does not choose an enterprise agent platform or design tenant isolation and tool ownership. If platform selection is still open, compare the options in Microsoft Foundry, Bedrock AgentCore, or Gemini Enterprise: Choosing an Enterprise Agent Platform. Treat the selected model deployment and agent execution path as inputs to this exercise.

Collect, where available, request counts by workload, timestamps, input and output token distributions, end-to-end and model-call durations, retry counts, and failure categories. Keep the time interval consistent. A daily average can conceal a sharp five-minute burst; a peak from a one-off test can exaggerate normal demand. If production data does not exist, label estimates as assumptions and assign someone to replace them with observations before raising limits.

Agree on what counts as a request. One user action may trigger several model calls, tool calls, or follow-up generations. Count the unit whose capacity you are planning—usually a user task and, separately, each model attempt. Also define the user-visible deadline, the maximum acceptable queue wait, and which work can be delayed, rejected, or completed in a reduced form.

Tip: Make a small workload inventory before doing arithmetic. “Agent traffic” is not one workload if interactive support, overnight document processing, and an internal analyst assistant have different urgency, token sizes, or retry behavior.

Step 1: Separate user tasks from model attempts

Start with a request map. For each workflow, record the number of model calls per task, the expected token volume for each call, whether calls are sequential or parallel, and what could cause a retry. A task that makes three sequential calls consumes three attempts but has a different latency profile from a task that launches three calls concurrently. Both details matter when you budget throughput and active work.

Use a compact table such as this, replacing the illustrative entries with your own observations or explicit assumptions:

WorkloadTasks per minute at busy periodModel calls per taskInput tokens per callOutput tokens per callUser deadlinePriority
Interactive answer1821,60045020 secondsHigh
Case summary932,40070090 secondsNormal
Nightly classification30190012030 minutesLow

These numbers are a worksheet example, not a description of a Foundry limit or a recommended workload. Their purpose is to expose differences that a single requests-per-minute figure hides. If one interactive task fans out into parallel model calls, mark the fan-out explicitly; counting only the parent task understates simultaneous demand.

For each workload, distinguish normal arrivals from bursts. Record the typical busy interval, a credible short peak, and how long that peak might last. “Twice the daily average” is not a useful peak assumption unless it is connected to a real event, such as a scheduled batch, a staff shift change, or a product notification. If the trigger is unknown, say so and choose a conservative test range rather than presenting a guess as a forecast.

Map retries as well as first attempts. A timeout may cause a client to retry while the original request is still running. A worker may retry after an ambiguous response. Either can increase model demand even though the number of user tasks has not changed. Capture the retry reason, attempt number, and parent task identifier in operational records, while avoiding unnecessary sensitive prompt content.

Model calls and tools should also be counted separately. A tool call may increase wall-clock duration without using model throughput; a model call made to interpret the tool result does consume model capacity. For a shared tool layer, define which component owns the call boundary and its versioning rather than assigning tool delay to the model budget. The separate design question is covered in A Shared MCP Tool Layer in Foundry: Ownership, Versions and Access.

Step 2: Build a throughput budget from demand

Throughput is the amount of work a system must complete over time. For model planning, do not use only user tasks per minute. Estimate model attempts per minute and token demand per minute, since two workloads with the same attempt count can have very different token sizes.

A first-pass calculation is:

Model attempts per minute = user tasks per minute × model calls per task × expected attempts per call.

For token demand, calculate input and output separately:

Input tokens per minute = attempts per minute × average input tokens per attempt.

Output tokens per minute = attempts per minute × average output tokens per attempt.

Use separate averages for workload classes. Combining a short classification prompt with a long case summary into one overall average can mislead you whenever the workload mix changes. If token distributions are skewed, record a median and a high percentile as well as the mean. The mean helps estimate aggregate consumption; the upper tail helps identify tasks likely to run long or create a burst of active work.

Consider this explicitly hypothetical workload during a busy minute: 18 interactive tasks each use two model calls, while nine case summaries each use three. Assuming one attempt per call, that is 36 plus 27, or 63 model attempts in the minute. If a planning assumption adds a 10 percent retry allowance, the arithmetic becomes 69.3 expected attempts per minute. Since a real minute cannot contain a fraction of an attempt, use the figure as an average planning estimate, not as a literal cap. Measure actual retry rates and keep burst headroom separate from this allowance.

For the same example, the interactive work produces 57,600 input tokens and 16,200 output tokens per minute: 36 calls multiplied by 1,600 input and 450 output tokens. Case summaries produce 64,800 input and 18,900 output tokens per minute: 27 calls multiplied by 2,400 input and 700 output tokens. The combined first-attempt estimate is 122,400 input and 35,100 output tokens in the busy minute. Those figures do not establish that a particular deployment will accept that volume. They describe the demand that must be compared with the account’s actual applicable limits and observed behavior.

Do not convert an account limit into a production guarantee. Limits can depend on the resource, model, and deployment, so check the current limits applicable to the account and chosen deployment rather than relying on a universal number. Microsoft’s Foundry quota and limits guidance is the starting point for that verification.

Plan for composition changes. If a product launch shifts the request mix toward summaries with longer prompts, attempt volume might remain stable while token demand rises. If a new workflow adds a model call to every task, token demand and attempts both rise. Put workload-specific assumptions in the worksheet so an owner can see which change invalidates the estimate.

Step 3: Convert throughput into a concurrency budget

Concurrency is the number of requests active at the same time. It is not interchangeable with throughput. A system can have modest requests per minute but high concurrency when each request takes a long time. Conversely, a fast request stream can achieve significant throughput with comparatively few requests in flight.

Use Little’s Law as a planning relationship, not as a capacity guarantee:

Average in-flight requests ≈ arrival rate per second × average time in seconds.

If a workload averages 1.05 model attempts per second and a model call takes 8 seconds, the illustrative average is about 8.4 calls in flight. If the same calls take 20 seconds, the estimate becomes 21. This is an average under stable conditions. It does not cover bursts, uneven service times, timeouts, or a queue that is growing because arrivals exceed completions.

Estimate concurrency separately by workload and call stage. A task may spend several seconds waiting on a tool, then issue a short model call. Keeping a worker occupied during tool delay is different from consuming a model request slot. Record both the number of active user tasks and active model attempts. If the architecture cannot observe them separately, treat that blind spot as a measurement problem to resolve before increasing concurrency.

Set a ceiling for each layer rather than one large global number. Candidate controls include the number of tasks accepted by the application, the number of worker tasks allowed to run, and the number of model calls those workers may issue concurrently. The exact mechanisms depend on your implementation; the design objective is to prevent a burst in one layer from becoming an uncontrolled burst in the next.

A useful initial rule is to choose a conservative concurrency cap below the highest level you believe the deployment could tolerate, then increase it through controlled observation. The cap should protect latency and stability, not merely maximize activity. A large cap can make a queue appear shorter while shifting overload downstream, where requests wait longer, fail together, or trigger retries.

Warning: Do not multiply an observed average duration by a peak arrival rate and label the result a safe concurrency limit. It is only a rough estimate. Validate behavior with representative traffic, including the long-running portion of the workload, and stop if latency or failures worsen.

Step 4: Choose the queue boundary and define backpressure

A queue separates accepting work from executing it. It can absorb a temporary mismatch between arrival and completion rates, but it cannot create capacity. If new work arrives faster than workers complete it for long enough, queue age and backlog grow until a limit is reached. Decide in advance whether the application should wait, defer, shed, or reject work when that happens.

Place the queue where the system can safely represent a task before asking the model to perform it. A queued task needs a stable task identifier, workload class, creation time, priority, deadline or expiry, and enough input reference to retrieve the work. Store only what the application needs to resume processing. Make the state transition from accepted to running explicit, and record completion or terminal failure so that a worker restart does not turn uncertainty into an untracked duplicate.

A queue can support different experiences for different work. An interactive request may wait briefly and then return a clear busy response or an accepted-for-later status. A batch task may remain queued and be processed when capacity becomes available. A deadline-sensitive task may be rejected before it becomes stale. Choose those behaviors with the product owner: a silent wait is not a neutral fallback if the user expects an immediate answer.

Backpressure is the mechanism that slows or refuses upstream work when downstream capacity is constrained. It can mean pausing batch producers, lowering the number of workers, limiting new interactive requests, or returning a retryable response. It is not the same as retrying everything. An overloaded dependency often needs fewer attempts, not more.

A simple control flow is shown below. “Capacity available?” must consider both the application’s concurrency budget and the currently applicable downstream constraints; an affirmative answer is not a promise that every request will succeed.

flowchart TD
    accTitle: Model workload admission and recovery
    accDescr: A task is admitted only when the queue and downstream budget allow it; otherwise it is delayed or declined. Admitted work runs within a cap, then completes or follows a bounded failure path.
    A["New task arrives"] --> B["Check deadline and demand"]
    B --> C{"Capacity available?"}
    C -->|"Yes"| D["Queue within budget"]
    C -->|"No"| E["Defer or decline"]
    D --> F["Run within concurrency cap"]
    F --> G{"Completed?"}
    G -->|"Yes"| H["Record result"]
    G -->|"No"| I["Bound retry or stop"]

The diagram is a logical design, not a claim about a built-in Foundry queue or traffic control. The queue, admission policy, and worker limits belong to the application architecture. Ensure that “defer or decline” produces a user-visible outcome and that a failed attempt does not loop indefinitely. A bounded retry should have a maximum attempt count, a delay policy, and a terminal state that an operator can inspect.

When deciding whether to retry, distinguish a transient failure from a request that is too large, invalid, expired, or persistently over capacity. Retrying a permanent or capacity-related failure without reducing pressure can amplify the incident. Preserve the original task deadline across attempts: a retry should not silently turn a 20-second interactive promise into an unbounded wait.

Step 5: Apply the budget to an Azure architecture

A practical Azure-specific design can be expressed without tying the capacity plan to one deployment product: an application accepts a task, applies an admission policy, places eligible work on an Azure queue, and a bounded worker pool consumes it. The worker calls the selected Foundry model deployment, records the attempt outcome, and returns or exposes the result through the application. Monitoring covers the queue, worker, and model-call boundaries so that a delay can be attributed to the correct stage.

Treat “Azure queue” and “worker pool” as architecture roles in this diagram and worksheet, not as a claim that a particular service or configuration is required. Choose concrete services only after checking their operational fit, deployment constraints, and applicable limits. If the agent uses custom code packaged as a container, Foundry describes that as a hosted agent; prompt agents use declarative prompts and tools. These are different execution approaches, not a built-in traffic-splitting mechanism. Hosted agent concepts provide the relevant distinction.

The application’s admission layer should decide whether to accept work before the worker dispatches a model call. The worker should enforce its own active-work ceiling as a second boundary, because upstream callers may be misconfigured or a queue consumer may restart. The model-call wrapper should record attempt start, end, outcome category, and workload class. Together, these controls make it possible to compare accepted tasks, queued tasks, active calls, and completed calls rather than inferring system health from a single request counter.

Keep the identity and network responsibilities visible, but do not let this capacity plan become a substitute for their design. The application team owns the task contract and user-facing admission behavior; the platform team owns the queue and worker deployment controls it operates; the model operations owner verifies the applicable deployment limits and observes model-call outcomes. The identity and network owners confirm that the intended paths and credentials are supported in the chosen design.

BoundaryCapacity responsibilityIdentity and network responsibilityEvidence to inspect during an incident
Application ingressApply task-level admission, priority, and deadline rulesAuthenticate the caller and allow the intended routeAccepted, deferred, and declined task counts; response latency
QueueBound backlog and expose age, not just depthRestrict access to intended producers and consumersOldest task age, depth trend, enqueue and dequeue rates
WorkerEnforce active-task and model-call ceilingsUse its approved access path to the model endpointActive workers, in-flight calls, duration, failure category
Foundry deploymentConfirm the applicable model and deployment limitsConfirm endpoint reachability and authorized accessCall outcomes and the current applicable limit information
Result handlingPrevent stale work from appearing complete; record terminal stateReturn results only through the intended application pathCompletion state, task age at completion, duplicate indicators

This matrix is deliberately about responsibility boundaries, not detailed identity configuration. Keep tenant separation in its own architecture decision; see Tenant Isolation for Enterprise Agents on Azure. Likewise, determine which identity may perform each action in the dedicated identity design, Entra Identity for AI Agents: Who Is Allowed to Do What?. For capacity planning, the operational requirement is to know which owner can change each limit and which signal proves that the change had its intended effect.

Step 6: Work through a hypothetical sizing decision

Assume a fictional service receives the interactive and case-summary traffic in Step 2 during a recurring 10-minute peak. Assume its business owner accepts a 20-second deadline for interactive answers and 90 seconds for summaries. These figures are illustrative only; they are not measurements, performance claims, or limits for any model. The team estimates 63 first-attempt model calls per busy minute, with the separate 10 percent retry allowance used for capacity planning.

First, separate urgency. Interactive calls cannot simply join an unlimited batch backlog without breaking the stated experience. Case summaries can wait, provided their age remains within the business deadline. The team therefore creates distinct workload classes and a policy that reserves capacity for interactive work while allowing summaries to use spare capacity. This is a proposed design choice, not a Foundry feature. The reservation should be tested against real traffic; setting it too high can leave capacity idle while batch work grows, and setting it too low can starve interactive users.

Second, estimate in-flight work from observed or assumed durations. If the 36 interactive calls in a busy minute take an average of eight seconds, their rough average concurrency is 4.8 calls: 0.6 attempts per second multiplied by eight seconds. If the 27 summary calls average 18 seconds, their estimate is 8.1 calls: 0.45 per second multiplied by 18 seconds. Combined, that is about 13 average active calls before the retry allowance. The arithmetic does not model a synchronized burst, a long tail, or downstream rejection. It is a starting estimate for selecting a conservative cap, not a safe setting by itself.

Third, assign a bounded queue budget to summaries. Suppose the service owner decides, illustratively, that a summary must begin within 30 seconds to finish inside its 90-second target. The team can monitor oldest queue age and stop admitting low-priority summaries as that threshold approaches. It should not accept an unbounded number of tasks merely because storage allows it. If the projected wait would consume the work’s deadline, return a clear deferred or declined outcome rather than leave stale tasks to consume capacity later.

Fourth, define what happens during a model-call slowdown. If the average call duration doubles, the same arrival rate roughly doubles average in-flight work under the relationship used above. The worker cap then limits active calls, while the queue absorbs only a short transient. The team pauses or reduces batch intake first, inspects whether interactive latency is affected, and avoids broad retries until it understands the failure category. If capacity remains constrained, the user-facing policy takes precedence over quietly accumulating work.

Finally, define an acceptance review before raising the cap. Record a baseline for queue age, active calls, completion time by workload, and retry rate. Change one control at a time, observe comparable traffic, and have a rollback value ready. A proposed increase is acceptable only if it improves the intended measure without breaching the deadline, increasing repeated failures, or allowing backlog to trend upward. This is a decision rule, not a claim that a particular test has been performed.

Step 7: Set operational controls and failure responses

A useful operating plan connects each signal to an action and an owner. Queue depth alone is insufficient: a large queue of short, recent tasks may be healthy, while a smaller queue containing expired work may not be. Track age of the oldest task, completion rate, arrival rate, active work, model-call duration, failure categories, and retries. Break these down by workload class so batch behavior does not hide interactive degradation.

Establish thresholds from the workload’s deadlines. For example, an interactive policy might stop accepting new work when its predicted wait would exceed the stated user deadline. A batch policy might pause its producer when the oldest task reaches a predefined share of the allowed completion window. These thresholds are application decisions; do not present them as universal Azure or Foundry defaults. Name an on-call owner and make the action executable: pause batch intake, lower worker concurrency, disable a nonessential workflow, or communicate a delay.

Plan for at least four failure patterns. First, sustained arrivals exceed completions, so backlog grows continuously. The response is to reduce or reject arrivals, prioritize eligible work, or add capacity only after checking downstream constraints. Second, call duration rises, increasing active concurrency and queue age even when arrivals are unchanged. The response is to cap dispatch and inspect the slow stage rather than blindly increasing workers.

Third, a temporary failure causes a retry wave. Use bounded retries with delay and a cap, and ensure that workers do not all repeat the same work immediately. If the error indicates a persistent capacity limit, stop or slow dispatch and reassess the budget. Fourth, a worker loses its connection after the model call may have completed but before the result is recorded. Treat the outcome as uncertain: reconcile task state and avoid assuming that repeating the operation is harmless. Do not promise exactly-once external effects. For any action outside the model call, design idempotency or an explicit reconciliation process appropriate to that action.

A queue backlog can also hide a product failure. If users receive an “accepted” response but no result before their deadline, acceptance counts may look healthy while the service is not meeting its contract. Measure from task creation through result availability, not only from worker start to model response. Expire work that can no longer satisfy its purpose, and communicate the terminal state rather than silently dropping it.

If the system uses an agent implementation with a separately deployed execution component, include its concurrency and scaling behavior in the same end-to-end budget. Do not assume that a model-call limit alone bounds agent work: worker tasks may be waiting on tools, holding resources, or executing non-model steps. Conversely, increasing worker count does not necessarily increase model throughput if the model deployment is already the constrained stage.

Step 8: Review the budget when workload or deployment changes

A capacity budget is a versioned set of assumptions, not a one-time approval. Revisit it when a model or deployment changes, when prompts become longer, when a workflow adds a model call, when a new user group arrives, or when a batch schedule moves into the interactive peak. A change that seems functionally small can alter attempts per task or the distribution of durations enough to invalidate the old concurrency estimate.

Review actual workload by class against the original worksheet. Compare tasks per interval, model calls per task, input and output token distributions, retry frequency, and time spent at each boundary. If the measured retry rate exceeds the planning allowance, determine whether the cause is client behavior, transient failures, or work that should not be retried. If actual durations have a long tail, calculate concurrency using representative upper-percentile behavior as well as the average, then validate with controlled traffic.

Check the applicable quota and deployment information whenever the model or resource changes and before a planned increase in traffic. The current limits are account- and deployment-dependent; a value copied from another environment is not a substitute for checking the one you will operate. Keep a record of who verified the limit, when, which deployment it applies to, and what assumptions the capacity calculation uses. Avoid embedding an unverified number in application logic as if it were permanent.

Capacity changes should have a clear owner, an observation window, and a reversal condition. If the queue age increases, failures cluster, or interactive deadlines are missed, return to the previous cap or reduce intake while investigating. Avoid changing worker count, retry policy, prompt size, and admission thresholds simultaneously; otherwise, it becomes difficult to identify which change caused the result.

Printable capacity worksheet

Use this worksheet in a design review or operational handoff. Fill it in per workload and record whether each value is measured, estimated, or still unknown. It is intended to be copied into your own working document; it is not a downloadable artifact.

  • Workload and owner: Name the user task, business owner, technical owner, and the user-visible completion condition.
  • Arrival profile: Record normal rate, busy-period rate, credible burst, burst duration, and the source of each estimate.
  • Call shape: Record model calls per task, sequential or parallel behavior, input and output token distributions, and non-model waits.
  • Retry assumptions: Record the retryable failure categories, maximum attempts, delay behavior, and observed or assumed retry rate.
  • Throughput estimate: Calculate model attempts and input/output token demand per interval for each workload class. Keep first attempts separate from retry allowance.
  • Concurrency estimate: Record call-duration averages and upper-tail observations, estimate in-flight calls, and choose a conservative initial cap for validation.
  • Queue policy: State which work may wait, maximum acceptable age, expiration behavior, priority treatment, and what users see when admission is deferred or declined.
  • Backpressure action: Name the signal, threshold, owner, and exact action for pausing, reducing, or refusing work.
  • Azure responsibility boundary: Name the application, queue, worker, model operations, identity, and network owners; identify the record each owner checks during an incident.
  • Applicable limits: Record the model and deployment being checked, the date of verification, and the current account-specific information used in the plan.
  • Acceptance and rollback: Define the baseline, observation window, success criteria, rollback condition, and person authorized to reverse the change.
  • Review trigger: List workload, prompt, model, deployment, schedule, and user-population changes that require recalculating the budget.

A completed worksheet should let an engineer answer three practical questions quickly: what work is allowed in, how much can run at once, and who acts when the allowed budget is exhausted. If any answer depends on a value that has not been checked or a behavior nobody owns, record that as a design gap rather than hiding it in a capacity estimate.

Summary: keep demand, concurrency, and waiting connected

Start by counting user tasks and model attempts separately, then estimate input and output token demand by workload. Convert arrival rates and realistic durations into an initial concurrency estimate, but treat the result as an approximation that needs observation. Apply explicit caps at admission, worker, and model-call boundaries so one layer cannot overwhelm the next.

Use queues to absorb short mismatches, not to disguise sustained overload. Set age and deadline limits, define what users see when work is deferred, and pause or decline work when it can no longer finish usefully. Bound retries, measure the full task journey, and make failure ownership clear across the Azure application, queue, worker, and model deployment.

Before increasing traffic or concurrency, verify the applicable deployment limits, compare the workload with the assumptions in the worksheet, and define a rollback condition. Teams that need help translating those boundaries into an operable Azure design can explore enterprise platform engineering as a focused next step.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call