SearchFIT.ai: Track and grow your brand in AI search
Back to Blog
How-to Guide 20 mins

Scaling Agent Workloads on AWS: Concurrency Is Only One Constraint

A practical AWS guide to queue design, quota-aware concurrency, downstream capacity, graceful degradation and recovery for agent workloads.

The PADISO Team ·

Prerequisites

Before changing concurrency settings, establish what an agent request does from admission to completion. You need an inventory of the model calls, tool calls, external services and state writes that may occur; a way to measure request rate, call rate, latency and errors; and an owner for each dependency’s capacity limits. Use representative traffic assumptions, not a peak number copied from a dashboard without context.

This guide treats the agent runtime as one part of a larger workload. A worker can accept more concurrent requests than a model quota, a connected service, or a business process can safely handle. The goal is therefore not maximum parallelism. It is a controlled rate of useful work, with a defined response when demand exceeds the capacity of any required dependency.

You should also be able to distinguish a request that is waiting, actively executing, retrying, completed, cancelled or permanently failed. If current records cannot tell those states apart, add that distinction before increasing throughput. Otherwise, a queue can conceal stalled work and retries can make a dependency incident worse.

Tip: Start with one representative workflow and its actual dependency map. A single agent may take different paths, so an average number of tool calls can hide a costly or unusually slow branch.

1. Map the workload before choosing a concurrency limit

An agent request is not necessarily one unit of downstream work. It may involve several model turns, a search, a read from an application service and a final write. A retry can repeat one of those calls, while a tool may itself fan out to additional services. Count work at the dependency boundary, not only at the user-request boundary.

Draw the path from request admission through execution and final result recording. For each edge, record whether it is synchronous or asynchronous, whether it can be repeated, what happens if it times out, and whether it has its own quota or capacity owner. Include both required calls and optional enrichment. This reveals which dependencies control the workload’s safe rate and which actions can be skipped when the system is under pressure.

A useful first-pass model has four separate limits. The first is how many requests may wait in the queue. The second is how many agent executions may run at once. The third is how many calls may be in flight to each dependency. The fourth is how much work each dependency can sustain over time. These limits interact, but they are not interchangeable. More workers cannot increase a downstream system’s capacity.

For each request class, estimate the number of calls to every constrained dependency. Keep a range where behavior varies. For example, ordinary requests might make two calls to an application service while a document-heavy path makes five. If the system admits traffic using only the average, the expensive path may consume the capacity reserved for everything else.

Separate throughput from latency. A service can have acceptable average response time while its tail latency is growing, tying up workers and increasing the number of in-flight calls. Conversely, a queue can grow during a short burst even when the service is healthy and has enough spare capacity to drain it afterward. A concurrency setting should be judged alongside arrival rate, service time, queue age and recovery rate.

2. Choose a queue and define what its states mean

A queue buffers temporary differences between incoming demand and the rate at which workers can complete work. It does not create capacity. If arrivals exceed completions for long enough, the backlog grows until the system must reject work, shed optional work, or accept increasing delay. Decide in advance which of those outcomes is acceptable for each request class.

Use explicit states such as accepted, queued, running, retry_wait, completed, failed and cancelled. Persist enough information to recover the work’s status if a worker stops unexpectedly. A state transition should have a timestamp and a reason, such as dependency_timeout or capacity_deferred, rather than a generic failure label. These details make it possible to distinguish slow work from work that will never progress.

Give each queued item a stable work identifier and a deduplication key appropriate to the business action. This is not a promise of exactly-once execution. A worker can complete an external action and fail before recording that completion, so a later attempt may repeat the action. Where repeated effects would be harmful, make the destination operation idempotent if possible, or put a separate reconciliation step around the effect.

Set a maximum queue age based on the value of the work, not just the maximum queue size. An interactive request that is no longer useful after a short wait should expire or return a clear deferred response; a background task may remain valuable for longer. When an item expires, record that disposition. Silently dropping old work makes the queue look healthy while leaving users and operators unable to explain missing outcomes.

A queue needs a bounded admission policy. Define a maximum backlog or oldest-item age for each class, then specify what happens at the boundary. Options include returning a retryable busy response, asking the caller to try later, or accepting only high-priority work. Do not let a queue grow without limit simply because storage is available; waiting work still consumes operational attention and may become stale.

Warning: Queue redelivery and application retries can multiply each other. Pick which layer owns each retry, cap attempts, and ensure that a repeated attempt cannot unknowingly repeat an external side effect.

3. Build a reference design with independent limits

The design below separates admission, agent execution and dependency access. A queue absorbs short bursts; a dispatcher admits only work that fits the current policy; workers run a bounded number of executions; and per-dependency controls protect downstream capacity. State records let operators see whether work is queued, active or finished. Identity should follow the workload boundary: the worker receives only the authority its task needs, and a tool connection should not be treated as a reason to grant broad access to every worker.

flowchart TD
  accDescr: Workflow stages and decisions: Request admission, Bounded work queue, Policy dispatcher, Agent worker pool, Per-service call limits, Downstream services, Work state and outcome. The adjacent text explains the conditions and exceptions.
  accTitle: Scaling Agent Workloads on AWS — Concurrency Is Only One Constraint workflow
    A["Request admission"] --> B["Bounded work queue"]
    B --> C["Policy dispatcher"]
    C --> D["Agent worker pool"]
    D --> E["Per-service call limits"]
    E --> F["Downstream services"]
    D --> G["Work state and outcome"]

accTitle: Queue-aware agent workload reference flow accDescr: Requests enter a bounded queue, pass through a policy dispatcher, and run in a bounded worker pool. Per-service limits control calls to downstream services, while workers record work state and outcomes.

The dispatcher is the place to apply admission decisions based on request class, queue age, worker capacity and dependency health. It should not infer that a request is safe merely because a worker slot is free. A worker can be idle while one critical downstream service is already at its safe call limit.

The per-service limit belongs close to the point where calls are made. A global worker limit is not enough if requests use different dependency paths. If one tool is slow, its calls should not occupy every slot needed by unrelated work. Conversely, creating separate worker pools without a shared dependency limit can allow each pool to overwhelm the same service independently.

Keep the state store and the execution queue conceptually distinct, even if a particular implementation combines storage responsibilities. The queue answers what work should be considered for execution; the state record answers what is known about a specific work item. Define how a worker claims work, how long a claim remains valid, and how a replacement worker determines whether to resume, retry or mark the item for review.

Where tools are connected through Amazon Bedrock AgentCore Gateway, Gateway connects agent tools and target integrations; inbound and outbound authorization are distinct concerns. Treat authorization on the incoming request path separately from the identity used to reach a target. Neither boundary substitutes for workload throttling or for the target service’s own capacity control. AWS Gateway concepts

This design is intentionally a boundary map rather than a prescribed list of AWS products. Select queue, compute and state services that fit the team’s operational model, workload duration and recovery needs. The architectural requirement is that admission, execution, downstream calls and outcome recording can be controlled and observed separately.

4. Calculate a starting operating budget

Start with the narrowest dependency, not with the number of workers you hope to run. Measure or obtain a defensible capacity figure for that dependency, then choose a lower operating budget that leaves room for variation and other consumers. AWS documents service quotas, but the usable rate for a particular workload also depends on the measured capacity of its downstream services; there is no universal scale guarantee for an agent workload. Check the applicable quota information and validate the whole dependency path before raising limits. Amazon Bedrock AgentCore service quotas

For each dependency, write down: its relevant request or concurrency limit, the share reserved for this workload, the expected calls per agent request, the measured or estimated call duration, and the action to take when the budget is reached. Keep separate budgets if different request classes have materially different call patterns. Revisit the figures when model behavior, tool routing, prompts or downstream service use changes.

A simple rate estimate is useful for a first configuration: sustainable agent requests per minute are approximately the dependency’s allocated calls per minute divided by the expected calls per agent request. Treat this as an estimate, not a capacity guarantee. Retries, uneven request paths, bursts and latency variation can all reduce useful throughput. Make those assumptions visible beside the limit rather than burying them in a configuration file.

Concurrency and rate are related through service time, but a concurrency cap is not a rate limit by itself. If calls take longer, the same number of active calls yields fewer completed calls per minute. If calls complete quickly, the same cap may issue requests faster than a downstream quota allows. Use a rate control for rate-based limits and an in-flight cap for concurrent-work limits; when both constraints matter, enforce both.

Document the source and date for every limit. Distinguish an AWS service quota from a target service’s limit, an internal allocation, and an observed operating threshold. A quota increase, if available, does not establish that a database, API, or human review process can absorb the resulting rate. Keep the lowest relevant boundary in the admission calculation.

5. Set concurrency and retries so they do not amplify overload

Apply at least two independent controls: a cap on active agent executions and a cap on calls to each constrained dependency. If tool calls are issued concurrently within one agent execution, count those calls against the dependency limit too. Otherwise, a modest number of agent workers can still generate a large burst of tool traffic.

Use separate pools or budgets for work that has different service-time or priority characteristics when contention would be harmful. A slow, optional research path should not occupy every execution slot needed by a short, time-sensitive lookup. Separation has a cost: more configuration and more risk that independently sized pools exceed a shared downstream budget. Preserve a common per-dependency ceiling across them.

Retries need an explicit policy for each failure class. A connection reset may justify a limited retry; a quota response should normally trigger reduced pressure and a wait; invalid input should not be retried; and an ambiguous write may require outcome reconciliation before another attempt. Use bounded attempts and increasing delays with jitter where retries are appropriate. Never retry every failure immediately, because synchronized retries can turn a brief outage into a larger burst.

Do not allow a request to hold a worker while sleeping through a long backoff if the architecture can return the work to a delayed state and free execution capacity. At the same time, avoid repeatedly cycling a failing item so quickly that the queue becomes a retry engine. Track original attempts separately from retry attempts and make retry exhaustion visible.

Treat quota responses, timeouts and downstream overload signals as feedback to reduce or pause dispatch to the affected dependency. A single global circuit breaker can be too blunt if a workflow has optional and required calls. Consider whether only one tool should be paused, whether requests can continue in a reduced mode, and what evidence is needed before restoring normal dispatch.

Pro tip: Reserve a small amount of capacity for health checks and recovery work if the dependency’s design permits it. A system that consumes every slot on new requests may have no room to finish, reconcile or diagnose work already in progress.

6. Define graceful degradation by user-visible outcome

Graceful degradation is a deliberate, bounded alternative to the full workflow. It is not a vague promise that the agent will keep working. For each request class, state which dependency is essential, which capability can be omitted, and what response the caller receives if the essential path is unavailable.

An optional enrichment call can be skipped while preserving a clearly labelled partial result. A required write should not be represented as successful if it was never completed. A request that cannot safely proceed may be deferred with its status and a reasonable next action. If the model returns a proposed action but the system cannot verify its execution, keep the business outcome unresolved rather than converting the model’s statement into a completion record.

Make degraded modes explicit in the result and in telemetry. A response can identify the unavailable capability, the time of the decision, and whether work remains queued. Avoid returning a normal-looking answer that conceals missing data or an unperformed action. That creates a downstream decision problem: operators and users cannot tell whether the system answered from complete inputs.

Prioritize by business value and dependency cost, not by arrival order alone if the workflow has known classes. A request needed to complete an existing transaction may deserve capacity before an optional new analysis. Define priority rules that are deterministic and observable, and prevent low-priority work from remaining queued forever. A priority policy should not silently bypass dependency limits.

For tools connected through a gateway, keep degradation decisions at the application workflow boundary as well as at the connection boundary. A tool being reachable does not mean it is currently safe to call at the desired volume. For deeper questions about the boundaries around browser or code execution, see AgentCore Browser and Code Execution: Designing Safe Boundaries. Keep that execution-safety design distinct from the queue and capacity controls described here.

Memory and long-lived state create a related but separate design problem. Keep queue state, execution status and retained agent memory distinguishable, so a backlog does not become an accidental memory store. For the deeper treatment of retention and tenant boundaries, see Agent Memory on AWS: Retention, Retrieval and Tenant Boundaries.

7. Instrument for backlog, saturation and recovery

A useful operations view answers four questions quickly: Is work arriving faster than it finishes? Which dependency is constraining progress? How old is the oldest useful work? Is the system recovering after pressure eases? Aggregate request success alone cannot answer these. A workflow may return quickly by rejecting most work, or appear successful while its queue grows behind the scenes.

Track arrivals, admissions, completions, queue depth, oldest-item age, active executions and queue wait time. Break down counts by request class and final disposition. At each dependency boundary, measure call rate, in-flight calls, latency distribution, timeout and quota-related errors, and retry volume. Where appropriate, correlate those measures with the dependency owner’s own health signals rather than assuming the agent service sees the whole picture.

Record a trace or correlation identifier across the request, queue item, worker attempt and downstream calls. Keep attempt identifiers distinct from the stable work identifier. That distinction helps explain whether one user request created multiple calls, whether a retry followed an uncertain outcome, and whether a final state was written after an external side effect.

Alert on conditions that predict user impact, not just raw queue size. Examples include oldest-item age crossing the request’s usefulness window, sustained arrivals exceeding completions, repeated quota responses, exhausted retries, and a dependency limit that remains saturated while the queue grows. Choose thresholds from the workload’s service expectations and observed behavior; do not copy a generic number into every alert.

Recovery is a separate operating condition. When a dependency becomes healthy, release work gradually rather than draining the entire backlog at once. Confirm that new arrivals and recovery traffic fit within the same downstream budget. Watch whether the oldest-item age falls, whether completion rate exceeds arrival rate, and whether retries are diminishing. A queue that is shrinking only because new work is rejected may need a different response from a queue that is draining through successful completions.

Maintain an operator action for each alert: reduce a dispatch limit, pause a request class, disable an optional branch, defer new work, or move affected items to review. Give those actions a defined owner and a way to restore the previous policy. Manual intervention without a recorded reason and expiry can leave a temporary emergency setting in place long after its assumptions are obsolete.

8. Work through an illustrative capacity case

Consider a hypothetical document-processing workflow. Assume it receives 80 requests per minute during a busy period, and each completed request makes three calls to one constrained downstream service. Assume that service has been measured for this workload at 400 calls per minute, and the team chooses an operating budget of 300 calls per minute to leave room for variation and other consumers. These figures are illustrative assumptions, not AWS quotas or a performance claim.

At three calls per request, the chosen budget supports an estimated 100 requests per minute before accounting for retries or uneven paths. At 80 incoming requests per minute, the nominal spare rate is 20 requests per minute. If the dependency becomes unavailable for 30 seconds and incoming work continues at 80 per minute, approximately 40 requests arrive during the interruption. When service returns, a controlled rate of 100 requests per minute against 80 new arrivals would take roughly two minutes to clear those 40 requests, assuming each request has the expected three calls, there are no retries, and the dependency continues to sustain the assumed budget.

That arithmetic gives a starting point for queue policy, not a forecast of real recovery time. If one request class makes five calls, it uses more of the budget. If requests take longer, in-flight calls can accumulate even when the average rate appears acceptable. If retries add 10 percent to the call volume, the effective request capacity falls. The dispatcher should therefore admit against the relevant call budget and current queue state, not rely on a fixed translation from requests per minute to calls per minute.

Suppose the team estimates a mean downstream call duration of 2.5 seconds. A rough concurrency estimate at 300 calls per minute is 12.5 in-flight calls: five calls per second multiplied by 2.5 seconds. That calculation uses an average duration and does not account for tail latency or burstiness. It suggests a starting in-flight cap near that range for measurement, not a setting to deploy unchanged. Use actual latency distributions and confirm that the rate limiter still enforces the calls-per-minute budget.

Now consider the counterexample: the team increases the worker pool from 20 to 100 because the queue is growing, while leaving downstream calls unconstrained. More requests begin, each creates several calls, and timeouts trigger retries. The queue may briefly shrink as work leaves it, but downstream latency and failures rise; workers remain occupied longer; retries add load; and completed business outcomes may fall. A lower queue depth at the start of this sequence is not evidence that the system has scaled successfully.

A better operating decision is to hold or reduce dispatch when downstream signals deteriorate, admit only the work that fits the dependency budget, and make a clear choice for the rest: defer it, reject it, or run an explicitly reduced path. Once the service recovers, increase traffic in controlled increments while checking both the dependency and queue recovery signals. The rate at which to increase is workload-specific; it should be established by observation, not asserted as a universal ramp schedule.

9. Implement the controls in a deliberate order

  1. Inventory the request paths. List each request class, its required and optional tool calls, expected call counts, external effects and state transitions. Include the retry path. Ask the dependency owner for the relevant limits and how they are measured. If an exact capacity figure is unavailable, mark it unknown and choose conservative admission behavior until it can be measured.

  2. Write down capacity budgets. For each constrained service, specify the allowed call rate and in-flight count, how much is allocated to this workload, and which other consumers share the limit. Record the assumptions about calls per request, service time and retries. Check applicable AgentCore quotas separately from downstream service limits; the documented quota is one boundary, not proof that the full workflow can sustain the same rate. Review the current AgentCore quota documentation.

  3. Define queue admission and expiry. Choose a maximum backlog or age for each class and set the outcome when that boundary is reached. Define whether callers receive a busy response, a deferred status or an accepted background-work identifier. Decide which queued work expires and how expired items are reported. Do not accept an item into a queue unless the system can later establish its state and disposition.

  4. Add bounded execution and dependency controls. Set an active-worker ceiling and enforce a separate rate and concurrency policy at each constrained dependency. Account for calls launched in parallel within an agent execution. Ensure separate worker pools cannot collectively exceed a shared downstream budget. Keep the limit configuration reviewable and tied to the assumptions recorded in step two.

  5. Make retries and uncertain outcomes explicit. Specify which failures can be retried, how many times, how long to wait, and when a human or reconciliation path is needed. Record attempts separately from work items. For actions that change external state, decide how the system detects a completed action whose confirmation was lost. Do not describe this design as exactly-once execution.

  6. Add degraded modes and operator actions. For every optional dependency, define whether the workflow can return a partial result and how it labels that result. For every required dependency, define how work pauses or fails safely. Write down who can pause a class, lower a limit or restore normal dispatch, and what signal supports that decision.

  7. Observe the whole path before raising limits. Review queue age and completion rate alongside dependency latency, call rate, in-flight work and retry volume. Increase capacity only when the relevant upstream quota, worker behavior and downstream service can support it. If the required platform controls, workload metrics or failure handling are not yet in place, cloud platform engineering can help design and implement those foundations.

  8. Reassess after meaningful changes. A new tool, prompt path, model behavior, request class or downstream integration can change calls per request and service time. Recalculate budgets and update dashboards and runbooks when those assumptions change. A limit calibrated to yesterday’s workflow is not automatically safe for tomorrow’s.

Warning: Do not make a production concurrency increase the first test of a new workflow path. Establish its dependency calls, failure behavior and measured capacity before admitting it at normal traffic.

10. Use this decision worksheet before changing a limit

The worksheet below is intended to be copied into a design review or operations runbook. Fill it in for each request class and each constrained dependency. Blank or unknown entries are decisions still to make, not implicit permission to scale.

DecisionRecord before release
Request class and ownerName the work and the team accountable for its outcome.
Required and optional callsList the dependency and expected call range per request.
Queue policyState maximum age or backlog and behavior when reached.
Execution policyState active worker limit and any pool separation.
Dependency budgetState allocated rate, in-flight limit and shared consumers.
Retry policyState retryable failures, attempt cap, delay and terminal path.
External effectsState how duplicate or uncertain outcomes are detected and reconciled.
Degraded resultState what is omitted, what remains valid and how the result is labelled.
Recovery signalState how operators know work is draining and normal operation is safe.
Change triggerState which workflow or dependency changes require recalibration.

A printable summary for a release review can be as short as this: Queue bounded? Yes or no. Oldest useful work defined? Yes or no. Worker and per-service limits separate? Yes or no. Retries bounded and observable? Yes or no. Partial or deferred outcomes explicit? Yes or no. Recovery owner and signals named? Yes or no. If any answer is no, assign an owner and resolve the gap before raising throughput.

Summary: scale the constrained path, not the worker count

A reliable agent workload needs bounded admission, visible work states, controlled execution, and independent limits at each constrained dependency. A queue can smooth a burst but cannot fix sustained overload. Worker concurrency can use available capacity, but it cannot manufacture downstream capacity. Retries can recover transient faults, but unbounded or synchronized retries can deepen an incident.

Start with the dependency map and measured capacity, translate those constraints into explicit rate and in-flight budgets, and decide what the system does when those budgets are reached. Instrument queue age and completion alongside dependency health. Make degraded outcomes honest, keep recovery gradual, and revisit the assumptions whenever the workflow changes. The appropriate scale is the rate at which the whole path can complete useful work without losing control of its backlog or downstream effects.

Want to talk through your situation?

Book a 30-minute call with Kevin (Founder/CEO). No pitch - direct advice on what to do next.

Book a 30-min call