Rate limiting is a capacity contract, not an exceptional edge case. A resilient integration accepts work without losing it, schedules outbound calls within the provider's documented rules, and slows down when the provider signals pressure. The queue protects intake; idempotency protects side effects; retry policy and observability determine whether the backlog recovers or becomes a retry storm.
Model limits at the scope the provider enforces
Record limits by endpoint, tenant, credential, region, and time window when the documentation distinguishes them. Capture response headers that report remaining capacity or reset time and honor `Retry-After` when supplied. Do not copy a limit from another plan or account. Keep the configured rate below the published ceiling until real traffic shows how bursts, pagination, and concurrent workers behave.
Decouple durable intake from outbound execution
Authenticate and validate the inbound envelope, persist it with a stable event ID, enqueue a reference, and return the acknowledgement required by the source provider. Workers should retrieve the durable record and perform one bounded destination operation. If the queue publish fails, the receiver must not report success unless another durable recovery mechanism exists.
A queue absorbs demand; it does not create downstream capacity. Admission control and worker pacing still decide whether the backlog converges.
Pace workers with shared, tenant-aware state
A token bucket or scheduled dispatcher can smooth calls, but the limiter must be shared by every worker using the same quota. Partition fairly so one noisy tenant cannot starve others. Cap concurrency as well as requests per interval, because long-running calls can exceed connection or provider capacity even when request counts appear compliant. Reduce the rate when throttling rises and recover gradually.
Retry only failures likely to recover
Retry throttling, timeouts, and documented transient server failures with exponential backoff, jitter, and a maximum attempt or elapsed-time budget. Do not retry validation, authentication, permission, or schema errors without a state change. Implement retry in one layer; nested SDK, worker, and workflow retries can multiply requests unexpectedly. Mutating calls must be idempotent before automatic retry is enabled.
Design the dead-letter queue as an operating process
Choose the redrive threshold from provider behavior and recovery objectives rather than a universal attempt count. Store event ID, tenant, destination operation, sanitized error, attempt history, first and last failure time, and code version. Alert when messages arrive or age beyond policy. Before redrive, classify the cause, repair configuration or data, and replay a small canary at a controlled velocity.
Batch only when semantics and APIs support it
Use a documented bulk endpoint when it preserves the required per-record result and idempotency behavior. Bound batches by the provider's current item and payload limits, isolate tenant data, and record which items succeeded. Split and retry only failed or transient records when the API supports partial results. Do not delay urgent work merely to fill a batch, and do not invent a batch wrapper around an endpoint designed for single writes.
Understand the retry storm before it is expensive
The retry storm is the failure mode that turns a temporary capacity squeeze into an outage, and its mechanics are worth internalizing. A provider throttles, the retry layer schedules many deferred attempts, those attempts arrive simultaneously after their backoff windows expire, they trigger more throttling, and each round triples or quadruples demand instead of draining it. A worker that retries every failed call on a fixed schedule against an endpoint processing ten requests per second can quietly generate five times that volume from a thirty-second incident, and the provider's documented ceiling becomes a soft floor for the noise. Three controls contain it. Backoff must include jitter so synchronized retries scatter rather than land together; a full-jitter policy keeps attempts evenly distributed across the window instead of clustering at window boundaries. Retries must have a hard per-message budget, after which the message enters the dead-letter queue instead of re-joining the fight. And the system must detect the storm early, using the throttle rate and retry amplification signals described below, so operators can apply admission control before the backlog diverges. A well-designed system recovers from a ten-minute incident in a few minutes; a storm-prone system carries the same incident for an hour.
Alert on leading indicators, not just depth
Queue depth is a lagging measure: by the time it crosses a threshold, the incident already happened. The useful alerts fire earlier. Throttle rate trending above a few percent of dispatch rate predicts a queue that will grow even under stable demand. Retry amplification, the ratio of retry attempts to original attempts, above roughly 1.2 indicates compounding demand before the backlog visibly swells. Queue age, the difference between the oldest message and the target service level, catches a slow divergence that depth never flags. Throttled-then-idle oscillation, where dispatch swings between zero and full pace, usually means the limiter or the backoff parameters are fighting the provider's real capacity rather than matching it. Configure alerts on these four signals with graduated severity, keep each alert actionable by naming the likely cause and the first mitigation, and review alert history monthly because a queue system that pages often and recovers slowly is a capacity design problem, not an operations problem.
Prove recovery under controlled load
Test sustained demand above the configured rate, sudden bursts, quota changes, 429 responses with and without retry guidance, timeouts after an unknown write result, worker crashes, duplicate delivery, poison messages, DLQ redrive, and one tenant monopolizing traffic. Monitor queue depth and age, dispatch rate, throttle rate, retry amplification, success latency, DLQ inflow, and reconciliation mismatches. Capacity is healthy only when the oldest work returns toward the target after the burst ends.


