AWS Lambda maximum concurrency for SQS and AI API jobs
Lambda's SQS maximum concurrency caps how many instances one queue can invoke, from 2 to 1,000. How to set it so AI jobs stay within API limits.

AWS Lambda's maximum concurrency is a setting on an SQS event source mapping that caps how many instances of your function that queue can invoke at once, from 2 to 1,000. It is separate from your account's concurrent executions quota, which defaults to 1,000, and it should never be set above the function's reserved concurrency.
The Lambda facts come from AWS's own pages, read on 2026-09-28 and listed under Sources. The AI API side uses Sume's Generation admission docs as the example of a paid API with a concurrency limit.
Which settings limit Lambda concurrency?
Three numbers interact: the account quota, the function's reserved concurrency, and the per-queue maximum. AWS says the last two are independent, and that reserved concurrency must be at least the total maximum concurrency of every SQS source on the function, or Lambda might throttle your messages.
- Maximum concurrency can't be combined with provisioned mode, which uses minimum and maximum event pollers instead.
- Configuring it is free.
| Setting | Value |
|---|---|
| Account concurrent executions | 1,000 by default |
| SQS maximum concurrency | 2 to 1,000 per event source; empty turns it off |
| SQS mapping without a maximum | Up to 1,250 concurrent instances by default |
| Batch size (standard queue) | Up to 10,000 records; 10 for FIFO |
| Function timeout | Up to 900 seconds (15 minutes) |
| Queue visibility timeout | At least six times the function timeout |
How do I set maximum concurrency?
In the console, open the function, choose the SQS trigger, choose Edit, and enter a number between 2 and 1,000 under Maximum concurrency. From the AWS CLI, use update-event-source-mapping with --scaling-config:
aws lambda update-event-source-mapping \
--uuid "<event-source-mapping-uuid>" \
--scaling-config '{"MaximumConcurrency":5}' \
--function-response-types "ReportBatchItemFailures"Does maximum concurrency cap jobs running at an AI API?
Only if each invocation lasts as long as its job. If your function submits a job and returns, the setting caps simultaneous submits, not the jobs still running at the API, and a large queue drains into the API as fast as Lambda can submit.
Sume, for example, accepts jobs past its processing limit as queued while queue capacity remains, and refuses new paid submits with 429 queue_full only when both are full (Generation admission). The dashboard Concurrency tab shows your workspace's configured concurrency_limit.
- Backlog that fits in accepted jobs: submit and return. Put a
callback_urlon eachPOST /v1/videosand collect results with a webhook, as in AWS Lambda Function URL webhook for AI video. - Larger backlog: hold the slot. Set batch size 1 and maximum concurrency to your
concurrency_limit, and let each invocation poll its job to a terminal status, so the queue waits in SQS instead.
| Plan | Processing | Queue | Accepted jobs |
|---|---|---|---|
| Free | 1 | 5 | 6 |
| Pro | 4 | 20 | 24 |
| Startup | 8 | 40 | 48 |
| Scale | 20 | 100 | 120 |
What does a hold-the-slot handler look like?
This Python handler submits with the SQS message id as the Idempotency-Key, so a redelivered message gets the original job back instead of a new one (Video generation). It polls every 30 seconds, the interval Sume's docs suggest. Keep the function timeout above your longest job; Lambda allows at most 900 seconds, so this design fits jobs that finish inside 15 minutes.
import json, os, time, urllib.request
API = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
"Content-Type": "application/json"}
def call(url, body=None, key=None):
headers = dict(AUTH, **({"Idempotency-Key": key} if key else {}))
data = json.dumps(body).encode() if body else None
with urllib.request.urlopen(urllib.request.Request(url, data, headers), timeout=30) as r:
return json.load(r)
def lambda_handler(event, context):
failures = []
for record in event["Records"]:
try:
prompt = json.loads(record["body"])["prompt"]
job = call(API, {"model": "seedance-2", "prompt": prompt}, record["messageId"])
while job["status"] not in ("completed", "failed", "cancelled"):
time.sleep(30)
job = call(job["polling_url"])
except Exception:
failures.append({"itemIdentifier": record["messageId"]})
return {"batchItemFailures": failures}What happens when a submit fails?
- Without partial batch responses, one error makes every message in the batch visible again, including the ones that succeeded.
ReportBatchItemFailuresreturns only the failed ids; if the function throws, the whole batch still counts as failed. - A
400,401or402won't fix itself on retry. Let those messages reach a dead-letter queue: see What is a dead letter queue?. - With the hold-the-slot design, a
429 queue_fullmeans something else is using the workspace's slots: lower the maximum concurrency. In current Sume code a same-key retry replays thequeue_full, so redeliveries of that message fail the same way; resubmit it from the dead-letter queue under a new key.
Sources
Related posts
More in Integrations
- Azure AI Foundry MCP tool: connect an agent to Sume
Add an MCP tool to an Azure AI Foundry agent: a Custom keys connection for Sume's API key, then server_url, allowed_tools, and require_approval.
- Azure DevOps scheduled pipeline: a nightly cron in UTC
Add a schedules block with a UTC cron to the pipeline YAML, set always: true to run without code changes, and key any paid API call to the date.
- BullMQ retry: exponential backoff for paid API jobs
Set attempts and an exponential backoff on BullMQ jobs, stop early with UnrecoverableError, and key each paid API call to the job so retries replay.
- Celery task retry: backoff and jitter for a paid API call
Retry a Celery task with autoretry_for, retry_backoff and max_retries, wait out retry-after on a 429, and send one Idempotency-Key on every retry.
Written by Sume