AWS Lambda maximum concurrency for SQS and AI API jobs

Lambda's SQS maximum concurrency caps how many instances one queue can invoke, from 2 to 1,000. How to set it so AI jobs stay within API limits.

6 min readSume
All posts

AWS Lambda's maximum concurrency is a setting on an SQS event source mapping that caps how many instances of your function that queue can invoke at once, from 2 to 1,000. It is separate from your account's concurrent executions quota, which defaults to 1,000, and it should never be set above the function's reserved concurrency.

The Lambda facts come from AWS's own pages, read on 2026-09-28 and listed under Sources. The AI API side uses Sume's Generation admission docs as the example of a paid API with a concurrency limit.

Which settings limit Lambda concurrency?

Three numbers interact: the account quota, the function's reserved concurrency, and the per-queue maximum. AWS says the last two are independent, and that reserved concurrency must be at least the total maximum concurrency of every SQS source on the function, or Lambda might throttle your messages.

  • Maximum concurrency can't be combined with provisioned mode, which uses minimum and maximum event pollers instead.
  • Configuring it is free.
From AWS's SQS scaling, SQS event source and Lambda quotas pages, read 2026-09-28.
SettingValue
Account concurrent executions1,000 by default
SQS maximum concurrency2 to 1,000 per event source; empty turns it off
SQS mapping without a maximumUp to 1,250 concurrent instances by default
Batch size (standard queue)Up to 10,000 records; 10 for FIFO
Function timeoutUp to 900 seconds (15 minutes)
Queue visibility timeoutAt least six times the function timeout

How do I set maximum concurrency?

In the console, open the function, choose the SQS trigger, choose Edit, and enter a number between 2 and 1,000 under Maximum concurrency. From the AWS CLI, use update-event-source-mapping with --scaling-config:

aws lambda update-event-source-mapping \
  --uuid "<event-source-mapping-uuid>" \
  --scaling-config '{"MaximumConcurrency":5}' \
  --function-response-types "ReportBatchItemFailures"

Does maximum concurrency cap jobs running at an AI API?

Only if each invocation lasts as long as its job. If your function submits a job and returns, the setting caps simultaneous submits, not the jobs still running at the API, and a large queue drains into the API as fast as Lambda can submit.

Sume, for example, accepts jobs past its processing limit as queued while queue capacity remains, and refuses new paid submits with 429 queue_full only when both are full (Generation admission). The dashboard Concurrency tab shows your workspace's configured concurrency_limit.

  • Backlog that fits in accepted jobs: submit and return. Put a callback_url on each POST /v1/videos and collect results with a webhook, as in AWS Lambda Function URL webhook for AI video.
  • Larger backlog: hold the slot. Set batch size 1 and maximum concurrency to your concurrency_limit, and let each invocation poll its job to a terminal status, so the queue waits in SQS instead.
Plan defaults from Sume's Generation admission, read 2026-09-28.
PlanProcessingQueueAccepted jobs
Free156
Pro42024
Startup84048
Scale20100120

What does a hold-the-slot handler look like?

This Python handler submits with the SQS message id as the Idempotency-Key, so a redelivered message gets the original job back instead of a new one (Video generation). It polls every 30 seconds, the interval Sume's docs suggest. Keep the function timeout above your longest job; Lambda allows at most 900 seconds, so this design fits jobs that finish inside 15 minutes.

import json, os, time, urllib.request

API = "https://api.sume.com/v1/videos"
AUTH = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}",
        "Content-Type": "application/json"}

def call(url, body=None, key=None):
    headers = dict(AUTH, **({"Idempotency-Key": key} if key else {}))
    data = json.dumps(body).encode() if body else None
    with urllib.request.urlopen(urllib.request.Request(url, data, headers), timeout=30) as r:
        return json.load(r)

def lambda_handler(event, context):
    failures = []
    for record in event["Records"]:
        try:
            prompt = json.loads(record["body"])["prompt"]
            job = call(API, {"model": "seedance-2", "prompt": prompt}, record["messageId"])
            while job["status"] not in ("completed", "failed", "cancelled"):
                time.sleep(30)
                job = call(job["polling_url"])
        except Exception:
            failures.append({"itemIdentifier": record["messageId"]})
    return {"batchItemFailures": failures}

What happens when a submit fails?

  • Without partial batch responses, one error makes every message in the batch visible again, including the ones that succeeded. ReportBatchItemFailures returns only the failed ids; if the function throws, the whole batch still counts as failed.
  • A 400, 401 or 402 won't fix itself on retry. Let those messages reach a dead-letter queue: see What is a dead letter queue?.
  • With the hold-the-slot design, a 429 queue_full means something else is using the workspace's slots: lower the maximum concurrency. In current Sume code a same-key retry replays the queue_full, so redeliveries of that message fail the same way; resubmit it from the dead-letter queue under a new key.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume