DynamoDB conditional put: dedupe Sume job_id webhook retries

A PutItem with attribute_not_exists(job_id) makes the first Sume webhook the only one that proceeds. Node sample with a two-state record and TTL.

5 min readSume
All posts

Use job_id as the partition key and write each webhook with PutItem and ConditionExpression: "attribute_not_exists(job_id)". A ConditionalCheckFailedException means the item exists, so this delivery is a duplicate: return 2xx and stop. The first delivery succeeds and proceeds.

Sume's webhook docs tell receivers to treat job_id as the idempotency key, because delivery makes up to 10 attempts and Redeliver can send the same terminal event again later.

Claim, then finish

The sample claims the job with state: "claimed", and a later step flips it to done. That separates the dedupe from the work, so a crash after the claim does not silently drop the job.

import { DynamoDBClient, PutItemCommand, UpdateItemCommand } from "@aws-sdk/client-dynamodb";
const db = new DynamoDBClient({});
const TableName = process.env.TABLE ?? "sume_jobs";
export async function claim(event) {
  const expires = Math.floor(Date.now() / 1000) + 30 * 86400;
  try {
    await db.send(new PutItemCommand({
      TableName,
      Item: {
        job_id: { S: event.job_id },
        state: { S: "claimed" },
        body: { S: JSON.stringify(event) },
        expires_at: { N: String(expires) },
      },
      ConditionExpression: "attribute_not_exists(job_id)",
    }));
    return true;
  } catch (e) {
    if (e.name === "ConditionalCheckFailedException") return false;
    throw e;
  }
}
export const finish = (jobId) => db.send(new UpdateItemCommand({
  TableName,
  Key: { job_id: { S: jobId } },
  UpdateExpression: "SET #s = :d",
  ExpressionAttributeNames: { "#s": "state" },
  ExpressionAttributeValues: { ":d": { S: "done" } },
}));

How the two states behave

A conditional put is atomic, so two simultaneous deliveries cannot both win. Enable DynamoDB TTL on expires_at if you want rows to age out, and pick a window longer than the time you may wait before pressing Redeliver. Sume's docs do not state a Redeliver horizon, so the 30 days in the sample is your choice, not a Sume number.

Item state and what a new delivery does (Sume docs, read 2026-10-04)
ItemDelivery arrivesAction
absentFirst deliveryclaim() returns true, do the work, call finish()
claimedRetry while work is in progressclaim() returns false, return 2xx
claimed, staleWorker crashed after claimingA sweeper re-runs the work from the stored body
doneRedeliver or late retryclaim() returns false, return 2xx

Caveats

  • Reserve state in expressions with a placeholder, as the sample does. It is a DynamoDB reserved word.
  • Return 2xx only after the put succeeds. If DynamoDB throttles you, return a 5xx so Sume retries.
  • Sume gives each attempt 10 seconds. Claim and respond, then do the slow work, such as downloading artifacts, outside the request.

Plan for the claim that never finishes

The case to design for is the claimed-but-never-finished item. Your receiver claimed the job, returned 2xx, and the worker died. Sume will not send that event again, because it saw a success. A small sweeper that scans for items still in claimed after a few minutes, and re-runs the work from the stored body, closes that hole.

The other recovery path is the API. A job that never produced a row is a job whose webhook never arrived or never verified, and GET /v1/jobs/{job_id}/status tells you the truth. The jobs docs say to keep those polls available, because a delivery that fails ten times leaves a job that still reached its real terminal state.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume