GitLab CI retry reruns the whole script: pin the Sume key

GitLab's retry keyword re-runs the full job script, so a retried job that calls Sume can submit and pay twice. Retry only runner failures and pin the key.

5 min readSume
All posts

Short answer

When a GitLab job is retried it re-runs the full script, not just the step that failed, per the GitLab CI/CD YAML reference. If that script submits a paid Sume generation and then fails while polling, the retry submits again and you pay twice. Restrict retry to runner failures, and give the submit an Idempotency-Key built from values that do not change between attempts.

What retry does

The retry keyword controls when and how many times a job is automatically retried after a failure, and it takes a when subkey for the failure types. The reference lists script_failure, runner_system_failure, stuck_or_timeout_failure, api_failure and always. The distinction matters for paid APIs. A runner_system_failure means the infrastructure broke, often before your script did anything. A script_failure means your own commands exited non-zero, which for a poll timeout can happen after Sume already accepted and started work.

Retry triggers and what they mean for a Sume job (read 2026-10-03)
retry:when valueMeaningRisk for a paid submit
runner_system_failureThe runner failedLow; the script may not have run
script_failureThe script exited non-zeroHigh; the submit may already be accepted
Other when valuesGitLab lists more in the retry referenceCheck each one against the same question: could the submit have been accepted?

Pin the key to the work, not the attempt

Sume lets a retried submit return the original job when it carries the same Idempotency-Key. On a Format run, the same key and body returns 200 with idempotency_hit true, and a different body under the same key is 409 idempotency_conflict. For generation routes, the jobs docs say that retrying the submit with the same key returns the original job instead of billing a second one.

So build the key from the intent: the commit SHA and the name of the asset, for example. Do not generate a random UUID inside the script, because each attempt would then be a new request. Keep the request body identical across attempts as well.

sume-smoke:
  retry:
    max: 2
    when: runner_system_failure
  script:
    - |
      curl -sS -X POST https://api.sume.com/v1/image-1.0/generate \
        -H "x-api-key: $SUME_API_KEY" \
        -H "Idempotency-Key: ci-$CI_COMMIT_SHA-hero" \
        -H "Content-Type: application/json" \
        -d '{"prompt":"A matte black bottle on marble","mode":"async"}'

Separate submit from wait

A long poll inside one CI job is the usual cause of script_failure. Split the work: submit with mode async in one step, store the job id as an artifact, and poll in a later step or job whose retry is safe because reads are idempotent. Poll the status route, honor next_poll_after_seconds, and stop when terminal is true. A job timeout in CI does not cancel the generation; the work in flight is still billed.

Do not retry validation errors

A 400, 401 or 402 means the request or the account is wrong, and another attempt changes nothing. Let the job fail fast and read error.code from the body. Keep the API key in a masked, protected CI variable, and read it from the environment as the sample does, never from a file in the repository.

A worked failure

Picture a pipeline that submits a video job, then polls for 15 minutes and hits the CI job timeout. The job fails, and if the job uses a plain retry with no when filter, GitLab retries the whole script and a naive script submits a second video. (A job that only retries on runner_system_failure would not retry here, which is the point of the filter.) With the key pinned to the commit and the asset name, the second submit returns the first job, and the poll simply continues. Without it you have two paid generations and one result you will use.

The same logic applies to scheduled pipelines. A nightly run that reuses the same key for the same body may keep returning the original job, so include the date in the key when you actually want a fresh generation each night.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume