GitLab CI retry reruns the whole script: pin the Sume key
GitLab's retry keyword re-runs the full job script, so a retried job that calls Sume can submit and pay twice. Retry only runner failures and pin the key.

Short answer
When a GitLab job is retried it re-runs the full script, not just the step that failed, per the GitLab CI/CD YAML reference. If that script submits a paid Sume generation and then fails while polling, the retry submits again and you pay twice. Restrict retry to runner failures, and give the submit an Idempotency-Key built from values that do not change between attempts.
What retry does
The retry keyword controls when and how many times a job is automatically retried after a failure, and it takes a when subkey for the failure types. The reference lists script_failure, runner_system_failure, stuck_or_timeout_failure, api_failure and always. The distinction matters for paid APIs. A runner_system_failure means the infrastructure broke, often before your script did anything. A script_failure means your own commands exited non-zero, which for a poll timeout can happen after Sume already accepted and started work.
| retry:when value | Meaning | Risk for a paid submit |
|---|---|---|
| runner_system_failure | The runner failed | Low; the script may not have run |
| script_failure | The script exited non-zero | High; the submit may already be accepted |
| Other when values | GitLab lists more in the retry reference | Check each one against the same question: could the submit have been accepted? |
Pin the key to the work, not the attempt
Sume lets a retried submit return the original job when it carries the same Idempotency-Key. On a Format run, the same key and body returns 200 with idempotency_hit true, and a different body under the same key is 409 idempotency_conflict. For generation routes, the jobs docs say that retrying the submit with the same key returns the original job instead of billing a second one.
So build the key from the intent: the commit SHA and the name of the asset, for example. Do not generate a random UUID inside the script, because each attempt would then be a new request. Keep the request body identical across attempts as well.
sume-smoke:
retry:
max: 2
when: runner_system_failure
script:
- |
curl -sS -X POST https://api.sume.com/v1/image-1.0/generate \
-H "x-api-key: $SUME_API_KEY" \
-H "Idempotency-Key: ci-$CI_COMMIT_SHA-hero" \
-H "Content-Type: application/json" \
-d '{"prompt":"A matte black bottle on marble","mode":"async"}'
Separate submit from wait
A long poll inside one CI job is the usual cause of script_failure. Split the work: submit with mode async in one step, store the job id as an artifact, and poll in a later step or job whose retry is safe because reads are idempotent. Poll the status route, honor next_poll_after_seconds, and stop when terminal is true. A job timeout in CI does not cancel the generation; the work in flight is still billed.
Do not retry validation errors
A 400, 401 or 402 means the request or the account is wrong, and another attempt changes nothing. Let the job fail fast and read error.code from the body. Keep the API key in a masked, protected CI variable, and read it from the environment as the sample does, never from a file in the repository.
A worked failure
Picture a pipeline that submits a video job, then polls for 15 minutes and hits the CI job timeout. The job fails, and if the job uses a plain retry with no when filter, GitLab retries the whole script and a naive script submits a second video. (A job that only retries on runner_system_failure would not retry here, which is the point of the filter.) With the key pinned to the commit and the asset name, the second submit returns the first job, and the poll simply continues. Without it you have two paid generations and one result you will use.
The same logic applies to scheduled pipelines. A nightly run that reuses the same key for the same body may keep returning the original job, so include the date in the key when you actually want a fresh generation each night.
Sources
Related posts
More in Developers
- Go 1.27 drains response bodies: a Sume job poll loop
Go 1.27 drains unread HTTP/1 body bytes on Close. A stdlib loop that polls GET /v1/jobs/:id/status and honors next_poll_after_seconds.
- Go http.Client Timeout covers the body: Sume job poll
Go's Client.Timeout includes redirects and reading the body and keeps running after Do returns. Use it per poll and put the job deadline on a context.
- Go webhook handler: verify the Sume sume-v1 signature
A stdlib Go verifier for x-sume-webhook-signature: HMAC-SHA256 over timestamp.raw_body, five-minute window, constant-time compare, empty secret refused.
- Google Images formats and filenames for Sume output
Google Search supports BMP, GIF, JPEG, PNG, WebP, SVG and AVIF in img src, and wants short filenames and real alt text. Rename Sume downloads first.
Written by Sume