OpenAI background mode vs Sume Agent Completions: polling

Both let you start a long job and poll for it. OpenAI background mode returns a response id; Sume returns an agrun_ receipt and needs a required spend cap.

4 min readSume
All posts

OpenAI's background mode and Sume Agent Completions both start a long job, return an id, and let you poll until it is done. The differences are in what you poll for, how you cancel, whether you can stream, and what a spend limit looks like.

This is a comparison of two asynchronous patterns, not of the work they do: one returns model output, the other returns generated video and other media.

Neither design is better in the abstract. They serve different jobs, and the useful comparison is about what you must build around each.

How each one starts

On OpenAI's page, read 2026-10-10, you set background to true on a Responses API call. Responses move through queued and in_progress and then reach a terminal state such as completed, failed, or cancelled. You poll with GET requests by response id while the state is queued or in_progress.

On Sume, POST /v1/agent/completions returns 202 with an agrun_ receipt. You poll GET /v1/agent-runs/{id} or its status route, or you supply communication.webhook_url and let a signed webhook arrive. See Agent Completions.

Cancel, stream, retain

OpenAI documents cancel as idempotent, with a second call returning the final response object. Sume's cancel is also idempotent, and you pay for generation completed before the cancel.

OpenAI supports streaming a background response by setting stream to true and resuming with a sequence_number and starting_after. Sume has no streaming on Agent Completions, so you either poll or wait for the webhook. OpenAI also documents data retention caveats for zero data retention projects; Sume's docs make no equivalent statement that we can cite here, so check your own requirements before relying on either.

Background jobs compared, OpenAI page and Sume docs both read 2026-10-10
QuestionOpenAI background modeSume Agent Completions
How you start itbackground: true on a responsePOST /v1/agent/completions
What you get backA response that is queued or in_progress202 receipt with an agrun_ id
Terminal statescompleted, failed, cancelledcompleted, failed, canceled, skipped on run families
StreamingYes, with sequence_numberNo
Spend limit on the requestNot described on that pagegeneration_spend_cap_usd, required
Push notificationNot described on that pageSigned webhook

Which to reach for

If you want text or structured model output and can tolerate a slower first token, background mode on the Responses API fits. If you want a finished video, images, or audio produced by a video agent, and a hard dollar ceiling per request, Agent Completions is the Sume path.

You can also use both: a model-side step that plans or summarizes, and a Sume step that produces the media, joined by your own code and a run index.

For a product, the practical difference is cost control. A background response is bounded by tokens and time; an agent run is bounded by the generation cap you set, which is the number your finance team cares about.

What to build either way

Poll with a backoff and a deadline, store the id before anything else, handle the cancel path, and decide what happens on timeout. With Sume, add a webhook with signature verification and dedupe on request_id, and keep the cap as a number you control in code.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume