xAI batch mixes chat, image and video in one file; Sume uses queues
xAI's Batch API now takes chat, image and video requests in one JSONL file. A Sume bulk queue targets exactly one Format. How to split a holiday job.

Can one batch file carry text, image and video work? On xAI, yes. The xAI release notes (read 2026-10-04) say the Batch API now supports image generation, image editing and video generation in addition to chat completions, and that uploading JSONL through the Files API supports all batch endpoints in one file. A Sume bulk queue is narrower on purpose: POST /v1/formats/{handle}/{slug}/bulk-runs is addressed to a single Format, and every item is a run of that Format.
So a holiday job that mixes copy, stills and clips becomes one xAI file or several Sume queues. The split is the design decision, and it is easier to make early.
One file versus one queue per recipe
A mixed JSONL file is convenient when the work is a flat list of independent calls. A Sume Format is a saved recipe, so a queue naturally represents one repeated job: 100 product spots from the same Format, each with its own input. The three kinds of work in a holiday campaign map to three queues, each with its own concurrency between 1 and 16.
The benefit is operational. Each queue has its own frq_ id, counts and finish time, so a failure in the still-image Format cannot hide inside a queue of video runs, and you can read counts.failed per recipe.
Failure handling differs with the shape
With a mixed file, one batch reports across very different request kinds, and you sort failures out per line. With queues, each Format is its own receipt, so a recipe that is misconfigured fails on its own queue and does not obscure the others. A bad item fails the create call with details.index before a queue exists, which means you find format-wide mistakes in the first request, not after the render.
Spend is the other split. Each item in a Sume queue can carry its own generation_spend_cap_usd, up to 500, so the still-image Format and the video Format can have different ceilings. The docs list no queue-level cap, which makes the per-item value the control you have.
Mapping a mixed list to queues
| Question | xAI Batch API | Sume bulk run |
|---|---|---|
| Request kinds per submission | Chat, image generation, image editing and video generation, in one JSONL file | One Format per queue |
| Submission unit | JSONL uploaded via the Files API | JSON body with 1 to 100 items |
| Progress | Batch-level, per the vendor | One queue with counts and per-item status |
| Mixed job | One file | One queue per Format |
Group rows by recipe
This snippet groups a mixed to-do list by Format address and chunks each group at 100, which is what you would POST. It runs as is.
from collections import defaultdict
todo = (
[("mybrand/holiday-spot", {"input": {"sku": f"s{i}"}}) for i in range(130)]
+ [("sume/sume-slideshow", {"input": {"sku": f"p{i}"}}) for i in range(20)]
)
groups = defaultdict(list)
for fmt, item in todo:
groups[fmt].append(item)
requests = []
for fmt, items in groups.items():
for i in range(0, len(items), 100):
requests.append((fmt, {"concurrency": 8, "items": items[i : i + 100]}))
for fmt, body in requests:
print(fmt, len(body["items"]))
assert sum(len(b["items"]) for _, b in requests) == len(todo)Keep one key per queue
Idempotency keys are scoped to one Format, so the same key text used on two different Formats does not collide, but it is clearer to put the Format slug in the key. Store each queue id with its Format and chunk number, because there is no list-queues endpoint to rebuild the map later.
Sources
Related posts
More in Comparisons
- xAI batch video URLs expire in 1 hour: size the download worker
xAI's release notes say image and video URLs in batch results expire after 1 hour. How many parallel downloads you need, and how Sume's durable URLs differ.
- xAI file URLs auto-expire; Sume fetches image_url once at create
xAI's Files API can serve public URLs that auto-expire. Sume fetches an attachment image_url at run creation and copies it. What that means for hosting.
- Sume vs Argil: AI avatar video and video agents compared
Argil makes AI-avatar and story videos with a chat agent, Director; Sume is a video agent with a multi-model API. Avatars, API, pricing, and limits compared.
- Sume vs fal: a generative media API or a video agent platform
fal runs 1,000+ image, video, and audio models behind one API. Sume adds a video agent, Formats, and avatars to a multi-model API. How the two surfaces differ.
Written by Sume