Together AI video API alternative: Sume /v1/videos compared

Together AI's video API is create-then-retrieve with five statuses. How that maps to Sume /v1/videos, and what each side does better.

5 min readSume
All posts

Is Sume a Together AI video API alternative?

Partly. Both give you an asynchronous video endpoint: you submit a prompt, get a job id, and poll until a video is ready. Together AI is an inference platform with a serverless catalog and dedicated model inference, so you get low-level knobs. Sume exposes a managed catalog behind POST /v1/videos, with Sume-hosted output files and a USD balance that is reserved on submit.

If you need to tune sampling per request, Together is the better fit. If you want one job lifecycle across video, image, avatar and music, with idempotent retries, Sume is closer to what you want. This page compares the two as of 2026-10-02.

How does the Together AI video flow work?

Together's video overview describes an asynchronous job model: you start generation with client.videos.create(), receive a job id, then call client.videos.retrieve(job.id) to check on it. The page lists five job statuses: queued, in_progress, completed, failed and cancelled.

The documented parameters include prompt, model, width, height, guidance_scale, negative_prompt, steps, fps and seconds. The page notes that negative_prompt is not supported by every model, and it points to the serverless and dedicated catalogs for per-model duration, resolution, FPS and keyframe specs. The examples use model ids such as minimax/video-01-director and minimax/hailuo-02.

How does the Sume flow differ?

Sume's /v1/videos route follows the OpenRouter video generation API, with a short list of documented differences. You submit with a catalog model and a prompt, and the response carries a job id and a polling_url. You poll GET /v1/videos/{jobId} until the status is completed, then download from the content URL. The same job is also visible through GET /v1/jobs/{id}/status and GET /v1/jobs/{id}/result.

On /v1/videos, Sume follows the OpenRouter vocabulary: pending, in_progress, completed, failed and cancelled. Only the first name differs from Together. The generic job endpoints use a different set (queued, processing, completed, failed, canceled), so pick one surface and map statuses once. Sume documents Idempotency-Key on this route, so a retry returns the original job instead of billing twice.

Status vocabulary side by side (read 2026-10-02)
MeaningTogether AISume
Waitingqueuedpending
Runningin_progressin_progress
Donecompletedcompleted
Errorfailedfailed
Stoppedcancelledcancelled

What does Together do better?

Together exposes the sampler. steps, guidance_scale, fps and explicit width and height are in its documented parameter list, and dedicated inference lets you run a model on capacity you control. Sume goes the other way on purpose: every model in its catalog reports supported_sizes: null, so a size field returns 400 unsupported_parameter, and no v1 model accepts seed. You pick resolution, aspect_ratio and duration from what GET /v1/videos/models lists.

Sume also rejects, rather than silently drops, a parameter the chosen model does not list. That is safer for catching mistakes, and less flexible if you want to pass provider-specific options: a non-empty provider.options returns 400 unsupported_parameter.

How do cost and concurrency compare?

Together's page shows an example response with a cost of 0.28 for one generated video, but the excerpt does not give a full price list, so check its catalog for rates. On Sume, the workspace USD balance is reserved on submit at provider list price times 1.25, and usage.cost is the billable amount. That markup is the price of the managed layer; if raw per-clip cost is your only metric, compare it with your own numbers.

Sume concurrency is plan-based: Free 1, Pro 4, Startup 8, Scale 20 jobs processing at once. Extra valid jobs wait as queued up to a queue capacity of max(3, concurrency x 5), then fail with 429 queue_full. Top-ups do not raise concurrency.

curl -X POST https://api.sume.com/v1/videos \
  -H "Authorization: Bearer $SUME_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: together-port-001" \
  -d '{"model":"seedance-2","prompt":"A paper boat drifts down a rain gutter, macro lens"}'

What should you check before migrating a Together client?

Most of the work is in the error and retry paths rather than the happy path. A Together client that polls retrieve until completed will port with a changed base URL and a status-name map, but a few behaviours differ and are easy to miss in a quick test.

Sume job webhooks fire on terminal events only (job.completed, job.failed, job.canceled), with no progress callbacks, so do not build a progress bar on them. Keep polling as a backup, and use exponential backoff. If a local process times out, do not resubmit the original paid request: look the job up by id, or replay the same Idempotency-Key.

  • Map Together's queued to Sume's pending on /v1/videos; the other four names match. If you also read /v1/jobs/{id}/status, that endpoint says queued, processing and canceled.
  • Remove steps, guidance_scale and size from request bodies; use resolution, aspect_ratio and duration from the model's catalog entry.
  • Call GET /v1/videos/models at startup to read per-model durations and resolutions instead of hard-coding them.
  • Send an Idempotency-Key on every create so a network retry cannot bill twice. Reusing a key for a different payload returns 409 idempotency_conflict.
  • Download finished videos from the Sume content URL; treat artifact URLs as opaque.

Which should you pick?

Pick Together AI when you need per-request sampling control or dedicated capacity. Pick Sume when you want a managed catalog, a stable job envelope, sume/auto routing, and the same lifecycle for the rest of your media calls. Migration is mostly one status rename and dropping steps, guidance_scale and size. See the Sume video guide and job lifecycle.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume