AI/ML API vs Sume: OpenAI-style base URL or media jobs

AI/ML API gives an OpenAI-compatible base URL across text, image, video and music. Sume is media jobs, and its agent endpoint is async, not choices[].

5 min readSume
All posts

Does Sume replace AI/ML API?

Only for media. AI/ML API is a multi-modality gateway: its docs list text and chat, image, video, music, speech and voice, 3D, vision and embeddings, from developers including Anthropic, OpenAI, Google, Meta and Mistral AI. Sume's public API covers Image, Video, Avatar, Music and the jobs around them, plus Formats and agent runs. It has no chat-completions route, so if you want one key for LLM calls too, AI/ML API is the better fit.

What does the OpenAI-compatible claim mean?

AI/ML API emphasises OpenAI compatibility: its example builds an OpenAI client with base_url="https://api.aimlapi.com/v1". Text models are documented with completion and chat completion, streaming, code generation, reasoning, function calling, vision input and web search. The page I read did not detail the video flow or pricing, so confirm those in its model pages before you migrate.

Sume is compatible in a narrower place. Its /v1/videos route follows the OpenRouter video generation API with a short list of documented differences, so a client written against those docs needs little more than a new base URL and key. Its agent endpoint, POST /v1/agent/completions, is different from an OpenAI chat completion: it is async only, returns 202 with a receipt, and does not return choices[].

What each base URL is for (read 2026-10-02)
NeedAI/ML APISume
Chat completions and streamingYes, OpenAI-compatibleNo chat-completions route in the API reference
Video generationModels from several developers listedPOST /v1/videos, OpenRouter-shaped
Agent that plans and assembles mediaNot covered in the page I readPOST /v1/agent/completions and Formats, async receipts
Result deliveryCheck model pagesJob envelope, poll or signed webhook, media.sume.com files

What breaks when you point an OpenAI client at Sume?

Most things that assume a synchronous response. An OpenAI SDK call expects the answer in the same response. Sume video and agent calls return a receipt first, then you poll status_url until the job leaves queued or processing and read result_url. Terminal job statuses are completed, failed and canceled.

Sume also differs from OpenRouter in a few places you will hit quickly: bare model ids such as seedance-2 instead of org/slug, size returning 400 unsupported_parameter because every v1 model reports supported_sizes: null, no seed, and Idempotency-Key honoured on the route so a replay returns the original job. Billing is in workspace USD, reserved on submit.

How do errors and limits differ on the Sume side?

Sume returns one public error envelope with a request id inside error, plus ratelimit-limit, ratelimit-remaining, ratelimit-reset and retry-after headers on public responses. The two 429s mean different things. rate_limited is request volume, so back off using retry-after. queue_full means the workspace has no room for another paid job until one finishes or is canceled.

A 402 insufficient_credits arrives before provider work starts, because the balance is checked and reserved at submit. A gateway client that treats every 4xx as a bad request will mis-handle these, so branch on the code field.

  • Branch on error.code, not only on the HTTP status.
  • Never retry a paid submit without an Idempotency-Key.
  • Treat provider_capacity_exceeded as retry later with the same key.
  • Keep the request id from the error body when you contact support; do not log API keys or signed URLs.

When does each make sense?

Keep AI/ML API if your app mixes LLM calls with media and you want one base URL and one key. Use Sume for the media side when you need an async job contract, idempotent retries and finished files hosted on media.sume.com. A split is common: LLM calls on a gateway, video and avatar renders on Sume.

Check the catalog before you commit: GET /v1/videos/models lists resolutions, aspect ratios, durations and whether a model generates audio, so you can compare against what your current gateway serves.

curl "https://api.sume.com/v1/videos/models" \
  -H "Authorization: Bearer $SUME_API_KEY"

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume