fal's hint pins a runner; Sume has no runner choice
fal's queue submit takes a hint that routes to the runner used before. Sume hides providers and runners. What to do when you need affinity.

fal's queue submit() accepts a hint string that, per fal's docs, tries to route a request to the same runner that handled an earlier request with the same hint. Sume has no counterpart: its public API is provider-neutral and does not let you choose, see or pin the machine that serves a job.
That is mostly irrelevant for plain media generation and matters only if you depend on runner-local state. This page explains the difference and what to do instead.
What does fal's hint do?
The fal queue page lists the submit() parameters: priority (normal or low), start_timeout, webhook_url, path, headers, and hint, which fal describes as a routing hint sent as the X-Fal-Runner-Hint header. When you pass one, fal tries to route the request to the same runner that handled a previous request with the same hint, which fal says is useful for keeping requests on a runner that already has a model or adapter loaded in memory. It is a best-effort affinity control, not a guarantee.
The same page documents automatic retries: failed requests (503, 504, connection errors) are retried up to 10 times and re-queued, and you can turn that off with the header X-Fal-No-Retry: 1.
What does Sume expose instead?
Sume's docs say public API responses are provider-neutral: they do not expose hidden provider names, raw provider task ids, raw provider URLs, internal workflow names or storage object keys. With model: "sume/auto", the poll response reports sume/auto and Sume does not disclose which family served the request, and the docs warn you not to infer it from the output.
What you control on Sume is the model (a catalog id), the request shape, and the idempotency key. A replay with the same key returns the original job, and the Auto resolution is a pure function of the normalized request and the catalog version, so a replay prices and routes identically.
| Control | fal queue | Sume public API |
|---|---|---|
| Runner affinity | hint string | None |
| Queue priority | priority: normal or low | None documented |
| Queue deadline | start_timeout (504 if not started) | None documented |
| Retry handling | Auto retries up to 10 times; X-Fal-No-Retry: 1 disables | You retry; Idempotency-Key makes replays safe |
| Provider or runner visible | Not described on this page | Hidden by design |
When would a missing hint matter?
If your workload is "send a prompt, get a file", it does not: each Sume job is self-contained and carries all its inputs as public HTTPS URLs. It would matter if you wanted successive requests to share something held on one machine. Sume's docs do not describe any such state, and Formats or Agent threads are the way Sume keeps context across steps, not runner affinity.
For multi-step work that needs continuity, Sume offers Format runs and Agent completions, where context lives in the run or thread, not on a runner. See the Jobs and results page for how a job can be recovered by id after your process restarts.
How do I get repeatable behavior on Sume?
Pin a catalog model instead of sume/auto, keep the same Idempotency-Key for exact retries, and read the capabilities from the catalog. Pinning gives you the same family every time; it does not pin hardware.
curl -X POST https://api.sume.com/v1/videos \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: scene-12-take-1" \
-d '{"model":"seedance-2","prompt":"Slow push-in on a ceramic mug","duration":6}'
# same key, same body: the original job comes back, no second charge.Which should you choose?
If you run your own serverless apps on fal and rely on routing a series of calls to one warm runner, hint is a real control that Sume cannot replace. If you consume hosted models and want the platform to own placement, Sume's hidden-provider design is simpler and gives you a stable contract. Compare the neighbouring retry behaviours in fal's no-retry header vs Sume's idempotency key and the overview in Sume vs fal.
Sources
Related posts
More in Comparisons
- fal queue has no size limit: what Sume does instead (queue_full)
fal's queue docs say there is no queue size limit and requests are never dropped. Sume caps accepted jobs per plan and returns 429 queue_full. What that means.
- fal low priority and start_timeout vs Sume queue-first admission
fal queues accept a low priority and a start_timeout that 504s. Sume has neither: it queues by plan and rejects with queue_full. What to do when work must wait.
- fal queue_position and logs vs Sume job status, events, usage
fal returns queue_position, runner logs and inference_time. Sume gives status, an events timeline and per-job usage. Which one tells you why a job is slow?
- fal webhook redirect 3xx is never retried: what Sume does instead
fal treats a 3xx from your webhook URL as a permanent failure. Sume does not follow redirects either, but counts a 3xx as a failed attempt, not an end.
Written by Sume