Let a browser poll your backend, not the Sume API: a proxy pattern
Keep SUME_API_KEY on the server: submit async, return the job id, and give the browser a read-only status route. A 27-line TypeScript route with the checks.

Do not call api.sume.com from the browser. Put two routes on your own backend: a POST that submits with mode: "async" and returns the job id, and a GET that reads /v1/jobs/{id} and returns only the status and the result. The key stays in a server environment variable, and the browser only ever sees your own ids.
The authentication docs describe the API key as a workspace credential, so anything that holds it can spend that workspace's credits. That is the whole reason for the proxy.
The two routes
This handler uses the Fetch Request and Response types, so it fits a Next.js route handler or any edge runtime. It typechecks with strict on. The orderId becomes the Idempotency-Key, so a double click or a client retry cannot create two jobs.
declare const process: { env: Record<string, string | undefined> };
const BASE = "https://api.sume.com";
const auth = () => ({ "x-api-key": process.env.SUME_API_KEY!, "Content-Type": "application/json" });
export async function POST(req: Request) {
const { prompt, orderId } = await req.json();
if (typeof prompt !== "string" || !prompt.trim() || typeof orderId !== "string")
return Response.json({ error: "bad_input" }, { status: 400 });
const r = await fetch(`${BASE}/v1/image-1.0/generate`, {
method: "POST",
headers: { ...auth(), "Idempotency-Key": `img-${orderId}` },
body: JSON.stringify({ prompt, mode: "async" }),
});
const body = await r.json();
if (!r.ok) return Response.json({ error: body.error?.code ?? "upstream" }, { status: r.status });
return Response.json({ jobId: body.data.request_id }, { status: 202 });
}
export async function GET(req: Request) {
const id = new URL(req.url).searchParams.get("jobId");
if (!id || !/^[\w-]+$/.test(id)) return Response.json({ error: "bad_id" }, { status: 400 });
const r = await fetch(`${BASE}/v1/jobs/${id}`, { headers: auth() });
const body = await r.json();
if (!r.ok) return Response.json({ error: body.error?.code ?? "upstream" }, { status: r.status });
const job = body.data.job;
return Response.json({ status: job.status, result: job.status === "completed" ? job.result : null });
}What the code does on purpose
- Validates input before spending anything, and returns a 400 of your own.
- Passes only
error.codeback to the browser, never the upstream body, so a request id or detail never leaks. - Checks the job id with a strict pattern before putting it in a URL path.
- Returns the
resultonly when the status iscompleted, so a half-finished job cannot be misread. - Uses async mode, so the POST returns in a normal request time and the 30-second sync cap never applies.
What you still need to add
The sample trusts that the caller may read any job id. A real service must store the mapping from your user to the job id at submit time and check it on the GET, because the Sume key sees the whole workspace. Add rate limiting on your own routes too, since each browser poll becomes one Sume read.
On the browser side, poll every two to three seconds with a visible backoff, and stop on any terminal status. If you want fewer reads, switch to a webhook and push the completion to the browser over your own channel.
Sources
Related posts
More in Developers
- Lint a Format package in CI before you PUT it
A short Python check for a Sume Format package: SKILL.md name equals the slug, allowed folders and extensions, file names, 100 MiB limits. Run it before PUT.
- Lip-sync a 2-minute monologue: split one TTS wav into H3 Max clips
MiniMax H3 Max lip-sync takes audio of 5 to 14.8 seconds. For a longer speech, render one TTS wav and cut it at sentence ends; the cost is worked below.
- Is there a lipsync-1.0 endpoint on Sume? Old paths 404; use Fabric
Sume's old /v1/lipsync-1.0 paths return 404 and its model ids return model_not_found. Send the same still and audio to veed/fabric-1.0 or H3 Max lip-sync.
- List Sume's music and TTS router engines in Python before deploy
A short Python script reads the music and TTS router model lists so a deploy check can confirm your pinned engine ids and fixed prices still exist.
Written by Sume