Agent Completions model field is sume-agent only: no Haiku or GLM
You cannot choose Claude Haiku 5.5 or GLM 5.3 in the Agent Completions model field. Sume accepts sume-agent and returns 400 for anything else.

The model field on Sume's POST /v1/agent/completions accepts only sume-agent. If you leave it out, you get the same agent. Any other value, including a Claude or GLM id, returns 400 invalid_request. New models such as Claude Haiku 5.5 and GLM-5.3 do not change that: the endpoint runs Sume's own agent with its own model choices, so the model is not a parameter you tune.
What the docs say about the endpoint
These are the documented request rules for Sume Agent Completions, read on 2026-10-08.
| Field | Rule |
|---|---|
| model | Optional; only sume-agent |
| instruction or messages | Send exactly one |
| generation_spend_cap_usd | Required; no default |
| assistant turns in messages | Rejected |
| attachments | Up to 30 images |
| Idempotency-Key | Same key returns original receipt; changed payload returns 409 |
| Success | 202 with an agent.run receipt |
Why that is the design
Sume's docs describe an Agent Completion as running the Sume agent on an ad-hoc prompt, in the same runtime as the Agents chat, with a sandbox, tools, an MCP bridge and media generation. Because the agent decides which tools and models to use, a client-chosen model name would not mean what it means on a plain chat API. The response is also not a chat completion: it is an async run receipt, and streaming is not available.
If you need to control the model
Run your own loop with the model you want and use Sume's hosted MCP server for the media work. There you control effort, reasoning level, retries and prompt size, and Sume controls what its tools accept. On Claude, output_config.effort is the knob, and Haiku 5.5 defaults to medium. On GLM-5.3, reasoning is always on at low, high or max.
The in-app agent picker is separate. Sume's catalog lists Haiku 5.5 and GLM 5.3 Flash for chat, behind a gate, and that list does not feed this API field. If a request fails with the model message, delete the field.
Related: older keys lack the Agent Completions scopes; see 403 insufficient_scope on an older key.
A correct request
A minimal valid body has an instruction and a generation_spend_cap_usd. Leave model out or set it to sume-agent.
The receipt returns model: "sume-agent" in the agent.run object, so you can assert on it in your client. Poll the status_url until next_action is not poll_status, and read usage.generation_spend_cap_usd_micros to confirm the cap was recorded.
- Body:
{"instruction": "...", "generation_spend_cap_usd": 2}. - Header:
Authorization: Bearer $SUME_API_KEY, plus anIdempotency-Keyfor safe retries. - Scope:
agent_completions:writeon the key.
Sources
Related posts
More in Developers
- Assemble a three-clip montage with fades through the Sume API
A working Timeline 1.0 request that joins three hosted clips over one audio spine with fade and dissolve transitions, plus the Python to poll it, for $0.10.
- Balance needed to submit 10 or 50 video jobs: reserve per model
Sume reserves each job's estimate at submit. A table of the balance 10 and 50 ten-second clips need on six video models, and where a 402 lands in a batch.
- callback_url on /v1/videos: the Sume job envelope that arrives
A /v1/videos callback_url delivers Sume's job.completed, job.failed or job.canceled envelope, not video.generation.* events. Payload, signature, checks.
- Cancel queued Sume jobs after queue_full; handle 409 already started
How to free capacity after a 429 queue_full: cancel queued jobs, read job_generation_already_started on running ones, and why a cancel releases the reserve.
Written by Sume