Claude rejects forced tool_choice: steer generate_video by description
Claude Sonnet 5.5 and Opus 5.5 return a 400 for tool_choice any or tool, and thinking cannot be disabled. Steer Sume tool calls with descriptions and a dry run.

You cannot make Claude Sonnet 5.5 or Opus 5.5 call a tool by forcing it. The Claude API release notes say tool_choice set to any or tool returns a 400 on both (Sonnet 5.5 from Sep 28, Opus 5.5 from Sep 22), and that thinking cannot be disabled. To get generate_video called, guide the model with the tool description and your prompt.
Opus 5.5 is listed at $4 and $20 per million tokens for input and output.
What changed
| Model | Date | Tool choice any or tool | Thinking |
|---|---|---|---|
| Opus 5.5 | Sep 22 | Returns 400 | Cannot be disabled |
| Sonnet 5.5 | Sep 28 | Returns 400 | Cannot be disabled |
Why a forced call was tempting
A forced call guarantees a tool result, which is convenient for a pipeline that wants exactly one render. With a paid tool, forcing was also risky: the model had no chance to pause, check the cost or ask a question.
The replacement is to make the right call the obvious one, and to put the safety in the arguments.
Steering with the tool contract
Sume's hosted MCP lets the agent read a tool's contract before using it. tools_schema with a name returns one tool contract, and the docs give an example instruction: call tools_schema for generate_image and explain idempotency_key and dry_run before submitting any paid generation.
- Say in the system prompt when to use the tool: for example, when the user asks for a clip.
- State which fields are mandatory:
idempotency_keyon every paid call. - Tell it to omit
payload.modelunless the user named a family; the router then usessume/auto. - Ask for
dry_run=truefirst on anything expensive, then a real call.
Checking that it called
Without forcing, your code has to verify the outcome. In a Messages API loop, look for a tool use block, and treat a text-only reply as a case to handle: re-prompt, or fall back to a plain REST submit.
If three or more independent calls of the same shape are needed, the docs point to script_run, which runs a short script on Sume's side that calls the tools in a loop or in parallel and returns one value. It is bounded by timeout_seconds, max_calls and max_paid_calls.
After the call
A paid call returns a job, not a finished video. Wait with jobs_wait in slices of at most 55 seconds, then read jobs_result. If a slice expires, call jobs_wait again with the same ids. Resubmitting would create and bill a second job.
Sources
Related posts
More in Developers
- Draft with GPT Image 2.5 Flare, finish with Sunburst: a two-pass edit
OpenAI pairs Flare with fast generation and Sunburst with editing precision. A Python two-pass on Sume's Image API that drafts, then refines the first result.
- Mcp-Name header rules: rate-limit paid render tools at the gateway
MCP 2026-07-28 requires Mcp-Method and Mcp-Name headers on Streamable HTTP POSTs. A gateway can rate-limit paid render tools by name without reading the body.
- GPT Image 2.5 on ElevenLabs: 14 ratios plus auto. Sume lists 17
ElevenLabs offers 14 fixed ratios plus auto for GPT Image 2.5. Sume's normalized list has 17 plus auto. Read what each model accepts before sending one.
- ElevenLabs TTS output_format: 192 kbps needs Creator, PCM needs Pro
The ElevenLabs text to speech reference ties 192 kbps MP3 to Creator and PCM or WAV to Pro. A format table, a fallback chooser in Python, and the seed range.
Written by Sume