Haiku 5.5 effort for an agent that polls Sume jobs

Claude Haiku 5.5 defaults to medium effort. What that means for a loop that calls Sume jobs_wait, and when to move to low or high.

5 min readSume
All posts

Leave Claude Haiku 5.5 at its default of medium effort for an agent that submits a Sume job and polls it with jobs_wait, and move to low only for the poll step itself. Anthropic's effort page says medium is the default on the Claude API and in Claude Code, that setting the default is the same as omitting the parameter, and that low is the cheapest and fastest level for "short tool tasks". The same page warns that in long agent prompts the model is more likely to skip a search, stop early, or skip a check at low.

What the vendor page says about Haiku 5.5

Haiku 5.5 supports all five levels: low, medium, high, xhigh and max. Effort applies to every output token, including tool calls and their arguments, so a lower level means fewer and terser tool calls. Thinking is on by default and counts toward max_tokens, so leave room for it.

Haiku 5.5 effort guidance, Claude API docs, read 2026-10-08
LevelVendor guidance for Haiku 5.5Fit for a Sume job loop
lowChat, short tool tasks, simple high-volume requests; may skip a check in long promptsOne poll step that only reads a status
mediumDefault; start here, including agentic codingPlan, submit, poll, summarise
highKnowledge work, longer agent tasks, strict instruction followingRuns that must follow spend rules exactly
xhigh / maxOnly where your evals show a gainRarely worth it for a status poll

Where Sume's side of the loop is fixed

Effort changes how much the model reasons. It does not change what Sume's hosted MCP server enforces. A paid tool still needs an idempotency_key, max_spend_usd is enforced only when you send it, and one jobs_wait call holds for at most 55 seconds on the remote server. A cheaper model does not make a render faster, so the poll loop looks the same at every effort level.

That is the argument for low on the poll step. Waiting is a mechanical step: read status, decide whether to call jobs_wait again. The step that deserves more thinking is the one before it, where the model chooses a tool, builds the payload and decides what cap to send.

A split that keeps the expensive thinking where it matters

Run the planning turn at the default medium and send the submit call with dry_run set to true first, so the model reads an admission preview before spending. Then let a low effort subagent or a later turn carry the polling. If a run skips a dry_run or drops max_spend_usd at low, take that as the signal Anthropic describes and raise the level for that step instead of adding more prompt text.

  • Plan and submit: medium, with dry_run first.
  • Poll with jobs_wait slices of 45 to 55 seconds: low is enough.
  • Never resubmit the paid create call to check progress; poll the job id.
  • Raise to high if the model ignores your spend rules in a long prompt.

Check it on your own runs

Anthropic states that the impact of effort varies by task and tells you to evaluate on your own use cases. Count three things per run: paid calls without max_spend_usd, calls without dry_run before a first submit, and jobs_wait calls that resubmitted a create. Those are the failures that cost money, and they are visible in the tool-call log. Tool timeouts for long jobs are covered in jobs_wait for long video jobs.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume