AI SDK 7 tool search with deferred tools and Sume tools
AI SDK 7.0.127 lets a search() callback rank eligible deferred tools. Load a few Sume tools up front and fetch the rest by tools_schema on demand.

AI SDK 7 can defer tools and rank them with a search() callback, so a Sume-connected agent does not have to carry the whole hosted inventory in every prompt. Give it the handful it uses on most turns, defer the rest, and keep the paid tools gated by idempotency_key.
The vercel/ai releases page (read 2026-10-02) lists ai@7.0.127 with "select and rank eligible deferred tools via search() callback in tool search" and a configurable maxResults for tool search results. The page is a release list, not a guide, so I did not confirm the exact option shape, and the code below avoids guessing it.
Which Sume tools are worth loading every turn?
Sume's hosted MCP exposes discovery, jobs, assets and paid generation tools (MCP tools and gates). A reasonable always-on set is the cheap read side, and the rest can be deferred.
| Load up front | Defer until searched |
|---|---|
| tools_list, tools_schema, mcp_health | music_create, stt_create |
| jobs_status, jobs_wait, jobs_result | image_upscale_create, video_upscale_create, rmbg_create |
| balance_get, usage_get | assets_upload_url, assets_complete |
| generate_image, generate_video | script_run |
How does an agent learn a deferred tool's contract?
Sume's docs say to call tools_schema with a name to fetch one tool contract, and to never assume HTTP API parity. So a deferred Sume tool is a two-step discovery: the search ranks it, then the agent reads its schema before the first call.
Tell the model to read idempotency_key and dry_run first. Sume says idempotency_key is required on write and paid tools and is "for transport/dedup, not human approval," so approval still has to come from your app, as in Vercel AI SDK tool approval for a paid video.
What breaks if the right tool is not ranked?
The agent either picks a lesser tool or answers without one. For paid work that is safer than a wrong call, but it costs a turn. Tune maxResults upward for a Sume tool family that is small, and write tool descriptions that name the job ("generate a video") rather than the id.
Visibility also moves with auth. Under OAuth mcp:read only, mutating and paid tools are hidden, so a search over that session will never surface generate_video. If the agent says Sume has no video tool, check scopes before the ranking.
How do I test that deferral works?
Write five prompts that each need a different Sume tool: one for balance, one for a status check, one for an image, one for a video, one for speech. Run each against an agent that has only the always-on set loaded, and log which tool the search surfaced. Any prompt where the right tool is missing from the top results tells you what to rename or re-describe.
Check the paid prompts for gates, too. Every paid call should carry an idempotency_key, ideally a dry_run=true preview first, and your own approval when money is involved. A deferred tool is not a looser tool; it passes through the same gates as one loaded up front.
Keep the always-on set small. The point of deferral is a lean prompt, so measure token counts before and after, and add a tool to the always-on set only when searches for it keep costing a turn.
What does Sume add that the SDK does not?
A spend cap and a preview. max_spend_usd is enforced only when provided, and dry_run=true returns an admission and cost preview without a job. Both are arguments on the Sume tool; the SDK's search does not set them for you. After a submit, wait with jobs_wait in slices of at most 55 seconds and retry with the same ids on wait_slice_expired. Before you ship, read the live contract for every route you call at https://api.sume.com/reference/json, which Sume's docs name as the schema source of truth, and re-read the linked docs pages: limits, scopes and error codes change faster than blog posts do. Treat any number in this post as a snapshot dated 2026-10-02, and prefer the effective fields your own responses return, such as generation_limits, over a static table.
Sources
Related posts
More in Developers
- Did the edit stay in its region? A pixel-diff check for GPT Image 2.5
After a GPT Image 2.5 edit, measure how much changed outside the area you meant to change. A Pillow script that diffs the result against the original.
- Verify a Sume avatar video webhook in Python (HMAC SHA-256)
A Python verifier for avatar video webhooks: timestamp tolerance, rotation-safe comparison, and a hard refusal when the signing secret is empty.
- low_confidence_long_video: why video_frames warns past 90 seconds
Sume's video_frames returns the low_confidence_long_video warning when the source runs over 90 s. The job still succeeds; the hard cap is 300 s. What to do.
- Check has_audio first: video_inspect frames false before STT or detach
A free probe-only video_inspect tells you probe.has_audio before you reserve STT or run audio detach, so silent clips never hit the no-audio errors.
Written by Sume