Mastra eager tool execution: dry-run Sume calls first

Mastra 1.71 can start a tool once its own arguments are complete. For paid Sume generation that means a stable idempotency_key and a dry run before any spend.

4 min readSume
All posts

Eager tool execution is safe for Sume's generation tools only if each call carries a stable idempotency_key and the first pass is a dry run. The Mastra releases page lists 1.71.0 with eager tool execution that starts a tool as soon as its own arguments are complete, which can put several image calls in flight while the model is still writing the rest of its answer.

That is a real speedup for a storyboard with a dozen stills, and a real way to overspend if nothing limits it.

What the release page lists

Relevant items only.

Mastra releases (read 2026-10-03)
ReleaseItem
1.71.0Eager tool execution once a tool's own arguments are complete
1.71.0Storage capability negotiation

What can go wrong with eager paid calls

A tool call that starts early may be one the agent would have revised, or one that is repeated when a step is retried. For a free read that is harmless. For a paid create it is a charge, because each accepted job reserves its estimated USD amount at submit time.

Sume's gates are built for this. On hosted MCP, write and paid tools require an idempotency_key, which Sume treats as a transport and dedup key rather than human approval. A replay with the same key returns the original job instead of billing a second one. dry_run=true returns an admission and cost preview without submitting, and max_spend_usd caps spend when you pass it.

Make each eager call idempotent

Derive the key from the content of the call, not from a counter. If two eager calls describe the same scene, they get the same key and collapse into one job; if the model changes the prompt, the key should change as well, since Sume returns 409 idempotency_conflict when a key is reused with a different payload.

import { createHash } from "node:crypto";

export function sumeKey(taskId: string, args: unknown): string {
  const h = createHash("sha256")
    .update(JSON.stringify(args))
    .digest("hex")
    .slice(0, 16);
  return `${taskId}-${h}`;
}

// sumeKey("deck-17-scene-4", { prompt, aspect_ratio })

A two-pass pattern

Treat the first pass as pricing and the second as spending.

  • Pass one: let eager execution run every call with dry_run=true. Nothing is submitted, so nothing is reserved.
  • Total the estimates and compare with the balance from balance_get. Ask the user to confirm if the sum is large.
  • Pass two: repeat the same calls without dry_run, with the same keys and, if you want a hard cap, max_spend_usd.
  • Follow all the resulting job ids with one batch wait rather than many single waits.

Fan-out within workspace limits

Concurrency on Sume is a dispatch limit, not a submit limit. A workspace on the Free plan processes 1 job at a time with queue capacity of 5; Pro is 4 and 20; Scale is 20 and 100. Extra valid jobs wait as queued, which is a normal state. When accepted capacity is full, a submit returns 429 queue_full, and the right response is to wait for running jobs or cancel queued ones, then retry with the same key.

Submit responses include a generation_limits snapshot. Read queue_capacity_remaining and stop adding eager work when it is low. The dashboard Concurrency tab, not the plan table, is the source of truth for your effective limit.

Collecting the results

jobs_wait accepts 1 to 20 job ids with wait_for set to all or any, holds at most 55 seconds per call (50 by default), and on wait_slice_expired you simply call it again with the same ids. Pass include_results: true and every completed job returns its result in the same answer. Never resubmit a paid create because a wait slice ended.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume