Parallel tool calls on a voice agent: one idempotency key each
ElevenLabs agents default enable_parallel_tool_calls to true. If tools start paid jobs, give each call its own key and wait on all job ids in one batch.

The ElevenLabs changelog for Sept 21 lists an agent setting, enable_parallel_tool_calls, with a default of true. If one of your tools starts a paid generation job, a single user turn can now start several at once, so every call needs its own idempotency key and your follow-up should wait on all the job ids together.
Below: what the setting changes for a tool that submits Sume jobs, a runnable submit-and-poll script, and the rule for keys.
What changes when calls run in parallel
With sequential calls, an agent finishes one tool before it starts the next, so a duplicate submit is rare. With parallel calls the model can emit several tool calls in one step, and a retry of any one of them can arrive while the others are still in flight. For free reads that is fine. For a tool that spends money it is the case idempotency keys exist for.
Sume's jobs docs are explicit about the retry rule: after a timeout, retrying the submit is fine if you reuse the same Idempotency-Key, and you must not submit a new paid job for the same intent.
| Item | Fact | Source |
|---|---|---|
enable_parallel_tool_calls | Default true | ElevenLabs changelog, Sept 21 |
| Submit retry | Reuse the same Idempotency-Key | Sume jobs docs |
| Wait on many jobs (MCP) | jobs_wait takes 1 to 20 job_ids | Sume jobs docs |
| Wait slice | Default 50 s, capped at 55 s per call | Sume jobs docs |
Submit three images in parallel, then poll
This script submits three prompts at once with mode: "async", derives each idempotency key from a run id plus the prompt's position, and polls each job to a terminal state. The key is stable across retries of the same tool call and different across calls. It uses the POST /v1/images route and the job status endpoint from the docs.
import os, time
from concurrent.futures import ThreadPoolExecutor
import requests
BASE = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
RUN_ID = "run-2026-10-04-a"
PROMPTS = ["a red kettle on a white table", "a blue kettle on a white table", "a green kettle on a white table"]
def submit(index, prompt):
headers = {**HEAD, "Idempotency-Key": f"{RUN_ID}-image-{index}"}
body = {"model": "sume/auto", "prompt": prompt, "mode": "async"}
r = requests.post(BASE + "/v1/images", headers=headers, json=body, timeout=60)
r.raise_for_status()
return r.json()["data"]["job"]["id"]
def wait(job_id):
while True:
s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=HEAD, timeout=60).json()
data = s.get("data", s)
if data.get("terminal"):
return job_id, data.get("sume_status")
time.sleep(data.get("next_poll_after_seconds") or 5)
with ThreadPoolExecutor(max_workers=3) as pool:
ids = list(pool.map(lambda a: submit(*a), enumerate(PROMPTS)))
for job_id, status in pool.map(wait, ids):
print(job_id, status)Rules to put in the tool description
Write these into the tool's own instructions, not only into code review notes.
- Each call builds its key from the conversation or run id plus a stable per-call index, never from a timestamp or a random value generated on retry.
- A tool returns the job id immediately. It does not hold the turn open for a render.
- After a fan-out, the agent makes one batch wait over every job id rather than one wait per id. Over MCP that is
jobs_waitwithjob_ids, and onwait_slice_expiredit repeats the same call instead of resubmitting. - Sume's docs say to reuse a key only for the same operation and payload, so keep the index in the key.
When to turn the setting off
If your tools depend on each other, for example a voice must be chosen before a line is generated, the setting gives you nothing and adds ordering risk. The changelog entry gives the default and nothing about when to change it, so decide per agent: leave it on for independent fan-outs and off where one result feeds the next call.
Sources
Related posts
More in Agents
- OpenAI Dots x Runway: brief a dot, it plans shots. The Sume loop
Runway's Oct 1 changelog lists OpenAI Dots x Runway for all plans: brief a dot, it plans shots, Runway shoots. The same loop with Sume MCP tools.
- Run the Sume video agent from your backend with Agent Completions
POST /v1/agent/completions runs the same agent as the Sume Agents chat, with tools and media generation, and returns an async run receipt you poll or webhook.
- Safe automation for AI agents that call paid APIs
Keep agents read-only by default, keep secrets out of logs, and on hosted MCP send an idempotency_key, preview with dry_run, and cap with max_spend_usd.
- Scheduled AI video agent runs: cron, API triggers, and receipts
A Sume schedule is a saved Agents automation that runs on a cron cadence and returns a run receipt. Author it in the dashboard; start and monitor runs by API.
Written by Sume