Parallel tool calls on a voice agent: one idempotency key each

ElevenLabs agents default enable_parallel_tool_calls to true. If tools start paid jobs, give each call its own key and wait on all job ids in one batch.

4 min readSume
All posts

The ElevenLabs changelog for Sept 21 lists an agent setting, enable_parallel_tool_calls, with a default of true. If one of your tools starts a paid generation job, a single user turn can now start several at once, so every call needs its own idempotency key and your follow-up should wait on all the job ids together.

Below: what the setting changes for a tool that submits Sume jobs, a runnable submit-and-poll script, and the rule for keys.

What changes when calls run in parallel

With sequential calls, an agent finishes one tool before it starts the next, so a duplicate submit is rare. With parallel calls the model can emit several tool calls in one step, and a retry of any one of them can arrive while the others are still in flight. For free reads that is fine. For a tool that spends money it is the case idempotency keys exist for.

Sume's jobs docs are explicit about the retry rule: after a timeout, retrying the submit is fine if you reuse the same Idempotency-Key, and you must not submit a new paid job for the same intent.

Parallel tool call settings and Sume job rules (read 2026-10-03)
ItemFactSource
enable_parallel_tool_callsDefault trueElevenLabs changelog, Sept 21
Submit retryReuse the same Idempotency-KeySume jobs docs
Wait on many jobs (MCP)jobs_wait takes 1 to 20 job_idsSume jobs docs
Wait sliceDefault 50 s, capped at 55 s per callSume jobs docs

Submit three images in parallel, then poll

This script submits three prompts at once with mode: "async", derives each idempotency key from a run id plus the prompt's position, and polls each job to a terminal state. The key is stable across retries of the same tool call and different across calls. It uses the POST /v1/images route and the job status endpoint from the docs.

import os, time
from concurrent.futures import ThreadPoolExecutor
import requests

BASE = "https://api.sume.com"
HEAD = {"Authorization": "Bearer " + os.environ["SUME_API_KEY"]}
RUN_ID = "run-2026-10-04-a"
PROMPTS = ["a red kettle on a white table", "a blue kettle on a white table", "a green kettle on a white table"]

def submit(index, prompt):
    headers = {**HEAD, "Idempotency-Key": f"{RUN_ID}-image-{index}"}
    body = {"model": "sume/auto", "prompt": prompt, "mode": "async"}
    r = requests.post(BASE + "/v1/images", headers=headers, json=body, timeout=60)
    r.raise_for_status()
    return r.json()["data"]["job"]["id"]

def wait(job_id):
    while True:
        s = requests.get(f"{BASE}/v1/jobs/{job_id}/status", headers=HEAD, timeout=60).json()
        data = s.get("data", s)
        if data.get("terminal"):
            return job_id, data.get("sume_status")
        time.sleep(data.get("next_poll_after_seconds") or 5)

with ThreadPoolExecutor(max_workers=3) as pool:
    ids = list(pool.map(lambda a: submit(*a), enumerate(PROMPTS)))
    for job_id, status in pool.map(wait, ids):
        print(job_id, status)

Rules to put in the tool description

Write these into the tool's own instructions, not only into code review notes.

  • Each call builds its key from the conversation or run id plus a stable per-call index, never from a timestamp or a random value generated on retry.
  • A tool returns the job id immediately. It does not hold the turn open for a render.
  • After a fan-out, the agent makes one batch wait over every job id rather than one wait per id. Over MCP that is jobs_wait with job_ids, and on wait_slice_expired it repeats the same call instead of resubmitting.
  • Sume's docs say to reuse a key only for the same operation and payload, so keep the index in the key.

When to turn the setting off

If your tools depend on each other, for example a voice must be chosen before a line is generated, the setting gives you nothing and adds ordering risk. The changelog entry gives the default and nothing about when to change it, so decide per agent: leave it on for independent fan-outs and off where one result feeds the next call.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume