elevenlabs say CLI: text to speech from a terminal, and Sume's route
ElevenLabs' CLI got a say command (Eleven v3 default) on 2026-09-07. Sume's CLI has no audio generation commands yet; the TTS job is a plain HTTP call.

Sume's CLI cannot say a sentence aloud: its docs state that image, video and music generators are not CLI subcommands yet, and the command reference lists no text-to-speech command. A changelog roundup reports that ElevenLabs added a CLI elevenlabs say command on 2026-09-07 with Eleven v3 as its default. To get speech from a terminal on Sume today, call the TTS route over HTTP and read the job result.
This page shows what the CLI does cover and where speech fits.
What does the Sume CLI cover?
The command reference lists account and auth commands, read commands such as sume jobs get, sume jobs result and sume assets list, and sume tools list and sume tools schema for discovering schemas at runtime. So the CLI is useful after a TTS job exists: you can check its status and read its result from the terminal.
| Task | Sume CLI | Where to do it |
|---|---|---|
| Speak a sentence | Not a CLI command | POST /v1/tts-1.0/generate |
| Check a job | sume jobs status <job_id> | CLI |
| Read the audio result | sume jobs result <job_id> | CLI |
| List generated files | sume assets list | CLI |
| ElevenLabs say (reported) | Not applicable | ElevenLabs CLI, default Eleven v3 |
How do I speak a line from the shell?
Submit the request with curl, using an Idempotency-Key, then follow the job with sume jobs status and fetch it with sume jobs result. The result lists an audio artifact with a hosted URL that you can download. The default output is MP3 at 44.1 kHz and 128 kbps; ask for wav if you plan to feed the file to an audio tool.
Is there a shorter path from an agent?
Sume's hosted MCP server has a tts_create tool, and the CLI docs point to sume tools list to see what a key can call. For a script or a build step, plain HTTP is the least code. See Jobs and results for the wait modes.
What should I do?
If you only want a quick listen, use the vendor's own CLI or a web playground. If speech is part of a build, wrap the HTTP call in a small shell function and keep the job id with the file. Check the CLI docs again later: the surface changes as new commands ship.
Sources
Related posts
More in Developers
- ElevenLabs v4 10,000-character limit vs Sume TTS 20,000 transcript
Higgsfield's changelog lists 10,000 characters per ElevenLabs v4 generation; Sume TTS 1.0 takes 20,000. How to split a long script.
- Export video model prices weekly: a catalog-to-CSV script
Prices and limits for AI video models change. A short script reads Sume's video catalog into a CSV so you can date each row and track changes.
- Facebook Reels API: start, upload, finish and the video_state values
Publishing a Facebook Reel takes three calls: upload_phase start, a file upload to rupload, then finish with video_state PUBLISHED, SCHEDULED or DRAFT.
- Fan out 40 image variants without hitting queue_full
Sume accepts jobs as queued up to your plan's capacity, then returns 429 queue_full. Here are the plan numbers and a wave loop that reads generation_limits.
Written by Sume