ElevenLabs use_pvc_as_ivc: what it does and Sume's voice selector
ElevenLabs' use_pvc_as_ivc flag swaps a professional voice clone for its instant version. Sume's TTS has no such switch: you pick an avatar or a voice id.

use_pvc_as_ivc is a boolean on ElevenLabs' text-to-speech requests, default false, that asks for the Instant Voice Clone version of a Professional Voice Clone. ElevenLabs says it "may improve expressiveness and reduce latency". Sume's text-to-speech API has no equivalent switch: you choose one voice and Sume synthesizes with it.
The flag came back into focus on September 28, 2026, when ElevenLabs' changelog said use_pvc_as_ivc is no longer deprecated on the convert, stream and timestamp variants of Text to Speech, and added it to Text to Dialogue as well. Both facts below come from ElevenLabs' Text to Speech convert reference and its September 28 changelog, read on 2026-10-02.
What does use_pvc_as_ivc do on ElevenLabs?
A Professional Voice Clone (PVC) is the heavier clone tier. The flag tells the API to use the instant-clone rendition of that same voice for this one request. The convert reference gives the trade as expressiveness and latency, not quality in general, and it does not publish a number.
Two practical points follow from the page. The default is false, so existing calls keep using the professional version. And because the changelog calls the field no longer deprecated, code that dropped it during the earlier deprecation can use it again.
Does Sume have a PVC or IVC choice?
No. Sume's TTS 1.0 request takes a transcript of up to 20,000 characters and one voice selector: a top-level avatar_id or avatar_handle, or a voice.id. The voice object has a single mode, id. A voice.id must be a TTS voice UUID or a Voices library id that starts with voi_; any other shape is refused with 400 invalid_voice_id before a job is queued or credits are reserved, per the live OpenAPI described on the API reference page.
There is also no engine picker on TTS 1.0. Sending model or model_id returns 400. To choose an engine you use the TTS Router, whose catalog is Cartesia Sonic ids (sonic-3.6, sonic-3.5, sonic-3, sonic-latest, sonic-preview). One of those catalog rows, sonic-preview, is a provider beta channel and rejects professional voice clones with voice_model_mismatch.
How do the two request shapes compare?
| Question | ElevenLabs | Sume |
|---|---|---|
| Pick the clone tier per request | use_pvc_as_ivc, default false | No field; one voice per selector |
| Voice selector | A voice id in the URL path | avatar_id, avatar_handle, or voice.id |
| Engine selector | model_id | None on TTS 1.0; model on /v1/tts-router/generate |
| Wrong voice id shape | Not stated on the page | 400 invalid_voice_id, nothing reserved |
What should you do if latency is why you want the flag?
Sume's TTS is a job, not a stream: you submit, then poll GET /v1/jobs/:id/status or take a webhook, and read the audio artifact from the result. That shape suits narration, ad reads and avatar audio; it is the wrong shape for a live phone agent where first-byte time matters. If you need the instant-clone latency trade for a conversation, ElevenLabs' flag is the right tool and Sume is not.
If your use is finished audio, compare engines on the Sume side instead. Run the same transcript on two Router ids, keep the voice fixed, and listen. The completed TTS job records model_id, voice, language, output_format, generation_config and speed, so you can reproduce the winning take, as described in Jobs and results.
What does Sume not do here?
If your pipeline depends on ElevenLabs voice tiers, keep that step on ElevenLabs and bring the finished audio to Sume only when you need Sume's media tools downstream.
- It does not let a request choose between a professional and an instant version of one voice.
- It does not stream audio during synthesis; TTS returns a finished artifact.
- It does not accept a voice name from another ecosystem in
voice.id. - It does not make any latency claim for the Sonic engines; measure on your own transcript.
Sources
Related posts
More in Comparisons
- ElevenLabs stability 0.5 and similarity 0.75 vs Sume generation_config
ElevenLabs defaults stability to 0.5 and similarity to 0.75. Sume has no such sliders: generation_config takes only volume, speed and an emotion string.
- fal Agent (early access) vs Sume Agent Completions and hosted MCP
fal's August 2026 agent is early access, with API, CLI and MCP. Sume's agent has Agent Completions over the API and a hosted MCP. What each documents today.
- fal cancel returns 202 or 400: how Sume's cancel 409 differs
fal's queue cancel answers 202 CANCELLATION_REQUESTED, 400 ALREADY_COMPLETED or 404. Sume's cancel works only before generation starts, else 409.
- fal webhook redirect 3xx is never retried: what Sume does instead
fal treats a 3xx from your webhook URL as a permanent failure. Sume does not follow redirects either, but counts a 3xx as a failed attempt, not an end.
Written by Sume