Vocabulary flashcard video: one still and one spoken word per card
Build a vocabulary flashcard video: one image and one TTS word per card, joined with Timeline audio concat, then re-based into render slots from segments[].

A vocabulary flashcard video is one picture and one spoken word per card, in order. On Sume you make the picture with an image job, the word with a TTS job, join the spoken words gapless with Timeline audio concat, and use the segments[] it returns as the start times for the stills in Timeline 1.0.
The reason to do it in that order is timing. The audio decides how long each card lasts, so the video slots should be computed from the audio, not guessed.
What are the pieces?
Each piece is a separate job that returns a media.sume.com URL, which is what the later steps require. The Models overview lists the families; the flow below uses image generation, TTS 1.0 (POST /v1/tts-1.0/generate), Timeline audio and Timeline render.
| Step | Endpoint | Output | Price |
|---|---|---|---|
| Card picture | POST /v1/image-1.0/generate (one per word) | Image artifact | Per image, see catalog |
| Spoken word | POST /v1/tts-1.0/generate (one per word) | Audio artifact | Per character, see catalog |
| Join the words | POST /v1/timeline-1.0/audio, operation concat | One audio_url plus segments[] | $0.01 flat |
| Assemble | POST /v1/timeline-1.0/render | MP4 | $0.10 per ceil(output minute) |
How do I join the words and read the timings?
Concat takes 1 to 20 ordered parts, each { url, source_in?, duration? }, and joins them in the sample domain: no re-synthesis and no silence at the seams. All parts must share one channel layout, or the job fails with audio_parts_channel_mismatch. The produced audio is capped at 1800 seconds.
Twenty cards is the ceiling for one concat. For a longer deck, concat in groups of 20 and then concat the group files, since a part is just a URL.
curl -X POST https://api.sume.com/v1/timeline-1.0/audio \
-H "Authorization: Bearer $SUME_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: flashcards-unit-4-audio-v1" \
-d '{
"operation": "concat",
"parts": [
{ "url": "https://media.sume.com/artifacts/artf_demo/word-01.wav" },
{ "url": "https://media.sume.com/artifacts/artf_demo/word-02.wav" }
]
}'The result has kind: timeline_audio, one audio_url, the total duration_seconds, and segments[] with index, start and duration_seconds for each part. Those offsets are exactly what Timeline's video[].start needs.
How do I turn segments into render slots?
One still per segment, starting where its word starts and lasting as long as the word. The render rules matter here: video[0].start must be 0, later starts must increase, each slot must be at least 0.2 seconds, and coverage may trail the audio by at most 0.5 seconds. This small script builds the body from a concat result, using sample data so you can run it as is.
import json
segments = [{"index": i, "start": round(i * 3.2, 1), "duration_seconds": 3.2} for i in range(12)]
stills = [f"https://media.sume.com/artifacts/artf_demo/card-{i:02d}.png" for i in range(12)]
total = round(sum(s["duration_seconds"] for s in segments), 1)
body = {
"audio": {"url": "https://media.sume.com/artifacts/artf_demo/words.wav", "duration_seconds": total},
"video": [
{
"source_url": stills[s["index"]],
"start": s["start"],
"duration": s["duration_seconds"],
**({"transition": {"type": "fade", "duration": 0.25}} if s["index"] else {}),
}
for s in segments
],
"output": {"width": 1920, "height": 1080},
}
print(json.dumps(body, indent=2)[:400])
print("slots:", len(body["video"]), "seconds:", total)Before spending anything, post the same body to POST /v1/timeline-1.0/plan, the unbilled compile preflight. It returns duration_seconds, segment_count and estimated_cost_usd_micros, and it needs no idempotency key. Then send it to /render with one. A 38-second deck reserves one output minute, so $0.10.
What are the limits for a teacher or course team?
Stills are static holds: Timeline accepts a motion field on them and ignores it with a motion_ignored warning, so there is no Ken Burns drift. Pronunciation is the TTS model's, and Sume has no quiz or interaction layer, so the output is a plain MP4 you upload to your LMS or video host.
For a word list you will reuse, keep the TTS wav files and the card images. Re-rendering a changed card order costs the $0.10 render, plus the $0.01 concat if the audio order changes, and no new speech.
Sources
Related posts
More in Use cases
- Vrbo photo rules: 6 photos, 1024x683, no overlays
Vrbo requires at least 6 published photos at 1024 x 683 or larger, with no text or watermark overlays. Plan a compliant gallery and use AI only for promos.
- Walmart image file names: GTIN-14 and 100x100 swatches in a batch
Walmart recommends the 14-digit GTIN in each image file name and lists swatches at 100x100 px. How to name and size a batch of generated images to match.
- Walmart image URLs: ports 80/443/8080/8443, no Dropbox or queries
Walmart accepts image URLs on ports 8080, 80, 443 or 8443 and rejects query strings, non-public links, HTML pages and Dropbox. Check Sume media URLs first.
- Walmart's AI content rule: accurate, truthful, rights-compliant
Walmart's image guide says AI-generated content must be accurate, truthful and rights-compliant, with no logos or non-English text. Review each output.
Written by Sume