TikTok comment reply photo carousel: nine images from Sume

TikTok comments can now carry a carousel of up to nine photos (reported). Plan nine reply images with the Sume image API and price them from $0.225.

5 min readSume
All posts

Metricool reported (read 2026-10-03) that TikTok's comment section now supports voice comments of up to 60 seconds, polls with up to five answers, photo carousels of up to nine images and Live Photos. Nine reply images cost $0.225 on Qwen Image or Grok Imagine at $0.025 each, $0.3375 on Flux 2 Pro, and $0.90 on Nano Banana 2 at its default size. Sume makes the pictures; you add them in the TikTok app yourself.

    What was reported

    The item comes from Metricool's TikTok news page, a third-party roundup rather than TikTok's own page, so check the feature in your app. This post uses only the nine-image figure from it.

      Plan nine images

      A carousel reply works like a short answer: the first image is the answer, the next ones show steps or proof. Most Sume image models return up to four images per call (Grok Imagine returns one), so nine different steps are nine calls, or fewer if you ask for variants with n.

      • Image 1: the one-line answer as a clean visual, no text rendered by the model.
      • Images 2 to 8: one step or example each, in the same style.
      • Image 9: the call to action or the result.
      • Ratio: 3:4 gives a tall look and is a listed aspect_ratio value in the image docs.

      Cost of nine

      Prices are the catalog per-image prices (read 2026-10-03); retakes are extra.

        Nine reply images on Sume (read 2026-10-03)
        ModelPer imageNine imagesThree takes of each
        Qwen Image$0.025$0.225$0.675
        Flux 2 Pro$0.0375$0.3375$1.0125
        Seedream 4.5$0.05$0.45$1.35
        Nano Banana 2$0.10$0.90$2.70

        A script

        It makes nine images, one call per step, with one shared style sentence so the set matches.

          import os, requests
          
          H = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
          STYLE = "flat vector illustration, warm palette, no text"
          STEPS = ["a kettle on a stove", "a tea bag in a cup", "water pouring", "steeping for three minutes",
                   "removing the bag", "adding honey", "stirring", "a hand lifting the cup", "a finished cup"]
          urls = []
          for step in STEPS:
              r = requests.post("https://api.sume.com/v1/images", headers=H, timeout=120, json={
                  "model": "qwen/qwen-image", "aspect_ratio": "3:4",
                  "prompt": f"{step}, {STYLE}"})
              if r.status_code != 200:
                  raise SystemExit(f"{step}: {r.status_code} {r.text[:200]}")
              urls.append(r.json()["data"][0]["url"])
          print(len(urls), "images")
          print("\n".join(urls))
          

          Words belong in the app

          OpenAI's image guide (read 2026-10-03) says its models can still struggle with precise text placement and clarity. Type your captions in the TikTok app, not in the render.

            Sources

            Related posts

            More in Use cases

            All Use cases posts

            Written by Sume