Sume docs for a coding assistant: llms.txt, llms-full.txt, .md
docs.sume.com serves a 73-link llms.txt index, one llms-full.txt file and any page as markdown by adding .md. Which to hand to a coding assistant, and when.

To give a coding assistant the Sume docs, point it at https://docs.sume.com/llms.txt for the index and fetch only the pages the task needs by adding .md to a docs URL. Use https://docs.sume.com/llms-full.txt only when the assistant has room for it: the file was about 530 KB when I fetched it on 2026-10-04, far larger than a typical working context.
The llms.txt convention is a markdown file at a site root with a title, a short summary and links to detail pages (llmstxt.org, read 2026-10-04). Sume's file follows it, and its own first lines say that individual pages are available as clean markdown by appending .md to any page URL.
What does each file contain?
| URL | Size when fetched | Use it for |
|---|---|---|
| docs.sume.com/llms.txt | 12,095 bytes, 73 linked pages | Letting the assistant choose pages |
| docs.sume.com/llms-full.txt | 541,809 bytes | A one-shot load into a large context or a search index |
| docs.sume.com/sdk.md | 6,449 bytes | SDK install, client, and a first run |
| docs.sume.com/workflows/webhooks.md | 6,816 bytes | Signature rules and delivery behavior |
Which approach works best for a specific task?
Fetch the index, then the two or three pages that match the job. For a webhook receiver that is Webhooks and Errors and rate limits; for a TypeScript client it is the SDK page. A short script keeps this repeatable.
import os, urllib.request
BASE = "https://docs.sume.com"
PAGES = ["/sdk", "/workflows/webhooks", "/workflows/errors-and-credits"]
def fetch(path: str) -> str:
req = urllib.request.Request(BASE + path + ".md",
headers={"user-agent": "docs-context/1.0"})
with urllib.request.urlopen(req, timeout=30) as r:
return r.read().decode("utf-8")
if __name__ == "__main__":
out = os.environ.get("DOCS_CONTEXT_OUT", "sume-context.md")
with open(out, "w", encoding="utf-8") as f:
for p in PAGES:
f.write(f"\n\n<!-- {p} -->\n" + fetch(p))
print(out, os.path.getsize(out), "bytes")When is the full file the right choice?
Use llms-full.txt when the tool indexes it for retrieval, for example in an editor that chunks and searches documentation, so only the matching passages reach the model. Pasting all 541,809 bytes into one prompt spends context on pages the task never touches. For a single endpoint question, the one .md page is both cheaper and easier to check against the source. Sume's own llms.txt says the full file contains the contents of the docs in a single file, so a refetch after a docs change replaces the whole thing.
What should you not rely on?
- Do not paste facts from an old copy: the files change as docs change, so refetch per task and check the page itself for limits and prices.
- Do not treat the docs as the OpenAPI spec. For exact request and response shapes, fetch
https://api.sume.com/reference/json. - An assistant that cannot fetch URLs needs the content pasted in; the per-page
.mdfiles are the smaller thing to paste.
Sources
Related posts
More in Developers
- Luma callback_url or polling: which Sume job mode matches
Luma's API docs list keyframes, loop and callback_url for ray-2. On Sume the equivalent choice is job mode: async, sync up to 30 s, subscribe or webhook.
- Luma API callback_url vs Sume callback_url: signing and retries
Luma's video API takes a callback_url, and so does Sume's /v1/videos. What Sume's callback is signed with, how often it retries, and a Python verifier.
- Luma callback: 3 retries, 5 s timeout, vs Sume webhooks
Luma retries a failed callback at most 3 times with a 5-second timeout. Sume retries job webhooks up to 10 times with a 10-second timeout and signs each body.
- Lyria 3.5 is single-turn: budget retakes, not edits, on Sume
Google says Lyria 3.5 is single-turn and varies between calls. Plan retakes instead of edits: what 1, 3, 5 and 10 attempts cost on Google and on Sume.
Written by Sume