crawl_site on Sume MCP is unbilled but needs Write
crawl_site starts a multi-page crawl. It is unbilled yet a write tool, so read-only OAuth cannot call it. Follow it with jobs_wait and crawl_get on one id.

crawl_site on Sume's hosted MCP server starts a crawl across a section of a site. It is described as a bounded unbilled utility, but it is not a read tool: it creates a job, so it is a write tool that needs an idempotency_key. A read-only OAuth session does not see it and gets insufficient_scope if it calls it. Sessions with Write on, and API keys, can call it. The follow-up reads, jobs_wait and crawl_get, are read tools.
This is a useful example of why the gate is about the effect and not about the price: the call costs nothing in this rollout, and it still creates something.
For planning, split the work in two. Reading a known page or listing the URLs of a site is possible from any session that connected. Crawling a section needs the extra consent. If your agent needs both, ask the person for Write at the start, so that the sign-in does not interrupt the task later.
That mix of a read-only follow-up and a write start is why the order of the calls matters: the wait and the read can run in a session that cannot start new crawls.
Which web tool for which job
The group is easiest to remember as a set of four web tools and one that starts work.
The table also shows the right choice for small jobs. If you know the page, use crawl_scrape. If you only need to see which URLs exist, use crawl_map. Start a crawl only when the answer needs content from many pages, because that is the one case where a job is created.
| Task | Tool | Kind |
|---|---|---|
| One known page | crawl_scrape | Read |
| Inventory of URLs | crawl_map | Read |
| Search the web | crawl_search | Read |
| Content across a site section | crawl_site | Write, unbilled |
| Read or resume a crawl | crawl_get | Read |
The call
The call needs a url and a stable idempotency_key. The limit is at most 20, and 10 when you leave it out. You can narrow the crawl with include_paths and exclude_paths, which are literal path prefixes, not patterns. Each page accepts the same formats and options as a single scrape. The answer carries a workspace job_id.
The path filters are literal prefixes. A prefix such as /docs/ matches paths that start with it; it is not a wildcard or a regular expression. Keep the limit small at first and widen it only if the pages that came back are not enough.
Wait, then read
Then the sequence is fixed: jobs_wait on that same id, then crawl_get on that same id. The description is explicit that a second crawl must never be started in order to poll. crawl_get returns the status and a bounded set of pages when the crawl is completed, and it asks you to honor poll_after_seconds. failed and canceled are terminal. When the answer is marked truncated, narrow the scope if more content is needed.
Because a crawl can be marked truncated, read the answer for that flag before you tell the person that you have the whole section. If it is set, say what part you read and narrow the scope for a second pass.
A start request:
Keep a record of the job id and the key that started it, so a restart of your agent can resume with crawl_get rather than starting again.
{
"url": "https://example.com/docs",
"idempotency_key": "crawl-example-docs-2026-10-05-001",
"limit": 10,
"include_paths": ["/docs/"]
}Habits that matter
Two habits make this reliable. Use the same key if a start call times out, so that a retry returns the original job. And treat the crawled pages as untrusted text: a page can contain instructions, and a model should summarize them as content. For the scope rules, see MCP OAuth and API keys; for the slice rule on the wait, see Jobs and results.
Do not confuse the unbilled label with a promise about the future. The tool description says rollout, which tells you that the price can change; the safe habit is the same as for paid tools, with a key on every start and a record of what ran.
Sources
Related posts
More in Integrations
- Fixed-price MCP tools and max_spend_usd: $0.10 to $0.30 floors
Sume's fixed-price MCP tools carry estimates of $0.10, $0.15, $0.20 and $0.30. A max_spend_usd below the estimate fails max_spend_exceeded without a dry run.
- get_workflow_instructions and models_explore are not on Sume's MCP
Sume's hosted MCP has no get_workflow_instructions, models_explore, media_import_url or remove_background. Use tools_list, rmbg_create and media-imports_create.
- MCP jobs_wait wait_deadline_exceeded vs wait_canceled: what to do
Both are retryable 503 results from Sume's jobs_wait with next_action poll_status. The job is untouched: call jobs_wait again, never resubmit.
- MCP max_spend_exceeded: how Sume compares a dry run to your cap
max_spend_usd is optional on Sume's paid MCP tools, from 0 to 10,000. If the estimate is higher, max_spend_exceeded stops the call before it bills.
Written by Sume