Firecrawl includePaths regex vs Sume's literal path prefixes
Firecrawl's includePaths are regex on the URL pathname. Sume's include_paths are literal prefixes starting with /, max 10 items. How to rewrite one.

Sume's include_paths and exclude_paths are literal path prefixes, not regular expressions. Each item must start with /, be at most 200 characters, and each list holds at most 10 items. Firecrawl's includePaths are RE2-style regex matched against the pathname, so a pattern like /blog/.* must become the prefix /blog/.
Firecrawl's behavior is from its crawl guide and Sume's from OpenAPI and the crawl_site description, read 2026-09-30.
How do I rewrite a regex as prefixes?
Take the fixed start of each pattern, and split alternations into separate items.
| Firecrawl pattern | Sume include_paths |
|---|---|
/blog/.* | ["/blog/"] |
/docs/(api|sdk)/.* | ["/docs/api/", "/docs/sdk/"] |
| A pattern ending in a file extension | No equivalent; filter the returned URLs |
regexOnFullURL: true | No equivalent; only paths are matched |
What are the limits on the lists?
Ten items per list and 200 characters per item, and each item must match the schema pattern ^/. A full URL such as https://example.com/blog/ fails validation, so pass only the path.
What does a request look like?
Crawl is a write call, so it needs an Idempotency-Key header.
const res = await fetch("https://api.sume.com/v1/firecrawl/crawl", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.SUME_API_KEY,
"Content-Type": "application/json",
"Idempotency-Key": "crawl-blog-001",
},
body: JSON.stringify({
url: "https://example.com",
limit: 20,
include_paths: ["/blog/"],
exclude_paths: ["/blog/tag/"],
}),
});
console.log(res.status, await res.json());What do I do with a pattern I cannot express?
Run crawl_map (limit 1-100) on the origin, which returns links with titles and descriptions, filter that list with your own regex in code, and crawl_scrape the survivors. That keeps the regex logic on your side and avoids crawling pages you will discard. The cap on crawl size is covered in the crawl limit post.
Sources
Related posts
More in Developers
- Firecrawl executeJavascript return value on Sume
Firecrawl returns executeJavascript results in actions.javascriptReturns. Sume names it javascript_returns and bounds it; scripts max 4,000 characters.
- Firecrawl map endpoint limit 5000 vs Sume crawl_map limit 100
Firecrawl's /map defaults to 5000 links and allows 100000. Sume's crawl_map takes limit 1 to 100, so pick URLs, then scrape them with crawl_scrape.
- Firecrawl map sitemap skip/include/only: Sume crawl_map has no switch
Firecrawl's /map has a sitemap option (skip, include, only) and includeSubdomains. Sume's crawl_map takes only url and limit, and maps one origin.
- Firecrawl MCP timeout: Sume crawl_scrape's 35, 40 and 45 s deadlines
A synchronous crawl_scrape call has three deadlines: client abort at 35 s, MCP at 40 s, host watchdog at 45 s. Narrow the page or use crawl_site.
Written by Sume