Spring AI MCP client request-timeout 20s vs Sume jobs_wait
Spring AI's MCP client defaults request-timeout to 20s, shorter than a 50s Sume jobs_wait. Raise it with a customizer or keep waits short and re-issue them.

Spring AI's Boot starter for MCP clients ships with a request timeout, and the default is shorter than Sume's default jobs_wait hold. According to the Spring AI MCP client page, request-timeout defaults to 20s, read 2026-10-03. Sume's tools and gates page says jobs_wait holds for 50 seconds by default and is capped at 55. A default Spring client will give up on a wait roughly 30 seconds before Sume answers.
Where the timeout lives
The starter artifacts are spring-ai-starter-mcp-client and a -webflux variant. A Streamable HTTP server is configured under spring.ai.mcp.client.streamable-http.connections.<name>.url, with an optional .endpoint that defaults to /mcp. Sume's hosted MCP server lives at https://mcp.sume.com/mcp, so the URL can be the host and the default endpoint fits.
The client type is SYNC by default, and the page says mixing sync and async clients is not allowed. Two settings matter for long jobs: request-timeout and initialized, which defaults to true.
spring:
ai:
mcp:
client:
request-timeout: 60s
streamable-http:
connections:
sume:
url: https://mcp.sume.comTwo ways to fit the wait
The page documents McpClientCustomizer<McpClient.SyncSpec> as the hook for per-client settings, including spec.requestTimeout(Duration). It does not cover request headers, so how you attach a Sume credential on this path is not something this post can confirm from the Spring page; check the starter's current docs before relying on a property name.
| Option | How | Trade-off |
|---|---|---|
| Raise the timeout | request-timeout: 60s or a McpClientCustomizer calling spec.requestTimeout(Duration) | One slow tool call blocks that thread for up to the cap |
| Shorten the wait | Pass a smaller wait_seconds-style value in the call | More round trips, each under 20 seconds |
| Both | Timeout above 55 s and re-issue on wait_slice_expired | Safest for unattended work |
The rule that does not change
A timeout on the Spring side is a client transport failure, not a job outcome. Per Sume's docs, on wait_slice_expired you re-issue the wait with the same job ids, and you never resubmit the paid create. Paid and write tools require an idempotency_key, so even a retried create returns the same job rather than a second charge when the key repeats.
If the application restarts mid-wait, the job is still running on Sume. Persist the job id the moment the create returns, then resume with jobs_wait or jobs_result after restart.
A checklist
- Set
request-timeoutabove 55 seconds, or keep every wait under 20. - Treat a 524, 522, 523 or 525 as a transport failure, not a failed job.
- Store job ids before waiting.
- Use
jobs_waitwithwait_forset toalloranyfor up to 20 ids at once.
Sources
Related posts
More in Developers
- SSML in text to speech: Sume takes a plain transcript, no ssml field
Does Sume's text to speech accept SSML? The tts_create body has a plain transcript and rejects unknown keys. What to use for speed, volume, emotion and pauses.
- Stippled AI graphics turn to gray mush when resized: a downscale test
A dotted AI graphic can lose most of its contrast when downscaled. A Pillow test of nearest, bilinear and Lanczos on a stipple, and what to request instead.
- Streaming transcripts: burn only final text into clip captions
MAI-Transcribe-2-Streaming returns partials in about 100 ms, then settled text. Burned-in captions need the final text. Filter it and pass it to Sume as cues.
- STT sentence segmentation: unpunctuated speech split on silence
Sume STT 1.0 segmentation mode sentence returns gapless sentence segments, splitting unpunctuated runs on silence. Options, boundary_lead_ms and caveats.
Written by Sume