Haiku 5.5 in Cursor with Sume MCP: a 150k request costs 8x a 90k one

Cursor adds Claude Haiku 5.5 from Settings > Models. Its price steps up above 100k input tokens: 90k costs $0.009 and 150k costs $0.075 before output.

4 min readSume
All posts

To use Claude Haiku 5.5 with Sume in Cursor, add the model under Cursor Settings > Models and keep the Sume server entry in mcp.json. The one number to watch is the price step: Cursor's page lists $0.10 per million input tokens up to 100k, and $0.50 above it. A 90,000-token request costs $0.009 in input and a 150,000-token request costs $0.075, about 8 times more for 1.7 times the tokens.

What Cursor's page says (read 2026-10-08)

Cursor's model page for Claude Haiku 5.5 says the model supports context windows up to 1M tokens, costs $0.10 per million input and $0.50 per million output for requests up to 100k input tokens, and is five times higher above that threshold ($0.50 input, $2.50 output). It says you add the model from Cursor Settings > Models. Anthropic's own page lists the same tiers and the API id claude-haiku-5-5.

Haiku 5.5 input cost per request at the listed rates, read 2026-10-08
Input tokensRate per millionInput cost
30,000$0.100.03 x 0.10 = $0.003
90,000$0.100.09 x 0.10 = $0.009
150,000$0.500.15 x 0.50 = $0.075
300,000$0.500.30 x 0.50 = $0.15

Connect Sume in the same Cursor window

Sume's hosted MCP endpoint is https://mcp.sume.com/mcp. Per Sume's quickstart, add it as a remote server in Cursor Settings > MCP or in your MCP config file, then complete the OAuth sign-in when Cursor prompts. The consent page is on the MCP host, not app.sume.com.

{
  "mcpServers": {
    "sume": {
      "url": "https://mcp.sume.com/mcp"
    }
  }
}

Keep the request under 100k

The price step applies per request, so what matters is how much each turn carries. Three habits from Sume's docs help:

  • Stay on OAuth with Write off while you explore. The session then lists only read-only tools, which shortens the tool list.
  • Wait for renders with one batch jobs_wait of up to 20 ids and include_results: true, instead of many separate calls whose results pile up in the chat.
  • Never paste signed URLs or OAuth tokens into the chat; they add tokens and are secrets.

A worked example: one render wave

Say an agent fans out 20 clips and then waits. Each separate wait returns a small result, but the whole conversation is re-sent on every turn. If the chat holds 12,000 tokens before the wave, 20 single waits that each add 800 tokens grow the context by 16,000 tokens in total, so the last request carries about 28,000 tokens. That is still far below 100k, so the Haiku rate is $0.10: 0.028 x 0.10 = $0.0028 for that final request's input.

The price step only bites on long sessions: a coding chat that already holds 120,000 tokens of files and tool output pays the $0.50 rate on every following turn, even for a one-line status check. In that case, start a fresh chat for the render wave, or keep the Sume work in its own chat, so that the small requests stay on the low rate.

What this post does not claim

Sume does not list Haiku 5.5 as a Sume model, so Sume does not bill for it; Cursor does, under your Cursor plan. Cursor's page also quotes benchmark scores, which are left out here. After you connect, call tools_list once to confirm that the session works.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume