fal MCP server: setup for Claude Code, Cursor, and billing

fal's MCP server runs at mcp.fal.ai/mcp with a fal API key or sign-in. The server is free; model runs bill to your fal account at API prices.

5 min readSume
All posts

fal's MCP server is a hosted endpoint at https://mcp.fal.ai/mcp that lets Claude Code, Cursor, Windsurf, and other Streamable HTTP clients search fal's models, check pricing, run inference, and upload files. You connect with a fal API key in an Authorization header, or with a fal sign-in through its plugin and OAuth connectors; the server itself is free, and model runs bill to your fal account.

The fal facts come from fal's Run MCP docs page and its docs MCP server, both read on 2026-09-28. The docs server at https://fal.ai/docs/mcp is a different, read-only server, covered below. Sume has no connection to either; Sume vs fal compares the platforms.

How do I add the fal MCP server to Claude Code, Cursor, or Codex?

fal's setup tabs cover ChatGPT and Codex, Claude Code, Cursor, Windsurf, and other MCP clients. For Claude Code with an API key, fal gives this command; replace YOUR_FAL_KEY with your fal API key, then run /mcp to check that fal-ai is connected:

  • Cursor: add a fal-ai entry to mcpServers in ~/.cursor/mcp.json, with "url": "https://mcp.fal.ai/mcp" and an Authorization: Bearer header.
  • Windsurf: the same entry in ~/.codeium/windsurf/mcp_config.json, with serverUrl in place of url.
  • ChatGPT and Codex: install the fal plugin and connect your fal account when prompted.
  • To test, ask: "Use fal to search for image generation models. Do not run a model." A working connection returns model search results.
claude mcp add --transport http fal-ai \
  https://mcp.fal.ai/mcp \
  --header "Authorization: Bearer YOUR_FAL_KEY"

What tools does the fal MCP server have?

fal lists 11 tools in three groups:

From fal's Run MCP docs, read 2026-09-28.
GroupToolsWhat fal says they do
Discoverysearch_models, get_model_schema, get_pricing, search_docsSearch the 1,000+ model catalog, read a model's inputs and outputs, check a model's cost before running, search the docs
Executionrun_model, submit_job, check_job, get_job_result, cancel_jobRun with a bounded wait, or submit and poll a long job; cancel a queued or running one
Utilityupload_file, recommend_modelUpload a public file URL or base64 data to fal's CDN; get model recommendations

Is fal.ai/docs/mcp the same server?

No. fal documents separate MCP servers with different jobs, and only the first one below runs models:

  • https://mcp.fal.ai/mcp is the hosted model server described above. It takes a fal API key or a fal sign-in, runs models, uploads files, and bills model runs to your fal account.
  • https://fal.ai/docs/mcp is a docs server. Its tools are search_fal and query_docs_filesystem_fal, which search and read fal's documentation pages and OpenAPI specs, plus submit_feedback for reporting a docs problem. fal says that apart from submit_feedback it is read-only and does not access anything beyond the published site content, so it cannot run a model.
  • The Run MCP page also points to a third server, the Platform MCP, for account operations and serverless debugging, and says you can use it together with the model server.

How does billing work on the fal MCP server?

fal's FAQ says the MCP server is free and you pay only for the model runs you trigger, at the same pricing as direct API calls. The server calls fal's Platform API with the connected account's permissions, and get_pricing returns a model's per-run cost before you run it.

  • API-key connections bill the account the key belongs to.
  • OAuth connections use your personal fal account until you choose another under Account settings with Use for MCP. If access to the selected account fails, operations stop instead of using your personal credits.
  • The MCP endpoint adds no rate limits of its own; the same concurrency limits as direct API calls apply.
  • Your key or OAuth token travels per request in the Authorization header; fal says the key is not persisted or logged as a credential.

Why does a video job come back unfinished?

run_model waits up to 45 seconds by default, then returns either the result or processing with a request ID. fal's docs say to use submit_job, check_job, and get_job_result from the start for video, 3D, and training. Don't resubmit to check progress: another run_model or submit_job call creates a new billable job.

How is Sume's hosted MCP server different?

It is the same kind of remote server with its own tools, sign-in, and wallet. Sume basics says hosted MCP still works but is not the primary path today. The differences that matter when you run both:

  • Auth: Sume's OAuth consent is read-only by default, and paid tools appear only after you turn Write on (mcp:write); an API key sees the full tool set (OAuth and API keys). Connect Claude Code, Cursor, or Codex to Sume covers setup, per-call spend controls, and billing.
  • Files: fal's upload_file takes base64 for small local files under 1 MB or a public file URL, and a local path sent to fal's hosted server fails. Sume's hosted MCP cannot read files from your laptop (Tools and gates), so generation inputs are public HTTPS URLs.
  • Long jobs: Sume's generation tools are waited on with jobs_wait; see MCP tool call timeouts on long video jobs.

Sources

Related posts

More in Integrations

All Integrations posts

Written by Sume