Clef-flash as a yes/no gate before a paid Sume Agent Completion
Cloudflare's Clef-flash returns typed answers with probabilities. Put one in front of POST /v1/agent/completions as a filter, keep the spend cap as the guard.

Yes. A small decision model such as Cloudflare's Clef-flash can sit in front of a paid Sume Agent Completion and answer one question: should this item start a run at all? It works as a cheap filter. It does not replace the guards on the paid call. Cloudflare describes Clef as producing "bounded structured outputs" with probabilities, and lists a median latency of 38.8 ms for Clef-flash against 209.3 ms for Clef on its launch post.
What each step returns
The two steps do different jobs and return different things. The decision is synchronous and tiny. The Agent Completion is asynchronous: it answers 202 with an agent.run receipt, and you poll status_url until next_action is no longer poll_status.
| Step | Question it answers | What comes back | Source |
|---|---|---|---|
| Clef-flash | Should this item start a run? | A typed answer with a probability | Cloudflare launch post |
| POST /v1/agent/completions | Do the task | 202 and an agent.run receipt | Sume Agent Completions docs |
| GET status_url | Is it done? | status and next_action | Sume Agent Completions docs |
Receipt usage | What was the ceiling? | generation_spend_cap_usd_micros | Sume Agent Completions docs |
The gate in Python
Compare the probability with a threshold that you own, and start the run only on a yes. Send an Idempotency-Key derived from the item, so that a retry of the whole step never starts a second run. The decide function below is a stub. Replace its body with your Clef-flash call.
import os
import requests
API = "https://api.sume.com/v1"
KEY = os.environ["SUME_API_KEY"]
def decide(text: str) -> float:
"""P(yes) for: does this ticket need a video reply?"""
return 0.97 # replace with your Clef-flash call
def start_if_yes(ticket_id: str, text: str, threshold: float) -> str | None:
if decide(text) < threshold:
return None
r = requests.post(
f"{API}/agent/completions",
headers={"Authorization": f"Bearer {KEY}",
"Idempotency-Key": f"ticket-{ticket_id}-v1"},
json={"instruction": "Make a 20-second video reply to the ticket in the input.",
"input": {"ticket": text},
"generation_spend_cap_usd": 2},
timeout=30,
)
r.raise_for_status()
return r.json()["data"]["id"]
What the gate cannot do
generation_spend_cap_usd has no default on Agent Completions. Without it, the request fails with 400 invalid_request. The cap replaces the approval prompt that a person sees in the chat UI, so a gate that says yes does not make a run safe. It only makes the run likely to be wanted.
Fail closed. If the decision call times out or returns something that you cannot read, skip the item and log it. A skipped item costs nothing. A run started on a guess is capped, but it is still spend. Sume's safe automation page asks agents to keep read-only work separate from spend, and the gate is the read-only half of that split.
Sources
Related posts
More in Agents
- A decision model's confidence is not a spend cap: Sume's real gates
Strands Decider and Clef return confidence scores. Only the Sume gates in this table, from idempotency_key to the required spend cap, limit what a run can cost.
- DeepSeek V4.1 Flash sees images: hand one to a Sume Agent Completion
V4.1 Flash is described as natively multimodal. To act on an image with Sume, pass it as an input_image attachment on an Agent Completion, up to 30 per run.
- Does the Sume run spend cap include the LLM turn? No, only generation
On a Sume run receipt, billable_amount_usd_micros is generation spend only; the agent's LLM turn bills a separate Agent wallet. How to budget both.
- Fire a Sume schedule from a decision, not a clock: API trigger
Set api_trigger_enabled on a Sume schedule, let Clef or Decider decide when it is worth running, then POST /v1/actions/{action_id}/runs with an idempotency key.
Written by Sume