Clef-flash as a yes/no gate before a paid Sume Agent Completion

Cloudflare's Clef-flash returns typed answers with probabilities. Put one in front of POST /v1/agent/completions as a filter, keep the spend cap as the guard.

5 min readSume
All posts

Yes. A small decision model such as Cloudflare's Clef-flash can sit in front of a paid Sume Agent Completion and answer one question: should this item start a run at all? It works as a cheap filter. It does not replace the guards on the paid call. Cloudflare describes Clef as producing "bounded structured outputs" with probabilities, and lists a median latency of 38.8 ms for Clef-flash against 209.3 ms for Clef on its launch post.

What each step returns

The two steps do different jobs and return different things. The decision is synchronous and tiny. The Agent Completion is asynchronous: it answers 202 with an agent.run receipt, and you poll status_url until next_action is no longer poll_status.

Gate and run compared (read 2026-10-05)
StepQuestion it answersWhat comes backSource
Clef-flashShould this item start a run?A typed answer with a probabilityCloudflare launch post
POST /v1/agent/completionsDo the task202 and an agent.run receiptSume Agent Completions docs
GET status_urlIs it done?status and next_actionSume Agent Completions docs
Receipt usageWhat was the ceiling?generation_spend_cap_usd_microsSume Agent Completions docs

The gate in Python

Compare the probability with a threshold that you own, and start the run only on a yes. Send an Idempotency-Key derived from the item, so that a retry of the whole step never starts a second run. The decide function below is a stub. Replace its body with your Clef-flash call.

import os
import requests

API = "https://api.sume.com/v1"
KEY = os.environ["SUME_API_KEY"]


def decide(text: str) -> float:
    """P(yes) for: does this ticket need a video reply?"""
    return 0.97  # replace with your Clef-flash call


def start_if_yes(ticket_id: str, text: str, threshold: float) -> str | None:
    if decide(text) < threshold:
        return None
    r = requests.post(
        f"{API}/agent/completions",
        headers={"Authorization": f"Bearer {KEY}",
                 "Idempotency-Key": f"ticket-{ticket_id}-v1"},
        json={"instruction": "Make a 20-second video reply to the ticket in the input.",
              "input": {"ticket": text},
              "generation_spend_cap_usd": 2},
        timeout=30,
    )
    r.raise_for_status()
    return r.json()["data"]["id"]

What the gate cannot do

generation_spend_cap_usd has no default on Agent Completions. Without it, the request fails with 400 invalid_request. The cap replaces the approval prompt that a person sees in the chat UI, so a gate that says yes does not make a run safe. It only makes the run likely to be wanted.

Fail closed. If the decision call times out or returns something that you cannot read, skip the item and log it. A skipped item costs nothing. A run started on a guess is capped, but it is still spend. Sume's safe automation page asks agents to keep read-only work separate from spend, and the gate is the read-only half of that split.

Sources

Related posts

More in Agents

All Agents posts

Written by Sume