Graph API X-App-Usage call_count: pause a Reel publisher early
Meta's X-App-Usage header reports call_count, total_cputime and total_time as percentages. Read it after each call and pause your publisher before the limit.

Meta's Graph API returns an X-App-Usage header with call_count, total_cputime and total_time. call_count is the percentage of calls your app made over a rolling one-hour period, and the others are percentages of the allotted CPU and total time for query processing. Read the header after every response and pause the publisher as it climbs.
Meta advises that once the limit is reached you stop making calls, because continuing to call increases the call count and extends recovery time.
What is in the header?
Meta shows the header like this: {"call_count": 28, "total_time": 25, "total_cputime": 25}.
| Field | Meaning per Meta |
|---|---|
call_count | Percentage of calls made by your app over a rolling one-hour period |
total_cputime | Percentage of CPU time allotted for query processing |
total_time | Percentage of total time allotted for query processing |
How do I react in code?
Pick your own thresholds; Meta does not publish a safe percentage in the page I read. The sample parses the header and returns how long to wait.
import json
def pause_seconds(header: str, soft: int = 70, hard: int = 90) -> int:
usage = json.loads(header or "{}")
worst = max(usage.get("call_count", 0), usage.get("total_cputime", 0), usage.get("total_time", 0))
if worst >= hard:
return 600
if worst >= soft:
return 60
return 0
if __name__ == "__main__":
print(pause_seconds('{"call_count": 28, "total_time": 25, "total_cputime": 25}'))
print(pause_seconds('{"call_count": 92, "total_time": 25, "total_cputime": 25}'))Where do Sume jobs fit?
Rendering is a separate budget. Sume's own limits are reported in ratelimit-* headers and generation_limits, so a Meta pause should hold only the publish queue; queued Sume jobs keep their results (Generation admission).
Sources
Related posts
More in Developers
- Griffin 10 ms audio packets vs Sume TTS: async jobs, no streaming
Tavus Griffin emits speech in packets as small as 10 ms. Sume's TTS Router lists streaming TTS as a non-goal, so audio arrives as a finished file for lip sync.
- Grok Imagine Image 2.0 makes 10 images per request; Sume's n cap
xAI says grok-imagine-image-2.0 returns up to 10 images per request at $0.04 each. Sume's n is 1-10 per request, with a lower per-model cap on x-ai/grok-image.
- H3 Max job: poll or webhook? A signed Python receiver
Poll GET /v1/jobs/:id/status for one clip, use a signed webhook for batches. Python receiver that refuses an empty secret and checks the sume-v1 signature.
- Lip sync for singing: fal can turn off guidance, Sume has no switch
fal's H3 Max Lip Sync is transcription-guided by default and can be switched off for singing or processed audio. Sume's documented body has no such field.
Written by Sume