spend_approval_queue_full 429: clear pending approvals, do not retry
A thread with too many pending spend approvals gets 429 spend_approval_queue_full. Resolve the pending ones first; a 503 store_misconfigured is for support.

spend_approval_queue_full is a 429 raised when a Studio Agent thread already has too many unresolved spend approvals and the API cannot queue another one. The 429 status invites a retry, and retryable is true in the envelope by the generic 429 rule, but waiting does not help. A person has to approve or reject the pending ones first.
Where it comes from
When a generation needs interactive confirmation, the API saves a pending approval so that a person can approve it later, and then answers 402 spend_confirmation_required. If saving that record fails because the thread's pending queue is full, the API answers 429 spend_approval_queue_full instead. The body includes studio_agent_thread_id, operation_type and requested_usd_micros, so you know which thread and which operation to look at.
Two nearby errors
Keep the three apart, because they send you to three different places:
| Code | Status | What to do |
|---|---|---|
| spend_confirmation_required | 402 | Ask a person to approve; resend with the same Idempotency-Key |
| spend_approval_queue_full | 429 | Resolve pending approvals on the thread, then resend |
| spend_approval_store_misconfigured | 503 | Not retryable by mapping; contact support |
A backoff that is not blind
Unlike a rate limit, the queue drains by human action. A client should stop submitting paid work on that thread, surface the pending count to the owner and resume when approvals are cleared:
import json
BODY = '''{"error": {"code": "spend_approval_queue_full", "retryable": true,
"details": {"studio_agent_thread_id": "thr_example",
"operation_type": "images.create", "requested_usd_micros": 4500000}}}'''
err = json.loads(BODY)["error"]
if err["code"] == "spend_approval_queue_full":
t = err["details"]["studio_agent_thread_id"]
print(f"PAUSE thread {t}: approvals are waiting on a person; do not loop")Show it to the person who can fix it
The useful output of this error is a message to a human. Say which thread is blocked, how much the refused operation would have cost (divide requested_usd_micros by 1,000,000), and what to do: open the thread and approve or reject the waiting items. Then resume with the same Idempotency-Key, so the resumed call is the same intent and cannot create a second job.
Why the 503 is different
spend_approval_store_misconfigured says that the approval store, which needs shared Redis, is not set up correctly. It is a platform problem and not a state of your thread, and the generic 503 branch marks it retryable: false with next_action: contact_support. Keep the request_id and report it.
Both live in the same gate as the per-run cap automation_generation_spend_cap_exceeded, which only applies to automation runs. See the related posts for those, and the Errors and credits page for the envelope.
If you operate an automation that can create many paid requests at once, cap your own parallel paid submits per thread below the pending limit. That keeps the queue from filling in the first place, and it makes a person's review of the approvals manageable, since each item is a real spend decision that should be read, not skimmed.
Sources
Related posts
More in Developers
- spend_confirmation_required 402 on Sume: why retrying will not help
A 402 spend_confirmation_required means a person must approve the spend first. It is not a balance error: retryable is false and next_action is fix_input.
- Split a 10-minute TikTok into parts for 3 and 5-minute accounts
TikTok's API allows up to 10 minutes, but an account may be limited to 3 or 5. Split a 600-second video into 4 or 2 parts with Sume trim at $0.02 a job.
- Split a 40-second brief into four Omni prompts, one character block
A Python script that turns one 40-second brief into four 10-second Gemini Omni 1.1 Flash request bodies for Sume, with one shared character block.
- Spread Graph API calls evenly: pace a nightly Reel batch
Meta advises spreading queries evenly to avoid traffic spikes. Space publish calls across the hour, and size the Sume render wave from generation_limits.
Written by Sume