OpenRouter /generation lookup vs Sume usage by job_id
OpenRouter looks up cost and native token counts by gen- id. Sume sums cost with GET /v1/usage?job_id=. What each returns and how to reconcile a bill.

How do I look up what one call cost?
On OpenRouter you call GET /generation?id=gen-.... On Sume you call GET /v1/usage?job_id=job_... and read summary.debited_usd. Both answer "what did this one request cost", but they differ in shape: OpenRouter returns a record of that generation, while Sume returns a summary folded over every ledger row the job caused.
What does OpenRouter's generation record contain?
Its reference page lists a required id matching gen- plus alphanumerics (128 characters max) and returns cost fields (total_cost, upstream_inference_cost, usage, cache_discount), standard and native token counts (including cached, reasoning and image tokens), performance fields (latency, generation_time, moderation_latency), and provider fields (provider_name, model, upstream_id and provider_responses with fallback attempt details). A missing generation returns 404, and rate limits apply.
This is an inference-accounting view: which provider served it, in what time, with what native tokens.
What does Sume return?
Sume's usage ledger records reserved, captured and refunded rows. Scoping by job_id (or run_id, or thread_id) adds a summary. The figure to quote is debited_usd (and debited_usd_micros): what the wallet actually deducted. held_usd_micros is a hold still open, and refunded_usd_micros is a hold given back after failure, cancellation or queue_full. Neither is spend. final turns true once no hold is open.
Sume's docs also say never to sum rows yourself, because a refunded row keeps its hold amount in billable_amount_usd_micros. Read the summary.
Compared
What each lookup is built to answer.
| Detail | OpenRouter | Sume |
|---|---|---|
| Endpoint | GET /generation?id= | GET /v1/usage?job_id= |
| Id shape | gen-... | job_... (also run_id, thread_id) |
| Cost field | total_cost, usage, upstream_inference_cost | summary.debited_usd |
| Hold and refund state | Not on the page read | held_usd_micros, refunded_usd_micros, final |
| Provider detail | provider_name, provider_responses | Raw provider ids are not exposed in public events |
| Token counts | Native and standard counts | Cost is in USD per job, not per token |
How do I reconcile a month of Sume spend?
Page through /v1/usage for the period, group by job id, and compare the summed debited_usd against the invoice. Use the cheap per-job summary to explain any outlier. This loop runs as written once SUME_API_KEY is set.
import asyncio, os
import httpx
async def main():
h = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
async with httpx.AsyncClient(base_url="https://api.sume.com") as c:
r = await c.get("/v1/usage", params={"job_id": "job_123"}, headers=h)
r.raise_for_status()
s = r.json().get("summary", {})
print(s.get("debited_usd"), s.get("final"))
asyncio.run(main())What should you store at submit time?
Both lookups need an id you kept. Store the id from the submit response before you do anything else, because the lookup is only as good as the record you hold.
- On Sume, keep the job id (
request_idin the submit response) and your own idempotency key together. - Add
run_idorthread_idto your record when the work came from a Format or Agent run, since/v1/usageaccepts those scopes too. - Quote
debited_usdto finance, not the sum of reserved rows; holds are not spend until captured.
Sources
Related posts
More in Comparisons
- OpenRouter models fallback array and 3-entry limit vs Sume
OpenRouter's models array tries the next model on downtime, rate limits or moderation; fallbacks allows 3. Sume's allow_fallbacks has no effect.
- OpenRouter presets vs a Sume Format: what each one pins
An OpenRouter preset stores model, provider rules, prompt and parameters under a slug. A Sume Format is a versioned recipe you run by handle. Which one you pin.
- OpenRouter provider.sort and max_price vs Sume's inert sort
OpenRouter's provider.sort picks price, throughput or latency and turns off load balancing. On Sume's image route, sort is accepted and changes nothing.
- Perso AI dubbing: 10 speakers, 2-speaker lip sync, vs Sume
Perso says it detects up to 10 speakers and lip-syncs two. Sume's avatar video uses one avatar per final video. What to use for a multi-speaker dub.
Written by Sume