OpenRouter /generation lookup vs Sume usage by job_id

OpenRouter looks up cost and native token counts by gen- id. Sume sums cost with GET /v1/usage?job_id=. What each returns and how to reconcile a bill.

5 min readSume
All posts

How do I look up what one call cost?

On OpenRouter you call GET /generation?id=gen-.... On Sume you call GET /v1/usage?job_id=job_... and read summary.debited_usd. Both answer "what did this one request cost", but they differ in shape: OpenRouter returns a record of that generation, while Sume returns a summary folded over every ledger row the job caused.

What does OpenRouter's generation record contain?

Its reference page lists a required id matching gen- plus alphanumerics (128 characters max) and returns cost fields (total_cost, upstream_inference_cost, usage, cache_discount), standard and native token counts (including cached, reasoning and image tokens), performance fields (latency, generation_time, moderation_latency), and provider fields (provider_name, model, upstream_id and provider_responses with fallback attempt details). A missing generation returns 404, and rate limits apply.

This is an inference-accounting view: which provider served it, in what time, with what native tokens.

What does Sume return?

Sume's usage ledger records reserved, captured and refunded rows. Scoping by job_id (or run_id, or thread_id) adds a summary. The figure to quote is debited_usd (and debited_usd_micros): what the wallet actually deducted. held_usd_micros is a hold still open, and refunded_usd_micros is a hold given back after failure, cancellation or queue_full. Neither is spend. final turns true once no hold is open.

Sume's docs also say never to sum rows yourself, because a refunded row keeps its hold amount in billable_amount_usd_micros. Read the summary.

Compared

What each lookup is built to answer.

Cost lookup per request, read 2026-10-02
DetailOpenRouterSume
EndpointGET /generation?id=GET /v1/usage?job_id=
Id shapegen-...job_... (also run_id, thread_id)
Cost fieldtotal_cost, usage, upstream_inference_costsummary.debited_usd
Hold and refund stateNot on the page readheld_usd_micros, refunded_usd_micros, final
Provider detailprovider_name, provider_responsesRaw provider ids are not exposed in public events
Token countsNative and standard countsCost is in USD per job, not per token

How do I reconcile a month of Sume spend?

Page through /v1/usage for the period, group by job id, and compare the summed debited_usd against the invoice. Use the cheap per-job summary to explain any outlier. This loop runs as written once SUME_API_KEY is set.

import asyncio, os
import httpx

async def main():
    h = {"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"}
    async with httpx.AsyncClient(base_url="https://api.sume.com") as c:
        r = await c.get("/v1/usage", params={"job_id": "job_123"}, headers=h)
        r.raise_for_status()
        s = r.json().get("summary", {})
        print(s.get("debited_usd"), s.get("final"))

asyncio.run(main())

What should you store at submit time?

Both lookups need an id you kept. Store the id from the submit response before you do anything else, because the lookup is only as good as the record you hold.

  • On Sume, keep the job id (request_id in the submit response) and your own idempotency key together.
  • Add run_id or thread_id to your record when the work came from a Format or Agent run, since /v1/usage accepts those scopes too.
  • Quote debited_usd to finance, not the sum of reserved rows; holds are not spend until captured.

Sources

Related posts

More in Comparisons

All Comparisons posts

Written by Sume