Gemini agent keeps memory for days; Sume API runs start fresh
Google's Gemini agent keeps one memory across devices and days. Sume's API runs start in a new thread, so state travels as input, files and previous_run_id.

Google's new Gemini agent is built to remember: Google's announcement says it keeps one set of memories and context across devices and keeps working for hours or days. Sume's API works the other way. An Agent Completion and a Scheduled run each start in a new thread, so anything the next run must know has to be handed to it as input, attached images or a durable media URL, and Format runs can additionally continue an earlier run with previous_run_id.
The Google facts below are from Gemini at Work 2026: Introducing Gemini agent, published October 8, 2026 and read 2026-10-10. The Sume facts are from the docs pages in the sources. This post compares designs; it does not claim Sume replaces a workplace agent.
What Google says the Gemini agent remembers
The announcement says the agent runs in the cloud and "maintains a single set of memories, context, and one personalization graph no matter what device or channel you access it on," and that work lasting hours or days keeps running after you close your laptop. It describes four kinds of memory: session, semantic, procedural and episodic.
It also describes coworker agents that get their own Workspace account, including an email address, calendar and Drive, and says you can assign work, schedule tasks, or have the agent respond to events. The post does not say whether the core agent is generally available, so treat availability as something to check with Google.
How a Sume run keeps (or loses) state
The Sume docs are explicit that an Agent Completion runs in a new thread each time. The receipt's thread_id identifies it, but you cannot send a thread_id back to continue the conversation, and assistant turns in messages[] are rejected rather than ignored. A schedule fires "as an Agent in a fresh thread" too, and a schedule run is not a job, so it does not appear in /v1/jobs.
Format runs are the one place with a documented continuation. A run is one agent turn; sending previous_run_id on the next POST continues the same conversation, replays what the agent produced, and lets it redo one part while keeping the rest. The continuation gets a new id, its own spend cap and its own webhook, and artifacts[] lists all media from the whole conversation while usage stays per run.
| Need | Google Gemini agent | Sume API |
|---|---|---|
| State across calls | One shared memory and context across devices | None by default; each Completion or schedule run is a new thread |
| Pick up where a task stopped | Work continues for hours or days | Format runs: previous_run_id on the next POST |
| Pass known facts in | Learned from questions and objectives | input (up to 64 properties, 2 MiB on schedules) and attachments |
| Long work | Keeps running after you close your laptop | Async receipt; poll status_url or take a signed webhook |
Carrying state yourself
Because the memory is yours to keep, design the hand-off. Four patterns come straight from the docs:
- Put durable facts in
input. Sume writes it, whole, to a file in the run workspace and tells the agent to read it as data, never as instructions. - Pass the last run's output URLs forward. Media URLs on
media.sume.comare durable and do not expire, so a later run can reference an earlier image or video. - Bind an
output_schemaso each run returns the fields the next run needs, such as a brand tone note or the chosen scene id. - For Formats, continue with
previous_run_idand bind the sameoutput_schemaagain, because the schema is per run and not inherited.
Schedules and events
Google says its agent can run scheduled tasks or respond to events. Sume has two triggers for a saved automation: a cron expression with an IANA timezone, or an API call from your service, which is how an external event starts work. The trigger type is fixed when the schedule is created. Overlap is handled with on_active_run: a schedule defaults to skip, which records a skipped run, while reject answers 409.
Schedules are created in the dashboard or by asking the Agent in chat. The Developer API can list them, read them, start runs and monitor runs, but it cannot create or edit one, and there is no run events endpoint: events_url on a schedule run is always empty.
Which to choose
If you need a coworker that remembers last week's conversation across tools, Google's announcement describes that product, and Sume's API does not provide it. If you need a metered media pipeline that your backend calls, each run needs its inputs spelled out, and in return every run has a spend cap, a receipt and durable files. The two are less alternatives than different layers: an agent with memory can decide what to make, then call a capped Sume run to make it.
Sources
Related posts
More in Agents
- Guardrails for an AI agent that calls paid media APIs
Five guardrails for an agent that spends money: required cap, read before paying, idempotency keys, no secrets in logs, and a cancel path. Drawn from Sume docs.
- Planlock in front of Sume MCP: approve the plan, then the calls
Planlock is an MCP proxy that enforces a human-approved plan. Where it fits in front of Sume's hosted MCP, and which Sume gates still do work behind it.
- Porting chat-completions code to Sume Agent Completions
Agent Completions takes system and user messages, rejects assistant turns, returns a 202 receipt, and does not stream. Here is what to change when porting.
- Scheduled run cost cap: a per-run cap can only lower it
A scheduled Sume run defaults to a $1.00 cap. A per-run cap can only lower the schedule cap, never raise it. Here is how that differs from a direct Format call.
Written by Sume