Lambda get_remaining_time_in_millis: stop polling a Sume job in time

A Lambda that polls a Sume job should check get_remaining_time_in_millis before each sleep and return the job id so a later invocation can resume. Python.

4 min readSume
All posts

When a Lambda function polls a Sume job, read context.get_remaining_time_in_millis() before every sleep, and if the time left is smaller than your next wait plus a safety margin, return the job id instead of sleeping into a timeout. The Lambda context docs describe that method as returning the milliseconds left before the execution times out. A function that is killed mid-poll tells nobody the job id, while a function that returns cleanly can be invoked again with it.

This keeps a polling Lambda inside the platform ceiling without guessing how long a render will take. The timeout page lists a default of 3 seconds and a maximum of 900 seconds, so even the longest setting is a budget, not an unlimited wait. Video jobs can run longer than that, so the function needs a way to hand off.

The loop in plain terms

Sume's GET /v1/jobs/{id}/status returns terminal, result_ready and next_poll_after_seconds, described in the jobs and results docs. The loop reads the status, stops when the job is terminal, waits at least the server's hint, and checks the clock before sleeping. If there is not enough time left, it returns a small object that your scheduler, a Step Functions state or an SQS delay can feed back into the same function.

Lambda timing facts and the loop's decision (read 2026-10-08)
ItemValueSource
Lambda default timeout3 secondsAWS docs
Lambda maximum timeout900 secondsAWS docs
Time left in a Python handlercontext.get_remaining_time_in_millis()AWS docs
Job finishedterminal true, then fetch result when result_ready is trueSume jobs docs
Server wait hintnext_poll_after_seconds, floor of 2 s in this loopSume jobs docs, our choice of floor

Steps

  • Set the function timeout to the longest you are willing to pay for, up to the 900 second maximum, and pick a safety margin for your own return path.
  • Pass the job id in the event. Never create a new job inside the poll function.
  • Return done: false with the job id when time is short, and have your scheduler call the function again later.
  • When result_ready is true, fetch GET /v1/jobs/{id}/result in the same or a later invocation.

Handler

The sample uses the standard library and reads the key from SUME_API_KEY. We ran it with a fake context object against a local stand-in for the API, covering a finished job and a queued job that hit the time limit.

import json
import os
import time
import urllib.request

BASE = os.environ.get("SUME_BASE", "https://api.sume.com")
SAFETY_MS = 15_000  # leave time to return a clean answer


def status(job_id: str) -> dict:
    req = urllib.request.Request(
        f"{BASE}/v1/jobs/{job_id}/status",
        headers={"Authorization": f"Bearer {os.environ['SUME_API_KEY']}"},
    )
    with urllib.request.urlopen(req, timeout=5) as res:
        return json.load(res)["data"]


def handler(event, context):
    job_id = event["job_id"]
    while True:
        doc = status(job_id)
        if doc["terminal"]:
            return {"job_id": job_id, "done": True, "result_ready": doc["result_ready"]}
        wait = max(float(doc.get("next_poll_after_seconds") or 2), 2.0)
        if context.get_remaining_time_in_millis() < SAFETY_MS + wait * 1000:
            return {"job_id": job_id, "done": False, "resume": "invoke again with this job_id"}
        time.sleep(wait)

What Sume does not do

Sume does not know your function's deadline and does not shorten next_poll_after_seconds for it. Stopping the poll does not stop the job; it keeps running and keeps its status, so polling again later with the same job id is safe. If you want push instead of poll, a signed webhook sends the terminal event, and polling stays as the backup.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume