Article 50(2) assistive editing: trim and crop vs generate
Article 50(2) exempts assistive editing that does not substantially alter input. Sort your Sume outputs into generated and file-processed before you rely on it.

Article 50(2) of the EU AI Act carves out systems that perform an assistive function for standard editing, or that do not substantially alter the input data they receive. A plain trim or crop of footage you already own sits much nearer that exception than a new clip a model generated from a prompt. Whether a particular edit qualifies is a legal reading, so the practical job for a video team is to record, per output, whether a generative model produced the pixels or a file tool only rearranged them.
What the article says about the exception
The text of Article 50, read on 2026-10-03, requires providers to mark synthetic outputs in a machine-readable format, detectable as artificially generated. It asks for solutions that are effective, interoperable, robust and reliable as far as technically feasible. The obligation does not apply to systems that perform an assistive function for standard editing, or that do not substantially alter the input data.
Two things follow. The exception is about what the system does to its input, not about which tool brand you used. And the word substantially leaves room for judgment, which is why a ledger of what each step did is more useful than a guess made after the fact.
Sorting Sume operations by what they do to the pixels
The Sume docs describe each media operation, and the split is visible in them. Generation goes through a model. The media tools run ffmpeg on a worker. The table sorts the documented operations by that fact only. It does not say which ones the Act exempts, because the docs make no claim about that and this post does not either.
| Operation | What the docs say it does | Model inference |
|---|---|---|
| POST /v1/videos | Generates video from a prompt, optional first or last frame or references | Yes |
| POST /v1/video-trim | Cuts a [start, end) range of one clip into a new MP4; exact mode re-encodes | No, worker ffmpeg only |
| POST /v1/timeline-1.0/render | Assembles one audio spine plus ordered video slots into one MP4 | No, worker ffmpeg only |
| POST /v1/video-inspect | Probes the clip and samples stills; never re-encodes the source | No MP4 produced |
Where the line is blurry
Some edits sit in the middle. A reframe from 16:9 to 9:16 by cropping discards most of the original frame, and a generated first frame that is then animated sends pixels through a model. A Timeline render that stitches three generated clips contains no new inference at assembly time, yet every slot in it is generated content.
So classify by the whole chain. If any slot came from a generation call, the output is not an untouched copy of your own footage, whatever the final assembly step was. Keep the classification at the level of the source assets, then derive the output status from them.
A small ledger you can keep per output
Store the operation, the source asset, and whether a generation call is in its ancestry. The record below is plain Python and runs as written. It marks an output as generated when any ancestor was.
from dataclasses import dataclass, field
@dataclass
class Asset:
name: str
operation: str # e.g. generate, trim, timeline, upload
parents: list = field(default_factory=list)
def has_generation(asset):
if asset.operation == "generate":
return True
return any(has_generation(p) for p in asset.parents)
raw = Asset("customer-footage.mp4", "upload")
clip = Asset("shot-1.mp4", "generate")
cut = Asset("cut.mp4", "trim", [raw])
final = Asset("final.mp4", "timeline", [cut, clip])
for a in (cut, final):
print(a.name, "generated ancestry:", has_generation(a))What to do with it
Outputs with generated ancestry go through your marking and disclosure workflow without further debate. Outputs with only upload and file-processing ancestry are the ones to take to counsel if you want to rely on the exception. The ledger makes that conversation short, because each case is already described by what happened to the pixels.
For the provider side of the Act, see the machine-readable marking post, and for the split between your duties and your vendor's, the two-column breakdown.
Sources
Related posts
More in Use cases
- Clean a transcript with a 2,000-character instruction, then caption
ElevenLabs speech to text can edit a transcript from a natural-language instruction up to 2,000 characters. A cleanup then captions workflow with Sume captions.
- Demand Gen carousels: 2 to 10 matching cards from one reference
Demand Gen carousels take 2 to 10 cards. Image assets run 4:5 or 9:16 at 5 MB. How to batch matching cards from one reference image with the Sume image API.
- Demand Gen logo 150 KB, images 5 MB: check sizes before upload
Demand Gen allows images up to 5 MB but logos only 150 KB. Check bytes on every Sume image before upload, since output compression is not served in v1.
- Demand Gen video ads: four ratios, a 5 s floor and the 4:5 gap
Demand Gen video takes 1:1, 16:9, 4:5 and 9:16, from 5 s. How to render the four-ratio set with Sume, including the 4:5 case the video API does not list.
Written by Sume