Mastodon media focal point: x and y from -1 to 1, 16:9 crops

Mastodon preview images are never cropped by the server, so apps crop them using a focal point from -1.0 to 1.0. How to compute it from a pixel in a Sume image.

4 min readSume
All posts

Mastodon's server never crops preview images, so each app crops them, and a focal point tells the app which part must stay in view. The point is a pair of numbers from -1.0 to 1.0, left to right and bottom to top, with (0, 0) at the centre. Mastodon says its front-end thumbnails are most commonly 16:9.

How the numbers map

From Mastodon's documentation page, last updated November 23, 2025. The page defers to a separate guide for the full definition, which I did not read, so the table covers only the description on this page.

Focal point to pixel, on a 1536x1024 image, computed (read 2026-10-04)
Focal point (x, y)Pixel (left, top)Meaning
(0, 0)(768, 512)centre
(0.5, 0.5)(1152, 256)centre of the upper-right quadrant
(-0.5, -0.5)(384, 768)centre of the lower-left quadrant
(-1, 1)(0, 0)top-left corner

The formula

A pixel (px, py) in a W by H image maps to x = px / W x 2 - 1 and y = 1 - py / H x 2. The y value flips because pixel rows count downward while the focal axis counts upward.

import asyncio

def focal(px, py, width, height):
    return {"x": round(px / width * 2 - 1, 3), "y": round(1 - py / height * 2, 3)}

async def main():
    # subject centre at pixel (1152, 256) in a 1536x1024 Sume image
    print(focal(1152, 256, 1536, 1024))  # {'x': 0.5, 'y': 0.5}
    print(focal(768, 512, 1536, 1024))   # {'x': 0.0, 'y': 0.0}

asyncio.run(main())

Why it matters for a 3:2 image

Sume can generate 1536x1024 at 3:2 (both edges multiples of 16, 1,572,864 pixels) for a GPT image. A 16:9 crop at full width is 1536x864, which leaves 160 pixels of vertical slack. If a client centres the 864-pixel band on the focal point and clamps it to the image, a subject at pixel row 256 puts the band at the very top, and a subject at row 512 puts it 80 pixels down. That clamping rule is my assumption for illustration, since Mastodon leaves cropping to each app.

Practice

  • Find the subject's pixel centre, from a face box or by eye.
  • Convert it with the formula and keep three decimals at most.
  • Do not rely on one client's crop. The doc says apps do the cropping.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume