Those screens belong to a B2B supply-chain SaaS, and they’re invoices, purchase orders and inventory, mostly dense tables and dialogs. In July I started a system that shoots each one at six widths, once on a merge request’s branch and once on main. One change to a single stylesheet widened a table column, and its hand-picked list named five screens.
| Scope | Screens shot | Changed screens among them |
|---|---|---|
| Hand-picked | 5 | 4 |
| From the diff | 14 | 13 |
The hand-picked run never looked at the other nine, and nothing in its output said so.
It began as a script that shot routes at six widths, to check whether I’d broken a layout. Comparing two deployed environments varied the code and the data at once, so detail pages diffed as noise, and in CI the database was empty. Nobody trusted shots that noisy or blank.
A build diffed against itself
The baseline is main, shot against the same data as the branch, so nobody approves it. Approved baselines rot on a UI that changes with every MR. Two pipeline runs of the seed differed on 103 of 155 routes, so one job seeds the app, mostly through its own APIs, and dumps the database for both sides. Each side refuses to shoot unless its row counts hash to the seed’s. The clock is pinned too, because the app shows relative dates.
Diffing a build against itself measures the noise that’s left, since every difference it finds is fake. In August that flagged 40 of the 318 screens under test, enough fake regressions to bury any real one. Webfonts loaded on one pass and fell back on the next, wrapping a label onto two lines, and now they’re blocked. Focus is cleared so no focus ring shows, and the run lets charts settle and waits out lists still adding rows until the element count holds steady twice. After that, 317 of 318 came back clean.
States a URL can’t reach
Dialogs, menus and rows in edit mode have no URL. Each became a scenario that clicks its way there. A backlog of 401 uncaptured states, dialogs the largest group, closed with none open. Of those, 333 were captured, 19 were retired or blocked, and 49 are covered by a shot of the same component, like the one shared dialog behind most “Are you sure” prompts.
Let the diff pick the screens
A full run of every screen against a live backend takes 57 min. A targeted run takes its scope from the diff, because a hand-picked list is the reviewers’ guess again. Each changed file maps to its component, and a global stylesheet to every screen. An index of what every screen renders picked 14 of 485 for the column change, a check that took 3 min 25 s. Both sides replayed recorded API traffic in place of a backend, and a replay matched a live run byte for byte on 131 of 131 shots.
A gallery that reads like the app
Screens in the gallery follow the app’s menu, as nesting them by URL had split screens the menu keeps together. Holding D lays main over the branch shot, a blink comparison for spotting layout shifts by eye. The H key paints tiles where a screen changed, since single changed pixels read as static; against a Figma frame, where a raw pixel count says little, Parity Bench names what’s off.
The change badge once measured the 1,440-pixel width only, and one screen read 0.0% there and 45% at 375 pixels. It now weighs the six widths equally, and its tooltip says which width broke.
Where it stands
Any MR can trigger it from its pipeline, and it advises without gating anything. I don’t count the regressions it catches per month, so its value is still a feeling. On my own MRs, a local worker runs it after each push, and QAnelita reads the result.