Parity Bench

2026 · Tooling

A judge for AI-built UI that compares components with their Figma frames and names what's off. One pipeline that looked 3× worse was within 1.4%.

Node, Playwright, Python

A pixel diff of one component against its Figma frame came back as 12,377, a number nobody could act on. I built Parity Bench for a legal-tech frontend team. This is one of its findings, as the tool reported it.

[high] column-gap  gutter 1 (measured from the image)
       reference 21px | build 36px | +15px vs the reference

The Figma frame had a 20-pixel gap between two columns, and the build had 30. A gap that moves 10 pixels drags a whole column of text sideways. Both readings in the finding come off the image, which is why they run wide.

The team had started letting AI turn Figma designs into components. Two agent skills, one that built components from Figma and one that audited them, ran capture, measurement, pixel diffing and even tolerance checks through an LLM. The tolerances and severity rules they applied were already fixed and written down.

Align before judging

Shift a card by a pixel and the whole comparison lights up red. Overlaid, one Figma export and its build showed every sentence twice, about 24 pixels apart, so the tool searches offset and scale, coarse to fine, before it diffs anything.

The build side is a Playwright screenshot of the component’s Storybook story, with the component pinned to a set width.

Pair Score (6×6 cells that differ) What it measured After
A full-bleed 1,408 px build, a 1,180 px frame 12,377 the resize 1,648 with the build pinned to 1,200 px
A @2x export, a 1x capture 4,196 the export’s scale 0 with @2x detection

Findings instead of a score

A gap survives a width mismatch, since 20 pixels is 20 pixels whether the row is 1,170 or 1,408 wide. Parity Bench pairs the component’s children by position, because Figma frame names and CSS class names have no reason to agree, and checks every gap between neighbors.

The tolerances allow 1 pixel either way on spacing and nothing on font size, line height, weight or color. Segoe UI rendering as San Francisco on a Mac doesn’t count. The tool applies these rules in code, with no AI in the loop. The pixel score now runs only from the command line, because every pixel number in the interface needed a paragraph of caveat that a reader couldn’t act on. My visual-regression system still runs on pixel diffs, with builds on both sides.

Figma’s API answered 429 for days

The API was rate-limited and plugins weren’t allowed, which left the clipboard. Copy a node in Figma and the clipboard’s HTML carries it as a binary blob in Evan Wallace’s Kiwi format, and I wrote a decoder for it in standard-library Python. The blob carries its own schema, so the decoder hardcodes nothing about Figma’s data model and survives a new schema version. A pasted node gives exact font sizes, line heights, colors and weights, and its colors get matched back to the theme’s design tokens.

Leave the placeholders out

Parity Bench also judged between two AI build pipelines, each building the same component once. One build scored 6,174 mismatched cells against the other’s 1,936. With one placeholder image left out, the two were within 1.4% of each other. My first blind eval of QAnelita, my QA agent, leaked the answer into the evidence, so its score couldn’t be trusted either.

Parity Bench is internal. Calibrated against the build’s DOM, the finding at the top now reads about 21 pixels against 30, where the truth is 20.