SCRYResearch

Scry / Research ·

10–100× Cheaper Design Diffs: A Judged 60-Screen Scorecard

In the last two posts we gave a review agent everything a coding agent has (the Figma layer tree, a Storybook capture at the exact frame size, the live DOM and the implementation source) and compared models on one dense screen. The best of them were good and expensive: Claude Opus 5.5 at xhigh effort costs about fifty cents a screen. A design system with a few hundred stories can't run that on every change.

This round asks a different question. How cheaply can you get the same result on many screens? The answer comes from a pipeline change, not a model: a zero-LLM pre-filter measures the Figma layers against the Storybook capture, and then one vision call checks its list. On 60 screens scored against judged truth, that cascade finds as many human-tagged defects as Opus 5.5 xhigh at 1/120th of the cost. The two diff tiers we now ship in Scry, Basic and Plus, are built on it.

Cost vs. human-tag recall, 60 screens

83 human-tagged groups visible in the Figma/Storybook pair · blind-judged, round 8 · one run per detector

Cost vs. human-tag recall, 60 screens14 detectors plotted by cost per screen (log scale) against human-tag recall. Frontier: Muse 1.3 cascade 73.5% at $0.0023; Luna cascade 85.5% at $0.0045; Scry Basic 89.2% at $0.0091; Scry Plus 94.0% at $0.051.Muse 1.3 cascadeLuna cascadeGPT-6 Luna mediumScry BasicScry PlusOpus 5.5 cascadeClaude Opus 5.5 xhigh + guard
  • Scry tier (production code)
  • Cascade: pre-filter + one verifier call
  • Single model
  • Pareto frontier
Cost per screen against human-tag recall for 14 detectors on 60 Figma/Storybook screens. Cascades and the Scry tiers sit on the frontier; Claude Opus 5.5 xhigh is far to the right.

The result

DetectorHuman-tag recallAll-real recallPrecisionCost / screen
Scry Plus (cascade, Opus 5.5 on busy screens)94.0% (78/83)72.5%84.5%$0.0513
Cascade → Claude Opus 5.5 medium92.8% (77/83)55.6%90.5%$0.0704
GPT-5.6 Terra max91.6% (76/83)64.6%96.9%n/a
Scry Basic (cascade, GPT-6 Luna)89.2% (74/83)68.8%86.7%$0.0091
GPT-5.6 Sol xhigh87.9% (73/83)63.2%96.5%n/a
Cascade → GPT-6 Luna medium85.5% (71/83)62.7%85.7%$0.0045
Claude Opus 5.5 xhigh + retry guard84.3% (70/83)64.9%98.0%$0.5398
Cascade → Muse Spark 1.3 Contributor73.5% (61/83)36.3%91.6%$0.0023
GPT-6 Luna xhigh71.1% (59/83)52.5%97.1%$0.0150
GPT-6 Luna medium63.9% (53/83)45.7%97.8%$0.0073
GLM 5.3 Flash xhigh57.8% (48/83)36.4%94.3%$0.0098

Human-tag recall is the share of the 83 human-tagged defects, still visible in the images the detectors saw, that a detector found. All-real recall is the share of every real difference we know of on those screens: the 83 plus 1,064 more that a blind judge confirmed among everyone's findings. The table is a selection; all 23 detectors, including the ablations, are on the leaderboard. GPT-5.6 ran through the Codex CLI, which reports no dollar cost.

Four things stand out:

  • The cascade matches Opus for a fraction of the price. GPT-6 Luna medium as the verifier finds 71 of 83 human-tagged defects for $0.0045 a screen. Opus 5.5 xhigh finds 70 for $0.5398, which is 120 times as much.
  • The same model does far better inside the cascade. GPT-6 Luna medium on its own, with the full evidence bundle, finds 53 of 83 (63.9%). As the verifier over the pre-filter's list it finds 71 (85.5%) and costs less, because it reads a compact candidate list instead of a layer tree and source files.
  • The shipped tiers beat every single model on recall. Basic, at $0.0091, finds 89.2% of human tags and 68.8% of all real defects; the best single model with a cost, Opus 5.5 xhigh, is at 84.3% and 64.9% for 59 times the price. Plus, at $0.0513, is the best row on both recall measures and still 10 times cheaper than Opus xhigh.
  • The price is precision. Opus 5.5 xhigh made 15 wrong claims across the 60 screens (98.0% precision). The Luna cascade made 120, about two per screen (85.7%), and the tiers 121 and 152. Single models at the top are 96–98% precise; cascades and tiers are 84–94%.

So "10–100× cheaper" is literal: Plus is 10.5× cheaper than Opus 5.5 xhigh, Basic 59× and the Luna cascade 120×, each with higher human-tag recall and lower precision.

How the cascade works

The pipeline has two steps.

  1. Deterministic pre-filter. No model, about one second of CPU. It walks the Figma node tree, finds each element in the Storybook capture by template matching, and emits candidates with pixel boxes: shifted rows, size and gap changes, text differences, fills, and a contact sheet of enlarged icon pairs. On its own it is noisy (46 findings a screen) but it measures things vision models guess at, such as a 4 px row gap.
  2. One verifier call. One vision model gets the Figma render, the Storybook capture, the icon contact sheet and the candidate list, about 5–7k input tokens. It confirms or rejects each candidate, rewrites it in words a developer can act on, merges duplicates and adds anything the list missed.

The pre-filter supplies the measurements and the model supplies the judgement. That split is where the savings come from. A frontier model at high effort spends most of its tokens finding and measuring differences, and the pre-filter does that part for free.

Scry Basic is this cascade with GPT-6 Luna medium as the verifier, merged with a plain GPT-6 Luna pass over the same images. Scry Plus adds a second verifier, Claude Opus 5.5 at medium effort, on busy screens: those where the pre-filter reports 49 or more candidates and icons, which is where the single cheap call misses most. On the 60 screens, 31 escalated. Both rows in the table replay the 60 screens through the production Worker code, with the inputs production actually has: the pre-filter is a TypeScript port running on WebAssembly OpenCV, and there is no DOM capture and no component source. An ablation on the benchmark cascade showed that dropping the DOM costs nothing measurable (85.5% human-tag recall with or without it).

What a screen costs depends on how busy it is:

RunsBasicPlusPlus escalated to Opus
60-screen replay$0.0091$0.051331 of 60
Stage, two busy screens$0.0089 / $0.0094$0.0697 / $0.09192 of 2
Stage, four small components$0.0044$0.00500 of 4

A small component costs well under a cent on either tier. A busy screen that escalates costs seven to nine cents on Plus.

How we scored 60 screens

Our earlier 60-screen results matched findings to 140 human tags by label and box overlap. On that metric Opus 5.5 xhigh found 42.9% of the tags, and we spent a round finding out why the number was so low. Sorting its 80 misses by cause:

  • 34 were not visible in the images it was given. The human tags were drawn on the original pair (a real app screenshot and a generated app), and the detectors see a native Figma rebuild of the real app against a Storybook capture of the generated code. The Figma rebuild often copied the generated look, such as a bold "to" or a green "+1", so the difference no longer existed.
  • 26 were found but not matched by the scorer. The box was offset, or one finding covered several repeated tags and could match only one of them.
  • 19 were visible and not reported, and 15 of those came from a harness bug. On 12 of the 60 screens the Opus main pass stopped after one issue (the missing status bar) even though its own summary listed many. It is stochastic, and a retry fixes it. We added a retry guard, which re-runs a collapsed pass up to twice. Only 4 of the 80 misses were subtle differences the model genuinely didn't see.

So we rebuilt the truth set:

  1. Human tags, filtered for visibility. The tags form 101 groups. A judge checked each against the Figma/Storybook pair: 61 are clearly visible, 22 partly and 18 not at all. The 83 visible or partly visible groups are the human-tag truth.
  2. Every finding, judged. We pooled all detectors' findings per screen and clustered them. Judges then ruled on every cluster, with detector names and vote counts hidden: it matches a human group, it is a real difference no human tagged, it duplicates another, or it is not real. The 1,064 distinct real untagged differences plus the 83 groups make the 1,147-defect all-real truth. When we add a detector, only its new clusters are judged; this is round 8.
  3. Scoring by verdict, not geometry. A detector's recall is the share of truth items at least one of its findings resolves to. Precision counts distinct findings judged real over all distinct judged findings, so ten copies of one claim count once.

Under the new scoring the retry guard is worth 8.4 points of human-tag recall to Opus 5.5 xhigh (75.9% without it, 84.3% with it) for three cents more a screen.

Where the tiers fall short

We set targets for the tiers before building them, and they miss one of them:

TierHuman-tag recallAll-real recallPrecision
Basic89.2% (target 85%)68.8% (target 72%)86.7% (target 85%)
Plus94.0% (target 92%)72.5% (target 78%)84.5%

Both clear the human-tag target and both miss all-real recall, Basic by 3.2 points and Plus by 5.5. The misses are among the 1,064 smaller differences nobody tagged. Our first build also had a precision trim that dropped low-confidence findings. It cost 9 to 10 points of all-real recall (Basic 59.4%, Plus 62.9%) and did not even raise precision (85.5% and 83.2%), so we removed it; the trimmed runs are on the leaderboard as ablations. The next lever we plan to try is giving the plain pass the Storybook story source, which the benchmark agents had and production doesn't.

Where to follow along

The leaderboard now opens on the 60-screen judged table, with the Scry tiers and a cost chart. It is built from the scoring output, so when we run and judge a new detector the board updates. The 311-pair screenshot benchmark is still in the paper, and the dataset is Scrymore/scry-design-diff-eval on Hugging Face.

Most of this round's gain came from engineering the pipeline rather than choosing a model. About 900 lines of OpenCV code do the measuring, and a cheap model is enough to judge the measurements. On the 60-screen set, checking a screen now costs under a cent, at a human-tag recall that took fifty cents with a frontier model alone.