<!-- canonical: https://blog.scrymore.com/blog/cheaper-design-diffs-cascade/ | published: 2026-09-25 -->

# 10–100× Cheaper Design Diffs: A Judged 60-Screen Scorecard

In the [last two](/blog/beyond-screenshots-figma-storybook-diff-agents) [posts](/blog/cheaper-diff-agents-five-more-models) we gave a review agent everything a coding agent has (the Figma layer tree, a Storybook capture at the exact frame size, the live DOM and the implementation source) and compared models on one dense screen. The best of them were good and expensive: Claude Opus 5.5 at xhigh effort costs about fifty cents a screen. A design system with a few hundred stories can't run that on every change.

This round asks a different question. How cheaply can you get the same result on many screens? The answer comes from a pipeline change, not a model: a zero-LLM pre-filter measures the Figma layers against the Storybook capture, and then one vision call checks its list. On 60 screens scored against judged truth, that cascade finds as many human-tagged defects as Opus 5.5 xhigh at 1/120th of the cost. The two diff tiers we now ship in Scry, Basic and Plus, are built on it.

{{chart:cost-recall-judged60}}
## The result

| Detector | Human-tag recall | All-real recall | Precision | Cost / screen |
| --- | ---: | ---: | ---: | ---: |
| **Scry Plus** (cascade, Opus 5.5 on busy screens) | **94.0%** (78/83) | **72.5%** | 84.5% | $0.0513 |
| Cascade → Claude Opus 5.5 medium | 92.8% (77/83) | 55.6% | 90.5% | $0.0704 |
| GPT-5.6 Terra max | 91.6% (76/83) | 64.6% | 96.9% | n/a |
| **Scry Basic** (cascade, GPT-6 Luna) | **89.2%** (74/83) | 68.8% | 86.7% | **$0.0091** |
| GPT-5.6 Sol xhigh | 87.9% (73/83) | 63.2% | 96.5% | n/a |
| Cascade → GPT-6 Luna medium | 85.5% (71/83) | 62.7% | 85.7% | **$0.0045** |
| Claude Opus 5.5 xhigh + retry guard | 84.3% (70/83) | 64.9% | **98.0%** | $0.5398 |
| Cascade → Muse Spark 1.3 Contributor | 73.5% (61/83) | 36.3% | 91.6% | $0.0023 |
| GPT-6 Luna xhigh | 71.1% (59/83) | 52.5% | 97.1% | $0.0150 |
| GPT-6 Luna medium | 63.9% (53/83) | 45.7% | 97.8% | $0.0073 |
| GLM 5.3 Flash xhigh | 57.8% (48/83) | 36.4% | 94.3% | $0.0098 |

Human-tag recall is the share of the 83 human-tagged defects, still visible in the images the detectors saw, that a detector found. All-real recall is the share of every real difference we know of on those screens: the 83 plus 1,064 more that a blind judge confirmed among everyone's findings. The table is a selection; all 23 detectors, including the ablations, are on the [leaderboard](/leaderboard/). GPT-5.6 ran through the Codex CLI, which reports no dollar cost.

Four things stand out:

- **The cascade matches Opus for a fraction of the price.** GPT-6 Luna medium as the verifier finds 71 of 83 human-tagged defects for $0.0045 a screen. Opus 5.5 xhigh finds 70 for $0.5398, which is 120 times as much.
- **The same model does far better inside the cascade.** GPT-6 Luna medium on its own, with the full evidence bundle, finds 53 of 83 (63.9%). As the verifier over the pre-filter's list it finds 71 (85.5%) and costs less, because it reads a compact candidate list instead of a layer tree and source files.
- **The shipped tiers beat every single model on recall.** Basic, at $0.0091, finds 89.2% of human tags and 68.8% of all real defects; the best single model with a cost, Opus 5.5 xhigh, is at 84.3% and 64.9% for 59 times the price. Plus, at $0.0513, is the best row on both recall measures and still 10 times cheaper than Opus xhigh.
- **The price is precision.** Opus 5.5 xhigh made 15 wrong claims across the 60 screens (98.0% precision). The Luna cascade made 120, about two per screen (85.7%), and the tiers 121 and 152. Single models at the top are 96–98% precise; cascades and tiers are 84–94%.

So "10–100× cheaper" is literal: Plus is 10.5× cheaper than Opus 5.5 xhigh, Basic 59× and the Luna cascade 120×, each with higher human-tag recall and lower precision.

## How the cascade works

The pipeline has two steps.

1. **Deterministic pre-filter.** No model, about one second of CPU. It walks the Figma node tree, finds each element in the Storybook capture by template matching, and emits candidates with pixel boxes: shifted rows, size and gap changes, text differences, fills, and a contact sheet of enlarged icon pairs. On its own it is noisy (46 findings a screen) but it measures things vision models guess at, such as a 4 px row gap.
2. **One verifier call.** One vision model gets the Figma render, the Storybook capture, the icon contact sheet and the candidate list, about 5–7k input tokens. It confirms or rejects each candidate, rewrites it in words a developer can act on, merges duplicates and adds anything the list missed.

The pre-filter supplies the measurements and the model supplies the judgement. That split is where the savings come from. A frontier model at high effort spends most of its tokens finding and measuring differences, and the pre-filter does that part for free.

**Scry Basic** is this cascade with GPT-6 Luna medium as the verifier, merged with a plain GPT-6 Luna pass over the same images. **Scry Plus** adds a second verifier, Claude Opus 5.5 at medium effort, on busy screens: those where the pre-filter reports 49 or more candidates and icons, which is where the single cheap call misses most. On the 60 screens, 31 escalated. Both rows in the table replay the 60 screens through the production Worker code, with the inputs production actually has: the pre-filter is a TypeScript port running on WebAssembly OpenCV, and there is no DOM capture and no component source. An ablation on the benchmark cascade showed that dropping the DOM costs nothing measurable (85.5% human-tag recall with or without it).

What a screen costs depends on how busy it is:

| Runs | Basic | Plus | Plus escalated to Opus |
| --- | ---: | ---: | ---: |
| 60-screen replay | $0.0091 | $0.0513 | 31 of 60 |
| Stage, two busy screens | $0.0089 / $0.0094 | $0.0697 / $0.0919 | 2 of 2 |
| Stage, four small components | $0.0044 | $0.0050 | 0 of 4 |

A small component costs well under a cent on either tier. A busy screen that escalates costs seven to nine cents on Plus.

## How we scored 60 screens

Our earlier 60-screen results matched findings to 140 human tags by label and box overlap. On that metric Opus 5.5 xhigh found 42.9% of the tags, and we spent a round finding out why the number was so low. Sorting its 80 misses by cause:

- **34 were not visible in the images it was given.** The human tags were drawn on the original pair (a real app screenshot and a generated app), and the detectors see a native Figma rebuild of the real app against a Storybook capture of the generated code. The Figma rebuild often copied the generated look, such as a bold "to" or a green "+1", so the difference no longer existed.
- **26 were found but not matched by the scorer.** The box was offset, or one finding covered several repeated tags and could match only one of them.
- **19 were visible and not reported, and 15 of those came from a harness bug.** On 12 of the 60 screens the Opus main pass stopped after one issue (the missing status bar) even though its own summary listed many. It is stochastic, and a retry fixes it. We added a retry guard, which re-runs a collapsed pass up to twice. Only 4 of the 80 misses were subtle differences the model genuinely didn't see.

So we rebuilt the truth set:

1. **Human tags, filtered for visibility.** The tags form 101 groups. A judge checked each against the Figma/Storybook pair: 61 are clearly visible, 22 partly and 18 not at all. The 83 visible or partly visible groups are the human-tag truth.
2. **Every finding, judged.** We pooled all detectors' findings per screen and clustered them. Judges then ruled on every cluster, with detector names and vote counts hidden: it matches a human group, it is a real difference no human tagged, it duplicates another, or it is not real. The 1,064 distinct real untagged differences plus the 83 groups make the 1,147-defect all-real truth. When we add a detector, only its new clusters are judged; this is round 8.
3. **Scoring by verdict, not geometry.** A detector's recall is the share of truth items at least one of its findings resolves to. Precision counts distinct findings judged real over all distinct judged findings, so ten copies of one claim count once.

Under the new scoring the retry guard is worth 8.4 points of human-tag recall to Opus 5.5 xhigh (75.9% without it, 84.3% with it) for three cents more a screen.

## Where the tiers fall short

We set targets for the tiers before building them, and they miss one of them:

| Tier | Human-tag recall | All-real recall | Precision |
| --- | ---: | ---: | ---: |
| Basic | 89.2% (target 85%) | **68.8% (target 72%)** | 86.7% (target 85%) |
| Plus | 94.0% (target 92%) | **72.5% (target 78%)** | 84.5% |

Both clear the human-tag target and both miss all-real recall, Basic by 3.2 points and Plus by 5.5. The misses are among the 1,064 smaller differences nobody tagged. Our first build also had a precision trim that dropped low-confidence findings. It cost 9 to 10 points of all-real recall (Basic 59.4%, Plus 62.9%) and did not even raise precision (85.5% and 83.2%), so we removed it; the trimmed runs are on the leaderboard as ablations. The next lever we plan to try is giving the plain pass the Storybook story source, which the benchmark agents had and production doesn't.

## Where to follow along

The [leaderboard](/leaderboard/) now opens on the 60-screen judged table, with the Scry tiers and a cost chart. It is built from the scoring output, so when we run and judge a new detector the board updates. The 311-pair screenshot benchmark is still in the [paper](/paper/), and the dataset is [Scrymore/scry-design-diff-eval](https://huggingface.co/datasets/Scrymore/scry-design-diff-eval) on Hugging Face.

Most of this round's gain came from engineering the pipeline rather than choosing a model. About 900 lines of OpenCV code do the measuring, and a cheap model is enough to judge the measurements. On the 60-screen set, checking a screen now costs under a cent, at a human-tag recall that took fifty cents with a frontier model alone.
