SCRYResearch

Scry Design Diff Eval

Design Diff Leaderboard

How well do AI agents catch the differences between a Figma design and the code that implements it, and what does that cost per screen? The main table runs every detector on the same 60 Figma/Storybook screens and scores it against a judged truth set; the other tab is the 311-pair screenshot benchmark. Updated whenever we run a new model.

Top human-tag recall94.0%Scry Plus tier, $0.0513/screen
Best single model with a cost84.3%Claude Opus 5.5 xhigh + retry guard, $0.540/screen
Same recall, cheapest$0.0045Cascade → GPT-6 Luna, 120× cheaper
Judged truth1,147real defects on 60 screens, 83 human-tagged
Cost vs human-tag recall, 60 screensScatter of 14 detectors with a reported USD cost: cost per screen on a log axis against recall of the 83 human-tagged defect groups still visible in the Figma/Storybook pair. The stepped line is the Pareto frontier: Cascade → Muse Spark 1.3 Contributor medium at $0.0023 and 73.5%, then Cascade → GPT-6 Luna medium at $0.0045 and 85.5%, then Scry Basic tier at $0.0091 and 89.2%, then Scry Plus tier at $0.0513 and 94.0%.
Cost vs human-tag recall, 60 screensScatter of 14 detectors with a reported USD cost: cost per screen on a log axis against recall of the 83 human-tagged defect groups still visible in the Figma/Storybook pair. The stepped line is the Pareto frontier: Cascade → Muse Spark 1.3 Contributor medium at $0.0023 and 73.5%, then Cascade → GPT-6 Luna medium at $0.0045 and 85.5%, then Scry Basic tier at $0.0091 and 89.2%, then Scry Plus tier at $0.0513 and 94.0%.
  1. Cascade → Muse Spark 1.3 Contributor medium $0.0023 · 73.5%
  2. Cascade → GPT-6 Luna medium $0.0045 · 85.5%
  3. Scry Basic tier $0.0091 · 89.2%
  4. Scry Plus tier $0.0513 · 94.0%

Frontier (filled dots on the stepped line): no other run is both cheaper and better.

Other runs (hollow dots). Hover, tap or tab to a dot for its name.

Not plotted (no USD cost: Codex CLI and Claude Code subagent runs): GPT-5.6 Sol xhigh; GPT-5.6 Terra (max, xhigh); Cascade → Claude Opus (Claude Code subagent) default.

Cost vs human-tag recall, 60 screens: all plotted runs, cheapest first
RunCost per screenHuman-tag recallOn frontier
Cascade → Muse Spark 1.3 Contributor medium$0.002373.5%yes
Cascade → GLM Flash medium$0.002563.9%no
Qwen 3.7 Flash xhigh$0.002643.4%no
Cascade → GPT-6 Luna medium$0.004585.5%yes
GPT-6 Luna medium$0.007363.9%no
Scry Basic tier$0.009189.2%yes
GLM 5.3 Flash xhigh$0.009857.8%no
Cascade → GPT-6 Luna Pro medium$0.011284.3%no
GPT-6 Luna Pro medium$0.015063.9%no
GPT-6 Luna xhigh$0.015071.1%no
Gemini 3.8 Flash low$0.036539.8%no
Scry Plus tier$0.051394.0%yes
Cascade → Claude Opus 5.5 medium$0.070492.8%no
Claude Opus 5.5 xhigh + retry guard$0.54084.3%no
Diff detectors on 60 Figma/Storybook screens, scored against judged truth
#
1Scry Plus tiershippedfrontierScry tier94.0%78/8372.5%831/114784.5%78.0%$0.0513
2Cascade → Claude Opus (Claude Code subagent)defaultcascade92.8%77/8363.9%733/114794.2%76.2%n/a
3Cascade → Claude Opus 5.5mediumcascade92.8%77/8355.6%638/114790.5%68.9%$0.0704
4Scry Plus tier, precision trimablationretiredablation92.8%77/8362.9%721/114783.2%71.6%$0.0513
5GPT-5.6 Terramaxsingle model91.6%76/8364.6%741/114796.9%77.5%n/a
6Scry Basic tiershippedfrontierScry tier89.2%74/8368.8%789/114786.7%76.7%$0.0091
7GPT-5.6 Solxhighsingle model87.9%73/8363.2%725/114796.5%76.4%n/a
8Scry Basic tier, precision trimablationretiredablation87.9%73/8359.4%681/114785.5%70.1%$0.0091
9Cascade → GPT-6 Lunafrontiermediumcascade85.5%71/8362.7%719/114785.7%72.4%$0.0045
10Cascade → GPT-6 Luna, no DOMablationmediumablation85.5%71/8362.3%714/114783.9%71.5%$0.0046
11Claude Opus 5.5xhigh + retry guardsingle model84.3%70/8364.9%744/114798.0%78.1%$0.540
12Cascade → GPT-6 Luna Promediumcascade84.3%70/8362.5%717/114785.5%72.2%$0.0112
13Cascade → GPT-6 Luna, no DOM, TS pre-filterablationmediumablation84.3%70/8362.4%716/114784.5%71.8%$0.0046
14GPT-5.6 Terraxhighsingle model83.1%69/8357.8%663/114796.5%72.3%n/a
15Claude Opus 5.5ablationxhigh, no guardablation75.9%63/8356.1%643/114798.0%71.3%$0.508
16Cascade → Muse Spark 1.3 Contributorfrontiermediumcascade73.5%61/8336.3%416/114791.6%52.0%$0.0023
17GPT-6 Lunaxhighsingle model71.1%59/8352.5%602/114797.1%68.1%$0.0150
18GPT-6 Lunamediumsingle model63.9%53/8345.7%524/114797.8%62.3%$0.0073
19GPT-6 Luna Promediumsingle model63.9%53/8347.3%543/114798.2%63.9%$0.0150
20Cascade → GLM Flashmediumcascade63.9%53/8335.1%403/114787.0%50.1%$0.0025
21GLM 5.3 Flashxhighsingle model57.8%48/8336.4%417/114794.3%52.5%$0.0098
22Qwen 3.7 Flashxhighsingle model43.4%36/8321.6%248/114772.1%33.3%$0.0026
23Gemini 3.8 Flashlowsingle model39.8%33/8324.8%284/114796.6%39.4%$0.0365

Click a column to sort. Human-tag recall counts the 83 human-tagged defect groups that are still visible in the image pair the detectors saw; all-real recall adds every other real difference a blind judge confirmed among the pooled findings (1147 in all). A cascade runs a zero-LLM pre-filter and then one vision call that confirms, rewrites or adds findings. “Shipped” rows are the Scry Basic and Plus tiers replayed through the production Worker code. Ablations are kept for reference and left out of the chart and frontier. GPT-5.6 runs went through the Codex CLI and the subagent cascade through Claude Code, so neither has a USD cost.

Data generated from reports/eval60-rescore/scores-r8.json, docs/study_results.md.

Scry Basic and Plus

The two diff tiers in Scry are cascades: a zero-LLM pre-filter measures the Figma layer tree against the Storybook capture, then a GPT-6 Luna call confirms, rewrites or adds findings, merged with a plain GPT-6 Luna pass. Plus also sends busy screens to a Claude Opus 5.5 verifier. Both rows below replay the 60 screens through the production Worker code, with the inputs production has (no DOM capture, no component source).

TierCost / screen (60-screen replay)Human-tag recallAll-real recallPrecision
Basic$0.009189.2% target 85%68.8% target 72%86.7% target 85%
Plus$0.051394.0% target 92%72.5% target 78%84.5%

Both tiers clear their human-tag targets and miss their all-real targets: Basic by 3.2 points, Plus by 5.5 points. What a screen costs depends on how busy it is:

RunsBasicPlusPlus escalated to Opus
eval60 replay (2026-09-23, 60 screens)$0.0091$0.051331/60 (52%)
Stage, busy eval screens (busy 49/64)$0.0089 / $0.0094$0.0697 / $0.09192/2
Stage, busy 47–$0.00780/1
Stage, fresh sample-storybook components (small)$0.0044$0.00500/4

Stage rows are single runs of the same code on our stage environment. Fresh components have no human tags, so they show cost, not quality.

Methodology

The 60 screens
Real-app screens rebuilt natively in Figma and implemented as React Native code captured in Storybook, 15 each with no, one, several and many human-tagged issues. Each detector gets the Figma render and layer tree and the Storybook capture. The benchmark runs also get the Storybook DOM (the single-model agents the component source as well); the Scry tier replays get neither, as in production.
Judged truth
Human tags were drawn on the original screenshot pair, not on the Figma and Storybook frames, and 18 of the 101 tagged groups are not visible in that pair (the Figma rebuild copied the generated look), so they are dropped. Every detector's findings are pooled per screen, and blind judges, with detector names hidden, rule each pooled finding a match to a human group, a real untagged difference, a duplicate or not real. The judged truth is the 83 visible groups plus the 1,064 real untagged differences. New detectors are judged in delta rounds (round 8 is current).
Recall and precision
Human-tag recall is the share of visible human groups a detector found. All-real recall is the share of the whole judged truth. Precision is distinct findings judged real over all distinct judged findings; uncertain calls are left out. All-real recall is relative to the pool: a difference no detector reported is not in the truth set.
Why not tag-and-box matching
Earlier versions of this board scored the 60 screens by matching labels and boxes against 140 human tags. That metric was retired: about a quarter of the tags are not visible in the pair the agents see, and a correct finding with an offset box, or one finding covering several repeated tags, scored as a miss.
Cost and frontier
Cost is the provider-reported USD per screen, all passes included (OpenRouter, or the Worker's metering for the Scry tiers). A detector is on the frontier when no other costed detector is both at least as cheap and at least as good, and strictly better on one. The chart's stepped line and the table's “frontier” badge come from the same computation in the data build.

Write-ups and data