Scry Design Diff Eval
Design Diff Leaderboard
How well do AI agents catch the differences between a Figma design and the code that implements it, and what does that cost per screen? The main table runs every detector on the same 60 Figma/Storybook screens and scores it against a judged truth set; the other tab is the 311-pair screenshot benchmark. Updated whenever we run a new model.
- Cascade → Muse Spark 1.3 Contributor medium $0.0023 · 73.5%
- Cascade → GPT-6 Luna medium $0.0045 · 85.5%
- Scry Basic tier $0.0091 · 89.2%
- Scry Plus tier $0.0513 · 94.0%
Frontier (filled dots on the stepped line): no other run is both cheaper and better.
Other runs (hollow dots). Hover, tap or tab to a dot for its name.
Not plotted (no USD cost: Codex CLI and Claude Code subagent runs): GPT-5.6 Sol xhigh; GPT-5.6 Terra (max, xhigh); Cascade → Claude Opus (Claude Code subagent) default.
| Run | Cost per screen | Human-tag recall | On frontier |
|---|---|---|---|
| Cascade → Muse Spark 1.3 Contributor medium | $0.0023 | 73.5% | yes |
| Cascade → GLM Flash medium | $0.0025 | 63.9% | no |
| Qwen 3.7 Flash xhigh | $0.0026 | 43.4% | no |
| Cascade → GPT-6 Luna medium | $0.0045 | 85.5% | yes |
| GPT-6 Luna medium | $0.0073 | 63.9% | no |
| Scry Basic tier | $0.0091 | 89.2% | yes |
| GLM 5.3 Flash xhigh | $0.0098 | 57.8% | no |
| Cascade → GPT-6 Luna Pro medium | $0.0112 | 84.3% | no |
| GPT-6 Luna Pro medium | $0.0150 | 63.9% | no |
| GPT-6 Luna xhigh | $0.0150 | 71.1% | no |
| Gemini 3.8 Flash low | $0.0365 | 39.8% | no |
| Scry Plus tier | $0.0513 | 94.0% | yes |
| Cascade → Claude Opus 5.5 medium | $0.0704 | 92.8% | no |
| Claude Opus 5.5 xhigh + retry guard | $0.540 | 84.3% | no |
| # | ||||||||
|---|---|---|---|---|---|---|---|---|
| 1 | Scry Plus tiershippedfrontier | Scry tier | 94.0%78/83 | 72.5%831/1147 | 84.5% | 78.0% | $0.0513 | |
| 2 | Cascade → Claude Opus (Claude Code subagent) | default | cascade | 92.8%77/83 | 63.9%733/1147 | 94.2% | 76.2% | n/a |
| 3 | Cascade → Claude Opus 5.5 | medium | cascade | 92.8%77/83 | 55.6%638/1147 | 90.5% | 68.9% | $0.0704 |
| 4 | Scry Plus tier, precision trimablation | retired | ablation | 92.8%77/83 | 62.9%721/1147 | 83.2% | 71.6% | $0.0513 |
| 5 | GPT-5.6 Terra | max | single model | 91.6%76/83 | 64.6%741/1147 | 96.9% | 77.5% | n/a |
| 6 | Scry Basic tiershippedfrontier | Scry tier | 89.2%74/83 | 68.8%789/1147 | 86.7% | 76.7% | $0.0091 | |
| 7 | GPT-5.6 Sol | xhigh | single model | 87.9%73/83 | 63.2%725/1147 | 96.5% | 76.4% | n/a |
| 8 | Scry Basic tier, precision trimablation | retired | ablation | 87.9%73/83 | 59.4%681/1147 | 85.5% | 70.1% | $0.0091 |
| 9 | Cascade → GPT-6 Lunafrontier | medium | cascade | 85.5%71/83 | 62.7%719/1147 | 85.7% | 72.4% | $0.0045 |
| 10 | Cascade → GPT-6 Luna, no DOMablation | medium | ablation | 85.5%71/83 | 62.3%714/1147 | 83.9% | 71.5% | $0.0046 |
| 11 | Claude Opus 5.5 | xhigh + retry guard | single model | 84.3%70/83 | 64.9%744/1147 | 98.0% | 78.1% | $0.540 |
| 12 | Cascade → GPT-6 Luna Pro | medium | cascade | 84.3%70/83 | 62.5%717/1147 | 85.5% | 72.2% | $0.0112 |
| 13 | Cascade → GPT-6 Luna, no DOM, TS pre-filterablation | medium | ablation | 84.3%70/83 | 62.4%716/1147 | 84.5% | 71.8% | $0.0046 |
| 14 | GPT-5.6 Terra | xhigh | single model | 83.1%69/83 | 57.8%663/1147 | 96.5% | 72.3% | n/a |
| 15 | Claude Opus 5.5ablation | xhigh, no guard | ablation | 75.9%63/83 | 56.1%643/1147 | 98.0% | 71.3% | $0.508 |
| 16 | Cascade → Muse Spark 1.3 Contributorfrontier | medium | cascade | 73.5%61/83 | 36.3%416/1147 | 91.6% | 52.0% | $0.0023 |
| 17 | GPT-6 Luna | xhigh | single model | 71.1%59/83 | 52.5%602/1147 | 97.1% | 68.1% | $0.0150 |
| 18 | GPT-6 Luna | medium | single model | 63.9%53/83 | 45.7%524/1147 | 97.8% | 62.3% | $0.0073 |
| 19 | GPT-6 Luna Pro | medium | single model | 63.9%53/83 | 47.3%543/1147 | 98.2% | 63.9% | $0.0150 |
| 20 | Cascade → GLM Flash | medium | cascade | 63.9%53/83 | 35.1%403/1147 | 87.0% | 50.1% | $0.0025 |
| 21 | GLM 5.3 Flash | xhigh | single model | 57.8%48/83 | 36.4%417/1147 | 94.3% | 52.5% | $0.0098 |
| 22 | Qwen 3.7 Flash | xhigh | single model | 43.4%36/83 | 21.6%248/1147 | 72.1% | 33.3% | $0.0026 |
| 23 | Gemini 3.8 Flash | low | single model | 39.8%33/83 | 24.8%284/1147 | 96.6% | 39.4% | $0.0365 |
Click a column to sort. Human-tag recall counts the 83 human-tagged defect groups that are still visible in the image pair the detectors saw; all-real recall adds every other real difference a blind judge confirmed among the pooled findings (1147 in all). A cascade runs a zero-LLM pre-filter and then one vision call that confirms, rewrites or adds findings. “Shipped” rows are the Scry Basic and Plus tiers replayed through the production Worker code. Ablations are kept for reference and left out of the chart and frontier. GPT-5.6 runs went through the Codex CLI and the subagent cascade through Claude Code, so neither has a USD cost.
| # | ||||||
|---|---|---|---|---|---|---|
| 1 | oracle-notescontrol | 100.0% | 100.0% | 0.0% | 100.0% | 100.0% |
| 2 | claude-fable-5 xhigh | 40.4% | 67.1% | 92.2% | 52.4% | 16.0% |
| 3 | moonshotai/kimi-k2.7-code + Together recovery | 38.2% | 65.8% | 100.0% | 49.5% | 15.9% |
| 4 | gemini-3.5-flash | 37.5% | 62.4% | 96.1% | 47.9% | 20.2% |
| 5 | codex-cli/gpt-5.5 xhigh | 37.3% | 64.5% | 98.7% | 48.9% | 13.6% |
| 6 | minimax/minimax-m3 | 21.9% | 40.6% | 93.5% | 32.2% | 11.5% |
| 7 | google/gemma-4-26b-a4b-it | 20.3% | 35.9% | 85.7% | 30.5% | 12.4% |
| 8 | google/gemma-4-31b-it | 17.8% | 32.5% | 79.2% | 29.6% | 13.3% |
| 9 | always-no-issuecontrol | 0.0% | 0.0% | 0.0% | 24.8% | 0.0% |
| 10 | always-generic-issuecontrol | 0.0% | 0.0% | 100.0% | 0.0% | 0.0% |
Screenshots only, strict tag-and-box matching over 557 tagged issues. Controls are pipeline checks, not models.
Data generated from reports/eval60-rescore/scores-r8.json, docs/study_results.md.
Scry Basic and Plus
The two diff tiers in Scry are cascades: a zero-LLM pre-filter measures the Figma layer tree against the Storybook capture, then a GPT-6 Luna call confirms, rewrites or adds findings, merged with a plain GPT-6 Luna pass. Plus also sends busy screens to a Claude Opus 5.5 verifier. Both rows below replay the 60 screens through the production Worker code, with the inputs production has (no DOM capture, no component source).
| Tier | Cost / screen (60-screen replay) | Human-tag recall | All-real recall | Precision |
|---|---|---|---|---|
| Basic | $0.0091 | 89.2% target 85% | 68.8% target 72% | 86.7% target 85% |
| Plus | $0.0513 | 94.0% target 92% | 72.5% target 78% | 84.5% |
Both tiers clear their human-tag targets and miss their all-real targets: Basic by 3.2 points, Plus by 5.5 points. What a screen costs depends on how busy it is:
| Runs | Basic | Plus | Plus escalated to Opus |
|---|---|---|---|
| eval60 replay (2026-09-23, 60 screens) | $0.0091 | $0.0513 | 31/60 (52%) |
| Stage, busy eval screens (busy 49/64) | $0.0089 / $0.0094 | $0.0697 / $0.0919 | 2/2 |
| Stage, busy 47 | – | $0.0078 | 0/1 |
| Stage, fresh sample-storybook components (small) | $0.0044 | $0.0050 | 0/4 |
Stage rows are single runs of the same code on our stage environment. Fresh components have no human tags, so they show cost, not quality.
Methodology
- The 60 screens
- Real-app screens rebuilt natively in Figma and implemented as React Native code captured in Storybook, 15 each with no, one, several and many human-tagged issues. Each detector gets the Figma render and layer tree and the Storybook capture. The benchmark runs also get the Storybook DOM (the single-model agents the component source as well); the Scry tier replays get neither, as in production.
- Judged truth
- Human tags were drawn on the original screenshot pair, not on the Figma and Storybook frames, and 18 of the 101 tagged groups are not visible in that pair (the Figma rebuild copied the generated look), so they are dropped. Every detector's findings are pooled per screen, and blind judges, with detector names hidden, rule each pooled finding a match to a human group, a real untagged difference, a duplicate or not real. The judged truth is the 83 visible groups plus the 1,064 real untagged differences. New detectors are judged in delta rounds (round 8 is current).
- Recall and precision
- Human-tag recall is the share of visible human groups a detector found. All-real recall is the share of the whole judged truth. Precision is distinct findings judged real over all distinct judged findings; uncertain calls are left out. All-real recall is relative to the pool: a difference no detector reported is not in the truth set.
- Why not tag-and-box matching
- Earlier versions of this board scored the 60 screens by matching labels and boxes against 140 human tags. That metric was retired: about a quarter of the tags are not visible in the pair the agents see, and a correct finding with an offset box, or one finding covering several repeated tags, scored as a miss.
- Cost and frontier
- Cost is the provider-reported USD per screen, all passes included (OpenRouter, or the Worker's metering for the Scry tiers). A detector is on the frontier when no other costed detector is both at least as cheap and at least as good, and strictly better on one. The chart's stepped line and the table's “frontier” badge come from the same computation in the data build.
Write-ups and data
- Benchmark paper — methodology and results for the 311-pair screenshot benchmark (557 tagged issues).
- Scrymore/scry-design-diff-eval on Hugging Face — the dataset, CC-BY-4.0.
- 10–100× Cheaper Design Diffs: A Judged 60-Screen Scorecard
- Link Figma to any Storybook, then see where they drift
- Component Requests: From a Figma Component to a GitHub Issue in One Click
- The Cost Floor Moved: Five More Figma-to-Storybook Diff Agents
- Beyond Screenshots: AI Agents for Figma-to-Storybook Diffs
- Can VLMs Review Mobile UI? Introducing Scry Design Diff Eval