Scry / Research
Scry Research
Benchmarks and cost/performance results for vision-language models and agents doing mobile UI and Figma-to-Storybook diff review.
Posts · 6
10–100× Cheaper Design Diffs: A Judged 60-Screen Scorecard
In the last two posts we gave a review agent everything a coding agent has (the Figma layer tree, a Storybook capture at the exact frame size, the live DOM and the implementation source) and compared models on one dense…
Link Figma to any Storybook, then see where they drift
Every team with a design system keeps it twice: once as Figma components, once as code with a story for each state. The two copies start identical and then drift apart one merged PR and one nudged auto-layout at a time.…
Component Requests: From a Figma Component to a GitHub Issue in One Click
Every design-system team has a component that fell through the cracks. A designer draws a new alert banner for the billing flow. It gets variants, it gets reviewed, it goes into the library. Then someone has to tell…
The Cost Floor Moved: Five More Figma-to-Storybook Diff Agents
Our previous post gave an agent the same evidence a coding agent would have — the raw Figma layer tree, a Storybook screenshot captured at the exact frame size, the live DOM with computed styles, and the implementation…
Beyond Screenshots: AI Agents for Figma-to-Storybook Diffs
In our first Scry Design Diff Eval post, we asked whether vision language models could review mobile UI from screenshots alone. The short answer was: somewhat, but not well enough to be competitive with a human reviewer.…
Can VLMs Review Mobile UI? Introducing Scry Design Diff Eval
In most mobile teams the flow of UI Development is often modeled as a one directional flow of Designer -> Frontend Dev. In reality there is often a ping ponging back and forth between the nodes on final implementation…
New results land here whenever a frontier model ships. Subscribe via RSS, or read the benchmark paper and the leaderboard.