CompletedJul 27, 2026
Frontend Design Benchmark
Frontend design research
A side-by-side benchmark comparing GPT/Codex and Claude Opus 5 across five identical frontend prompts, with and without design skills.
View benchmarkWhat we tested
Five product prompts were given verbatim to GPT/Codex and Claude Opus 5. Each model built every prompt twice: once without design guidance and once with frontend design skills loaded.
Why we built it
Frontend quality is difficult to judge from isolated screenshots or broad claims. Putting the builds together makes differences in hierarchy, composition, typography, and product judgment easier to inspect.
How to explore it
Open any build on its own, or use the side-by-side view to compare models and toggle skills on and off. A legacy section keeps the earlier Claude Opus 4.8 results available for reference.
Start something