CompletedJul 27, 2026

Frontend Design Benchmark

Frontend design research

A side-by-side benchmark comparing GPT/Codex and Claude Opus 5 across five identical frontend prompts, with and without design skills.

View benchmark

What we tested

Five product prompts were given verbatim to GPT/Codex and Claude Opus 5. Each model built every prompt twice: once without design guidance and once with frontend design skills loaded.

Why we built it

Frontend quality is difficult to judge from isolated screenshots or broad claims. Putting the builds together makes differences in hierarchy, composition, typography, and product judgment easier to inspect.

How to explore it

Open any build on its own, or use the side-by-side view to compare models and toggle skills on and off. A legacy section keeps the earlier Claude Opus 4.8 results available for reference.

Start something

Have a workflow worth building?

Contact Diethos