
Three flagship models shipped inside four weeks this September. Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Claude Opus 5.5 on 22 September, with Gemini 3.1 Pro already in place. Hand any of them the same page to build and the results land within a hair of each other.
That is the answer, and it is not the answer the question expects. The model stopped being the variable. What varies is how many design decisions were already made by the time the prompt arrived.
The short call, before the detail:
- Building a whole marketing site in one pass?
Claude Opus 5.5, for the1Mtoken context. - Writing copy and layout in the same breath?
Claude Fable 5.1. - Already inside ChatGPT?
GPT-6 Astraneeds nothing extra from you. - Generating at volume on cheap tokens?
Gemini 3.8 FlashorClaude Haiku 4.5. - Want the page to stop reading as generated? That is a brief problem, and Thanor sells the brief.
Every flagship clears the same bar#
Correct modern CSS is table stakes now. All four write grid without being asked, size type with clamp(), respect prefers-reduced-motion, and produce a sticky header that does not jump on scroll. The floor of the category moved up, quietly, some time around the middle of the year.
The bar itself moved recently, which is worth stating. A year ago a generated page needed hand repair before anybody saw it: a nav that broke at 768px, focus states missing from custom controls, transforms that ignored the accessibility query. All four now get those right on the first pass, so code quality has stopped being a reason to prefer one over another.
What none of them does is decide. Ask for "a modern SaaS landing page" and each one fills the same empty slots in roughly the same way, because the gap gets filled from the same distribution:
- Typeface: a neutral geometric sans, very often
Inter. - Neutral: a grey sitting almost exactly on the axis,
#6B7280and its immediate neighbours. - Motion: opacity
0to1plustranslateY(20px), around300ms,ease-out, applied to every section alike. - Skeleton: hero, three feature cards, testimonial, call to action.
Read four outputs side by side and the family resemblance is not a flaw in any one model. It is what any model does with a slot nobody filled. The convergence is downstream of the instruction, which is the whole diagnosis of why generated sites rhyme.
So the useful comparison is not which model designs better. It is which six decisions you are willing to make before you press return.
The six decisions that change the page#
Six slots decide whether a page reads as art-directed or assembled. Every one of them is a value, and every one of them is yours.
| Decision | Left blank, you get | Named in the brief |
|---|---|---|
| Type scale | 16px body, headings by feel | 17px body, 1.26 ratio to 86px |
| Neutral | grey on the axis | #8A8577, hue pulled toward amber |
| Accent | a blue, or a violet gradient | #D8492B, one use only |
| Reveal | translateY(20px), 300ms | cubic-bezier(0.22, 1, 0.36, 1) at 620ms |
| Section order | hero, cards, testimonial, CTA | seven named blocks, stated in order |
| Density | whatever the framework ships | 64ch measure, 1.62 line height |
Six lines of specification, and the same model that produced a template produces something with a point of view. The accent used once rather than six times is a decision a model will never make unprompted, because restraint is not a default anywhere in the training data.
Two of the six rows are caps rather than values, and a cap is the thing a model never writes for itself:
- The accent appears on one element per section, never on body copy.
- The measure stops at
64chhowever wide the grid is allowed to run.
The last column is the entire product category. A prompt that carries those values has made the choices; a prompt that says "clean, modern, premium" has made none, and the vocabulary gap between the two is measurable.
Which one to reach for, by job#
Tier matters more than brand once the brief is written.
The full site in one pass#
Claude Opus 5.5, released 22 September, carries 1M tokens of context. That is enough to hold a design system, eight sections of copy and the existing component library in the same conversation, so section six still obeys the curve you named in section one.
One section, or a component#
Claude Sonnet 5, Claude Haiku 4.5 and Gemini 3.8 Flash all execute an exact value faithfully. A hero with six named values and a stated refusal list does not need a flagship, and the cheap tiers turn a specification around fast enough to iterate twice in the time a flagship takes once.
Copy and layout together#
Claude Fable 5.1 is Anthropic's top model and the one to point at a page where the words and the composition have to agree. GPT-6 Astra sits in the same place in OpenAI's lineup, with GPT-6 Sol and GPT-6 Luna below it since 22 September.

Claude Mythos exists and access to it is limited, so it is news rather than a tool you can plan around. Claude Design generates production HTML, CSS and JS from a prompt, which makes it a surface for a specification rather than a competitor to one.
What no model decides for you#
Refusals. Every real design is defined as much by what it will not do as by what it does, and no model has ever volunteered a refusal.
A brief that names three of them changes the output more than any model upgrade released this year:
- One accent, one place. Name the single element allowed to carry
#D8492Band the page stops looking like a swatch demo. - No gradient behind text. A gradient arrives wherever depth was never decided, and it reads as filler because it is.
- Nothing reveals twice. Motion on the first viewport only, and the rest of the page holds still.
- No third typeface, no icon set. Both arrive to fill a gap and both add a vocabulary you did not choose.
Refusals have to be written per design rather than as general taste advice, because the useful ones contradict each other. A page whose argument is density refuses whitespace; a page whose argument is calm refuses the counter and the draw-on. Generic advice cannot hold both positions, which is why a library that states them per design produces work that diverges.
The second thing no model decides is the signature moment: one mechanism, on one section, described as physics rather than mood. A counter that eases with cubic-bezier(0.32, 0.72, 0, 1) over 240ms, a card that lifts 3px on hover, a mark that draws its own stroke once on load. Ask for "delightful micro-interactions" and you get four of them fighting.
Write the refusals down next to the six values and the brief is finished. That is roughly a page of text, and it is the same page whichever model you paste it into.
The verdict, named#
Pick the model you already have a subscription to. Claude Fable 5.1, Claude Opus 5.5, GPT-6 Astra and Gemini 3.1 Pro are close enough on design work that the choice is a workflow question, not a quality one. Then put your effort where it moves the result.
The ranking that matters runs the other way:
- A specified brief in
Claude Haiku 4.5beats an adjective in any flagship. - The same brief in a flagship beats it in a cheap tier, by a margin you have to look for.
- Two different briefs in the same model produce pages that look unrelated.
That last line is the one worth keeping. The spread between two prompts is wider than the spread between two models, which puts the advantage in the document rather than the dropdown.
It also decides what to do with the next launch. A release raises the floor of the code and leaves the design exactly where your brief left it, so the honest response to a new model id is to paste the same page of values into it and check that the values survived. That is a ten minute job, and it is the whole upgrade process. For the newest release put through this exact test, read the Opus 5.5 write-up.
If writing that document is the part you would rather skip, that is the job Thanor does: exact values per design, for personal and client work. Browse the Thanor library and count the decisions in one free design before you decide whether it beats the brief you would have written.
Questions this raises
Which AI model is best for web design?
All four current flagships produce a page of comparable quality from the same instruction. `Claude Opus 5.5` handles the largest single pass because of its `1M` token context. `Claude Fable 5.1` is Anthropic's top model. `GPT-6 Astra` is OpenAI's flagship. `Gemini 3.1 Pro` leads the Gemini 3 series. Pick on where you already work, then spend your effort on the brief.
Does a better model make a better looking website?
Not past a point the category already cleared. Every flagship writes correct grid, `clamp()` sizing and a reduced-motion query without being asked. None of them chooses your type scale, your neutral's hue bias or the duration of your reveal. Those are the parts a visitor reacts to, and they come from the prompt.
What should a web design prompt actually contain?
Six things, as values rather than adjectives: four to six hex codes with a role each, two typefaces with a scale in px, one named easing curve with a duration in milliseconds, the section order, the page's density, and a short list of what the design refuses to do.
Is a cheap model good enough for a landing page?
For a single section against a written specification, often yes. `Claude Haiku 4.5` and `Gemini 3.8 Flash` follow exact values as faithfully as the flagships do. Where they fall behind is holding a long page together and keeping one decision consistent across eight sections.
Where do I get a specification if I cannot write one?
That is what the Thanor library is: full art direction per design in exact values, covering UI, app and SaaS interfaces, page sections and animated backgrounds. `$89` for three months, `$189` a year, `$299` once, and free accounts open the designs marked free.
Part of
Every model, on design
Each new model put on the same job, designing a website, and judged on what it actually produces. Which one to reach for, and the part of the result that does not depend on the model at all.
