
Every flagship shipping today writes frontend code that works. Claude Fable 5.1, Claude Opus 5.5, GPT-6 Astra and Gemini 3.1 Pro all produce accessible markup, keyboard focus states, container queries and a reduced-motion branch without being asked for any of it.
So the interesting question is not which one codes better. It is which one returns an interface somebody designed, and that answer does not live in the model at all. It lives in whether a token file arrived with the request.
The short call:
- Long build, one context:
Claude Opus 5.5,1Mtokens. - One component against a spec:
Claude Haiku 4.5orGemini 3.8 Flash. - Top of each lineup:
Claude Fable 5.1andGPT-6 Astra. - The values that make the component look art-directed: written once, in the Thanor library, and pasted into whichever of the four you use.
The ceiling is your token file#
A component has about fourteen decisions in it, and a typical prompt names none of them. Radius, border colour, the ramp the padding comes off, the hover transform, the duration of that transform, the disabled treatment, the focus ring offset, the shadow or the deliberate absence of one.
Leave them out and the model supplies the category median, every time:
border-radius: 8px, on everything from the button to the modal.box-shadow: 0 1px 3px rgba(0,0,0,0.1), the Tailwind default, unchanged.transition: all 0.2s ease, which animates properties you did not intend.- A
1pxborder in a grey with no relationship to the surface behind it.
None of that is wrong code. All of it is unowned code, and a designer reading the result can tell inside two seconds, because the values agree with the framework rather than with each other.
The clearest way to see it is to open two components from the same session. The button transitions on all 0.2s, the card on transform 0.3s, and nothing in the file says which one was intended. Two durations that close to each other read as a mistake rather than as a decision, because that is what they are.
The fix is boring. Write fourteen values down once, keep them in a file, and paste them above every request. The model stops guessing and starts executing, which is the only behaviour change worth chasing. That is the same mechanism as the gap between a paid prompt and a free component: one carries decisions, the other carries markup.
What a frontend brief has to name#
Seven groups, and a page of text covers all of them.
| Group | Blank leaves you with | Named looks like |
|---|---|---|
| Surface | #FFFFFF on #F9FAFB | #14141A raised on #0B0B0F |
| Border | 1px solid #E5E7EB | 1px solid rgba(242,239,233,0.08) |
| Radius | 8px everywhere | 2px controls, 14px cards |
| Spacing | arbitrary multiples of 4 | ramp 4 / 8 / 12 / 20 / 32 / 52 |
| Type | 16px and browser defaults | 17px at 1.62, scale ratio 1.26 |
| Response | all 0.2s ease | cubic-bezier(0.32, 0.72, 0, 1) at 240ms |
| States | hover only | rest, hover, focus, active, disabled, loading |
The last row is the one that separates a mock from an interface. Six states per interactive element, named as a requirement, and the model builds all six because you asked for six. Ask for "a polished button" and you get hover.
Radius is the cheapest tell to fix. A 2px control next to a 14px card is a decision somebody made; 8px on both is the framework talking.
The spacing ramp is the one people leave out most often and miss most. With a fixed set of six values and a rule that nothing may sit off it, every gap in the interface becomes a choice between six options rather than a free number, and the vertical rhythm holds across screens built weeks apart. Without it you get 18px here and 22px there, both defensible, neither related.
Model by model, on frontend work#
The tiers differ by how much they can hold, not by how well they write a flex container.
Holding a long build together#
Claude Opus 5.5, released 22 September with 1M tokens of context, keeps a token file, a component inventory and the existing code in view at once. The practical effect is consistency: component nineteen still uses the 240ms response curve from component one, because component one is still in the window.
Volume work on a written spec#
Claude Haiku 4.5 and Gemini 3.8 Flash follow exact values as faithfully as anything above them. Twenty small components against a token file is the job they are for, and the saved budget buys a second iteration.
Everything in between#
Claude Sonnet 5 is the default for a page at a time. Claude Fable 5.1 is Anthropic's top model, released 1 September, and GPT-6 Astra, released 3 September, leads OpenAI's lineup with GPT-6 Sol and GPT-6 Luna underneath it. Gemini 3.1 Pro is the current Gemini 3 series flagship, with Gemini 3.8 Flash as the latest stable Flash.
Pick one and stay there for a project. Switching mid-build is not a quality risk, it is a consistency risk: two models given the same token file agree, and two models given a half-remembered conversation do not.

Where the models actually diverge#
Three places, and none of them is design taste.
- Context length. A
1Mwindow is a different kind of tool from a200Kone when the build is a whole application. Nothing about the CSS improves; the fifteenth file just stops contradicting the second. - Instruction density tolerance. Hand a model forty explicit constraints and some tiers start dropping the ones near the end. The flagships hold more of them, which is why a dense brief is worth a dense model.
- Refactoring an existing file. Editing code you did not write rewards reasoning depth more than generating a greenfield component does.
Notice what is absent from that list. Colour sense, type judgement, spacing rhythm, motion feel: none of it varies by model, because none of it is being produced by the model. It is being copied from the brief, or invented from the median when the brief is silent.
That is worth saying plainly, because the upgrade cycle invites the opposite belief. A new release raises the floor of the code and leaves the design exactly where your prompt left it. The same pattern shows up in every launch this cluster covers.
There is a practical test. Keep one component as a fixture: a card with your radius, your border, your ramp padding, your 240ms response curve and all six states. On the day a new model ships, paste the token file and ask for that card. If the fourteen values come back intact, the release changes nothing about your design work, and you have learned it in four minutes.
The verdict for frontend work#
Use Claude Opus 5.5 for a build you want held in one context, and a cheap tier for anything you can specify tightly enough to hand over in one message. Both produce the same quality of component when both get the same token file.
The order of operations matters more than the choice. Write the tokens, build one component, check all six states by keyboard, then batch the rest against that component as the reference. Doing it the other way round, generating twenty components and reconciling them afterwards, costs more time than the tier ever saves.
Ranked by how much each choice actually changes the interface:
- The token file: everything.
- The spacing ramp: whether screens built weeks apart share a rhythm.
- The state list: the difference between a mock and a working control.
- The refusal list, meaning what this design will not do: the reason two projects off one library look unrelated.
- The model: a workflow preference.
Thanor writes those files: exact values per design, covering UI, app and SaaS interfaces, page sections and animated backgrounds, for personal and client work. $89 for three months, $189 a year, $299 once, and free accounts open the designs marked free.
Browse the Thanor library and read one free design as a token file before you decide whether a model upgrade was the thing you needed. Fourteen values, one afternoon, and the same file keeps working through the next three releases.
Questions this raises
What is the best AI model for frontend development?
For a long build held in one context, `Claude Opus 5.5` and its `1M` token window. For single components against a written specification, `Claude Sonnet 5`, `Claude Haiku 4.5` or `Gemini 3.8 Flash` produce the same result far cheaper. `Claude Fable 5.1` and `GPT-6 Astra` sit at the top of their respective lineups.
Do frontend results differ much between models?
Less than the marketing suggests and less than the prompt does. All the current flagships write accessible markup, correct focus states and container queries unasked. The divergence shows up in consistency across a long file, not in whether the code works.
What should a frontend design brief contain?
A token file. Colour with a role per value, a spacing ramp, a type scale in px, radii, border colours as `rgba()`, one easing curve with a duration, and the states every interactive element must implement. Values, not adjectives.
Why does AI frontend output look the same across projects?
Because unnamed values get filled from the same distribution every time: `Inter`, a grey on the neutral axis, `translateY(20px)` at roughly `300ms`, and an `8px` radius on everything. Name those four and the resemblance disappears.
Can a cheap model build a component library?
It can build the components. What it struggles with is keeping the twentieth component obedient to a decision made in the first, which is a context problem rather than a reasoning one. Write the tokens down once and paste them with every request.
Part of
Every model, on design
Each new model put on the same job, designing a website, and judged on what it actually produces. Which one to reach for, and the part of the result that does not depend on the model at all.
