Best model for UI design, by tier

Which Claude tier to point at an interface, and why the tier changes the cost of the job far more than it changes the design that comes back.

A cluster of small frosted glass interface panels floating at different depths, each holding an abstract arrangement of rounded blocks and circles
Four sizes, one layout. The etching is the part you supply.

Claude Fable 5.1 is the model to point at an interface, and the reason has almost nothing to do with visual judgement. Released 1 September and generally available, it holds the most constraints at once, which is the only property of a model that reliably shows up in a UI.

Tier changes cost and consistency. It does not change taste. Hand Claude Haiku 4.5 and Claude Fable 5.1 the same fully specified card component and the two results are difficult to tell apart. Hand both the word "clean" and they are identical, in the same disappointing way.

The short call:

  • Full interface, many screens: Claude Fable 5.1.
  • Design system plus existing code in one context: Claude Opus 5.5, 1M tokens.
  • A screen at a time: Claude Sonnet 5.
  • One component against a spec: Claude Haiku 4.5.
  • The spec itself: written already, in the Thanor library.

UI output is a decision problem#

An interface is roughly thirty decisions repeated across every screen. A model asked to design one without those decisions does not refuse; it supplies the median and moves on.

The median is recognisable, and once you can name it you see it everywhere:

  • Panel on page, both near-white, separated by a 1px grey line.
  • 8px radius on the card, the input, the button and the avatar alike.
  • Body copy at 14px, because dashboards default small, with no stated measure.
  • A single hover state, no focus ring, no loading treatment, no empty state.
  • Icons from one set at 20px, floating without an alignment rule.

That output is not broken. It is unowned, and a user feels that as flatness rather than as a bug. Interfaces earn their premium read from density and restraint, which are two of the first things a blank prompt discards.

Density is the one worth chasing first, because it is almost free. Body copy at 17px with a 1.62 line height and a 64ch measure, on a #0B0B0F ground with a #14141A raised panel, already reads as a considered product before a single component has been styled. The same layout at 14px on near-white with a grey hairline reads as an admin tool, and nothing else about it changed.

The repair is a page of values. Thirty decisions, written once, reused on every screen, which is the same discipline a design system encodes and the same one a real prompt library ships in place of adjectives.

The four Claude tiers, ranked by job#

Pick the tier by what has to stay in view, not by how good you want the design to be.

TierReach for it whenPractical limit
Claude Fable 5.1a whole interface, many screenscost per pass
Claude Opus 5.5system plus existing code togethercost per pass
Claude Sonnet 5one screen, fully specifiedlong multi-file edits
Claude Haiku 4.5one component, exact valuesconsistency across twenty

Claude Opus 5.5, released 22 September, is the one to use when the existing codebase is part of the problem. A 1M token window holds the token file, the component inventory and the screens already built, so screen nine obeys the decision made on screen one rather than reinventing it.

Two cautions before you configure anything. Claude Mythos is real and access to it is limited, so it belongs in a news item rather than in a workflow. And at least one competing guide documents an Opus version that was never shipped: the current Opus release is Claude Opus 5.5, and a pipeline pinned to an invented model id fails at the first call.

The tier you pick also sets how you work. A flagship rewards one long, dense session where the whole system stays in view. A cheap tier rewards many short requests, each carrying the token block again, which is a different discipline and often a faster one.

What a UI brief names that a model will not#

Five groups. Every one of them is a value, and every one of them is the difference between a mock and an interface somebody owns.

  1. Surfaces, as a stack. #0B0B0F page, #14141A raised, #1C1C24 overlay, and a border of 1px solid rgba(242,239,233,0.08) rather than a grey line with no relationship to either.
  2. Density, as numbers. 17px body at 1.62 line height, a 64ch measure on prose, spacing off a fixed ramp of 4 / 8 / 12 / 20 / 32 / 52.
  3. Response, as one curve. cubic-bezier(0.32, 0.72, 0, 1) at 240ms for anything a user triggers, 620ms and cubic-bezier(0.22, 1, 0.36, 1) for anything the page does on its own.
  4. Six states per control. Rest, hover, focus, active, disabled, loading. Ask for six and you get six.
  5. Refusals. No gradient behind text. The accent, #D8492B, appears on one element per screen. Nothing animates on scroll below the first viewport.
Four glass interface panels in a tight fan, each showing the same abstract layout rendered at a different level of polish
One curve, one duration, everywhere a user touches. Write it down and it stops being a mood.

The refusal list is the half almost nobody writes, and it is the half that makes two interfaces built from the same library look unrelated to each other. Write it per design rather than as general advice, because the useful refusals contradict one another: an interface whose argument is density refuses whitespace, and one whose argument is calm refuses the counter and the draw-on.

Where cheap tiers hold up#

A fully specified component does not need a flagship. Claude Haiku 4.5 given a hex value, a radius, a ramp position and a 240ms duration returns those exact values, because following an instruction is cheap and inventing a taste is not being asked for.

Where the cheap tiers slip is memory across a batch:

  • Component four uses the 14px radius you named. Component eighteen quietly drifts to 12px.
  • The focus ring offset is right in the button and missing in the select.
  • The loading state exists on the form and not on the table.

All three are context failures, and all three have the same fix: paste the token block with every single request rather than relying on the conversation to remember it. A tier upgrade buys the same reliability at a higher price, which is a legitimate trade when the batch is large.

There is a cheap check for all three. Build one component, then ask the same tier for a second component that has to sit next to it, pasting the token block again. If the radius, the border colour and the focus offset match without being corrected, the block is complete. If they drift, the block is missing a value, and it is faster to find out on component two than on component twenty.

Claude Design, Anthropic's prompt-to-prototype tool, generates production HTML, CSS and JS, and it rewards the identical discipline. What it does with a real specification is the clearest demonstration of the point this article keeps making.

The verdict for interface work#

Claude Fable 5.1 for a whole interface, Claude Opus 5.5 when existing code has to stay in the window, Claude Sonnet 5 for a screen, Claude Haiku 4.5 for a component. That is the entire tier question, and it is a budget decision.

Anybody arriving at this question because an interface came back looking ordinary is diagnosing the wrong part. The tier decided what the session could hold. The brief decided what the screen looks like, and the brief is the artefact you can change this afternoon.

The design decision sits elsewhere:

  • Thirty values, written once. Surfaces, ramp, scale, radii, one curve, six states.
  • Three refusals, stated per design. What this interface does not do.
  • Pasted with every request, so no tier has to remember anything.

Do that and the tier becomes what it should be, a line on an invoice. The same test run on the newest release lands in exactly the same place, which is the reason this cluster treats the brief as the durable asset and the model name as the part that expires.

Thanor sells those values as finished art direction per design, covering UI, app and SaaS interfaces, page sections and animated backgrounds, for personal and client work: $89 for three months, $189 a year, $299 once, and free accounts open the designs marked free.

Browse the Thanor library and read one free design as a token block before you spend another afternoon choosing a tier.

Questions this raises

What is the best Claude model for UI design?

`Claude Fable 5.1` is Anthropic's top model and the one to reach for on a full interface. `Claude Opus 5.5` carries `1M` tokens of context, which matters when the design system and the existing screens have to stay in view. `Claude Sonnet 5` handles a screen at a time and `Claude Haiku 4.5` is enough for a single component against a written spec.

Does the top tier design a better interface?

It holds more constraints at once and stays consistent over a longer file. It does not have better taste, because taste is not what it is producing. Every tier copies the values in the brief and invents the values the brief left out.

How long should a UI design prompt be?

Long enough to carry the values and no longer. A spacing ramp, a type scale in px, surface and border colours, radii, one easing curve with a duration, the six states per control and a short refusal list. That is roughly a page.

Is Claude Haiku 4.5 good enough for UI work?

For one component against an exact specification, yes. It follows a hex value and a `240ms` duration as faithfully as anything above it. Its limit is holding a decision steady across twenty components, which is a context problem you solve by pasting the tokens every time.

What does a specification change about the result?

The parts a user reacts to: the density, the neutral's hue bias, whether the accent appears once or six times, how fast a panel settles. Left blank, all four tiers produce the same median interface, which is why the brief is the variable and the tier is a budget line.

Part of

Every model, on design

Each new model put on the same job, designing a website, and judged on what it actually produces. Which one to reach for, and the part of the result that does not depend on the model at all.