Prompting harder does not work

The objection is that models ignore style guidelines. They do, because a guideline is phrasing. Here is the same brief as specification, counted.

Chaotic waves crashing against an unmoved glass panel that stays stubbornly generic, while beside it a precise glass frame quietly reshapes an identical panel
Left: four hundred words, zero decisions. Right: ninety words, thirty of them.

The objection to every prompt library is stated best by the people making it. "It's kind of wild in terms of how it will use different random designs, even given a specific style guideline." And in the same territory: "even if you constrain it to a pretty prescriptive toolkit it still does a pretty janky job of things." And the conclusion drawn from both: "LLMs don't really know how to give designs personality."

All three observations are accurate. The inference from them is wrong, and the difference is testable. What those style guidelines contained was phrasing, not specification, and the two behave nothing alike.

What a style guideline usually contains#

Open almost any brand or style guide written for a website and count the values in it. A typical one runs to several pages and contains four: two hex codes and two typeface names. Everything else is description.

  • A modern, clean aesthetic with plenty of white space. No spacing unit, no section padding, no measure.
  • Bold, confident typography with a clear hierarchy. No base size, no ratio, no line heights, no tracking.
  • Smooth, delightful micro-interactions. No duration, no curve, no list of what stays still.
  • A premium, trustworthy feel. Nothing at all.
  • Our primary colour is #4F46E5. One value, with no role assigned and no neutral, surface, border or muted text to go with it.

A model handed that document has almost nothing to obey. Six of the seven sentences describe a preference about decisions without making any of them, so the decisions arrive at the model still open, and an open decision is filled from the middle of the distribution. That is not the model going off brief. That is the model completing a brief that was mostly blank, in the only way a blank can be completed.

The objector's word random is worth examining too. The output is not random, it is highly predictable: the same gradient, the same three cards, the same 150ms curve, again and again. It looks random from the inside because it varies between runs while ignoring the guide, but it varies inside a very small space, which is exactly what convergence looks like from close up.

Phrasing versus specification, line by line#

Same intent on both sides. The left column is what briefs usually say, the right column is what closes the decision.

PhrasingSpecification
Bold, confident type hierarchy1.333 from 17px: 17, 23, 30, 40, 54, 72
Plenty of white space8px unit, section padding from {64, 96, 128} only
A premium colour paletteground #0E1014, surface #17191E, border #262A31, accent #D8623A
Neutral greys#F2EEE7, #A39C8F, #4A463F, all biased warm, never #808080
Smooth animations560ms, cubic-bezier(0.22, 1, 0.36, 1), headings only
Subtle depth1px border, no shadow, one radial light at 0.06 opacity
Readable body copy17px, line height 1.6, measure 66ch
Modern, minimal layouthero, proof, three-part explainer, pricing, FAQ, close
Clean and unclutteredno gradient, no marquee, no loop, nothing animates below fold two

Nine rows. The left column is four hundred words when written out in prose; the right column is about thirty values and fits on half a screen. Length was never the variable.

The last two rows are the ones that separate a real brief from a good one. A section order is a decision that almost no style guide contains, and a refusal list is a decision that none of them contain, yet those two rows do more to make two pages look unrelated than any amount of palette work. Colour is the decision people write down because it is the easiest one to name, not because it is the one carrying the page.

Why the prescriptive toolkit still failed#

The second objection is the more interesting one, because the objector did constrain the model and still got janky output. That happens, and it has a cause that is easy to verify.

A prescriptive toolkit usually means a component library: buttons, cards, inputs, a spacing helper. It specifies what things are made of and says nothing about which ones appear, in what order, at what size, with how much space between them. Those are composition decisions, and they are the ones a reader is actually reacting to.

Count the open decisions in a page built from a fully specified component library:

  1. Which sections exist, and in what order.
  2. How many items in the one grid on the page.
  3. Which single element on the page is allowed to be large.
  4. Where the vertical rhythm breaks on purpose, and where it holds.
  5. What the page refuses to do.
  6. Which one moment carries motion, and what stays still.

Six decisions, none of them touched by a toolkit, and all six visible from the first screen. A model closes all six towards the average, and the average is three cards in a row under a centred headline. So the toolkit was honoured and the page still looked generated, which is the outcome the objection describes exactly.

Two glass panels, the left surrounded by a dense cloud of vague swirling mist, the right held inside a precise armature of exact measured rules and guides
Same components, same colours. The right one had a section order and a size rule.

The count is the proof#

A brief's quality is measurable before you send it, which makes this a claim you can check rather than a claim you have to accept. Count the values in it.

  • Under 10 values. The output will be a default page wearing your copy. Nothing about the phrasing changes this.
  • 10 to 20 values. Recognisably yours in colour and type, still default in composition and motion.
  • 25 to 35 values. The page reads as directed. This is the target, and it is roughly one screen of text.
  • Over 60 values. Now you are specifying things that do not matter, and the brief starts to contradict itself.

Thirty is the number, and its composition matters as much as its size: six type sizes, one measure, five colours with roles, a spacing unit plus its allowed subset, one elevation rule, one curve with a duration, a section order, and a refusal list. Two of those eight are decisions about absence, which is the part that no component library and no free snippet will ever hand over.

Run the count on the last brief you wrote. If it comes back under ten, the model was never the problem, and the next rewrite of the same adjectives will return the same page it returned last time.

One caveat on the count, because it is a tool rather than a law. A value only counts if it is attached to something: #D8623A with no role assigned is not a specification, it is a colour sitting in a document. Every number in a real brief answers the question for what, and a brief of thirty unattached numbers performs about as well as no brief at all.

The verdict on prompting harder#

Rewriting a vague brief more forcefully changes nothing, because forcefulness is not a value. The objectors are right that models ignore style guidelines, and the reason is that the guidelines were not instructions: they were descriptions of instructions that were never written.

Specification is boring, short and mechanical. Thirty numbers, one screen, and the same thirty work in Claude, ChatGPT, Gemini, Lovable, v0, Bolt and Cursor without adjustment, because each of them is setting the same CSS properties and the same library arguments underneath.

Three things follow, and they are worth stating plainly:

  1. A model that ignored your guide was obeying it. There was nothing in it to disobey.
  2. The fix is upstream of the prompt. Decide the values, then write them down, then paste them. Deciding is the work; pasting is thirty seconds.
  3. The same thirty values are reusable across every page of a site, which is why this is a one-time cost rather than a per-page one.

The part that takes real time is deciding the thirty. Thanor is a subscription library of them: full art direction per design, written as values, covering UI, app and SaaS interfaces, page sections, animated backgrounds and three dimensional scenes, for personal and client work. $89 for three months, $189 a year, $299 once, and free accounts open the designs marked free.

Browse the Thanor library and count the values in one design, then count them in the brief you would have written. For the tells this fixes, tell by tell, there is the practical list, and one free design is enough to judge the format.

Questions this raises

Why does AI ignore my style guide?

Because most style guides are adjectives. Modern, clean, premium and minimal do not map to any value a renderer can set, so every decision they describe is still open when the model starts, and an open decision gets filled with the most common answer.

What is the difference between phrasing and specification?

Phrasing describes a decision. Specification makes it. Use a bold, confident type hierarchy is phrasing. A 1.333 scale from 17px giving 17, 23, 30, 40, 54 and 72, with line heights per step, is specification.

How many values does a brief need?

About thirty for a landing page. Six type sizes, one measure, five colours with roles, a spacing unit and its allowed subset, one elevation rule, one curve with a duration, and the section order. Count the numbers in your brief against thirty.

Do longer prompts produce better designs?

No. Length and specificity are different axes. Four hundred words of atmosphere leave every decision open, and ninety words of values close them all. The count that predicts the output is the number of values, not the number of words.

Part of

Hard questions

The questions a prompt library has to answer before anybody pays for one. Why every generated site looks the same, why a paid prompt beats a free component, and what the difference is actually made of.