AI website slop, and how to measure it

One audit ran Playwright over 1,590 Show HN landing pages and scored 22 percent heavy slop. Here are the ten checks, and the value that clears each.

A careless heap of identical glass panels dumped on each other, every one carrying the same blue and purple gradient and the same three card arrangement
Nearly a quarter of the grid scores as heavy slop. The outliers were specified.

Somebody pointed Playwright at 1,590 Show HN landing pages and scored them. Twenty-two percent came back as heavy slop, thirty-two percent as mild, forty-six percent as clean. That is a self-selected technical audience shipping their own launches, and slightly more than half the sample showed signs.

The audit stopped there, as this genre usually does. A number is a diagnosis, and a diagnosis has never fixed a page. Below is the detector, and beside each check the value that clears it.

Slop, counted#

The reaction to these pages is stronger than the design failure alone would explain, and the reason is in what readers say rather than in the numbers.

most AI developed stuff is just so insanely soulless crap I can instantly tell and instinctively close the tab

a comment on Hacker News

The operative word is instantly. Nobody in that thread describes reading the copy and finding it thin. They describe a judgment made from the first screen, in under a second, on the basis of visual pattern alone, and then closing the tab. Whatever the product does never enters the decision.

Another comment in the same territory is more damaging to a founder than any design critique: "It's average. Painfully average, to the point of it being easily mistaken for a scam." Average is the accusation. Not ugly, not broken, average, and average has been re-priced by volume. The same argument appears again elsewhere in one line: a cover like a hundred others is zero effort, and trust is built on invested effort.

That is why a slop score matters commercially rather than aesthetically. A visitor cannot inspect your engineering from the landing page, so they use the landing page as the sample of your work that they can inspect. If it looks unattended, the reasonable inference about everything behind it is unflattering, and the convergence that produces it is now well documented. Twenty-two percent of a technical audience's own launches are paying that price without knowing the checks exist.

The detector, ten checks#

Run these in order. Each one is visible in seconds, each has a binary failing condition, and each has a value that clears it. The first three catch most of the twenty-two percent on their own.

CheckFails whenValue that clears it
Backgroundany multi-stop gradient behind the heroone flat ground, #0E1014, plus a single accent
Neutral huea grey sits on the axis, #808080 or #EEEEEEneutrals biased: #F2EEE7, #A39C8F, #4A463F
Type stepstwo sizes within 6px of each otherratio 1.333 from 17px: 17, 23, 30, 40, 54, 72
Section paddingmore than three distinct vertical paddings64, 96, 128 only, on an 8px unit
Elevationshadows on some cards, not othersone rule: 1px border #262A31, no shadow, or shadow on all
Curvesevery transition at 150msone entrance curve, cubic-bezier(0.22, 1, 0.36, 1), 560ms
Entrance scopeevery section fades and slides upheadings only, once, never on scroll back
Card countexactly three, of anythingcount set by the argument: two, four or five
Measurebody copy spanning the full viewport66ch, held at every width
Icon setline icons at three different stroke weightsone set, one weight, 1.5px

Nine failures out of ten is a heavy score. Three to five is mild. Zero to two is clean, and clean here means specified rather than tasteful: every value above is arithmetic, and none of it requires an eye.

The gradient check is first because it is the tell readers name most. One comment puts it exactly: "The overuse of blue and purple gradient fills on the landing page is a telltale sign of AI slop."

Why the gradient carries so much weight#

A gradient is the default answer to a question nobody asked, and that is the whole of its meaning to a reader. Faced with a background and no palette, a model produces a transition between two safe hues, because a flat colour would expose that no colour was chosen.

What to write instead, as five values:

  • Ground. One flat value. #0E1014 for a dark page, #F7F4EE for a light one. Both are off-axis on purpose.
  • Surface. One step from the ground, #17191E on dark, #FFFFFF on light, and never a third level.
  • Accent. One hue, used for at most two things on the page. #D8623A warm, #2F6F5E cool. Anything more and the accent stops being an accent.
  • Border. One value, #262A31 on dark, at 1px. This replaces most of what people use shadows for.
  • Text. Two values, primary and muted, at a measured contrast ratio of at least 4.5:1 for body and 3:1 for large display type.

Five values, and the gradient has nothing left to do. If a design genuinely wants depth, the honest mechanism is a single soft radial light at very low opacity behind one element, not a two stop wash across the viewport, and the difference is that one of those was decided.

A single blue and purple gradient glass panel lifted out of a dim heap of identical ones and held under a hard white inspection beam
Five values on the right. The left one is what a model writes when it has none.

What the threads agree on#

Complaints about this appear in near-identical language across Hacker News and the design subreddits, and that agreement is what makes them useful: a critique one person makes is taste, and a critique thirty people make in the same words is a specification waiting to be written. Three observations recur.

The first is that the tells are compositional, not technical. "the paddings and margins are inconsistent and don't convey visual rhythm" is a spacing unit problem, and "why so many different font sizes with no hierarchy?" is a type scale problem. Neither is a code quality issue, which is why better engineering does not help.

serif headline with mono kickers. Sprinkle in em dashes and punchy writing and voila, you've got yourself a Claude coded page.

a comment on Hacker News

The second observation is that the tool leaves a signature, and the comment above names it precisely. That combination was distinctive once, and a year of volume turned it into an identifier, which is the fate of every default eventually. The lesson is not to avoid serif headlines or mono kickers. It is that any pairing adopted by default becomes a tell, so the pairing has to be chosen for this project rather than inherited from the last thousand.

The third observation is the sharpest of the three: "it looks and behaves like the AI default." Behaves is the important word. Readers are reading the motion as well as the surface, and the motion default is even more uniform than the visual one, because framework documentation argues about colour and almost never publishes a duration.

The verdict on slop#

Slop is not a quality of output, it is a quantity of unmade decisions, and it is measurable. Twenty-five values remove it: six type sizes from a ratio, one measure, five colours with roles, a spacing unit and its allowed subset, one elevation rule, one curve with a duration, and the section order.

Every one of those is a number or a named constant. None of them requires design training to write down, and all of them can be pasted into any model.

Three moves, in order:

  1. Score your own page against the ten checks before anybody else does, and write down which ones fail rather than arguing with them.
  2. Write the value that clears each failure, taking the numbers from the table above or from your own brand if you have one.
  3. Rebuild from the values, not from the page. A model given twenty-five numbers produces a different page than one given the old page and a request to improve it.

That third step is the one people skip, and skipping it is why a rebuild often lands back in the mild band. Editing a slop page keeps its structure, and the structure is where most of the score lives.

Thanor is a subscription library of exactly those value sets: full art direction per design, covering UI, app and SaaS interfaces, page sections and animated backgrounds, for personal and client work. $89 for three months, $189 a year, $299 once, and free accounts open the designs marked free.

Browse the Thanor library and run the ten checks against one design's direction. For the fix written tell by tell there is the companion piece, and one free design shows how the refusals are stated.

Questions this raises

What counts as AI website slop?

A page built entirely from unmade decisions: a gradient standing in for a palette, a single framework curve on every element, sizes too close together to form a hierarchy, and spacing invented per section. It is not bad design, it is absent design.

Is there an AI website slop detector?

There is a reliable set of checks you can run by eye in about two minutes, and most of them can be automated with a headless browser. Ten of them are listed in this article, each with the failing condition and the value that clears it.

How common is it really?

One audit of 1,590 Show HN landing pages scored 22 percent as heavy, 32 percent as mild and 46 percent as clean. Slightly more than half of a technical, self-selected sample showed at least mild signs.

Why do readers react so strongly to it?

Because a page that looks like a hundred others carries no evidence that anybody worked on it, and readers convert absent effort into absent trust. Several describe closing the tab before reading a word.

Part of

Hard questions

The questions a prompt library has to answer before anybody pays for one. Why every generated site looks the same, why a paid prompt beats a free component, and what the difference is actually made of.