
The most common three dimensional website prompt on the internet is eleven words long: make me a 3D website, modern and clean, with smooth animations. Every model answers it the same way, because those eleven words contain no decisions to honour.
A prompt that produces a directed scene has seven fields, and six of them are numbers. The whole thing fits in under a hundred words, which is shorter than most of the atmospheric paragraphs it replaces.
The prompt everybody pastes#
Send the eleven word version to Claude, ChatGPT or Gemini and the reply is close to identical across all three. A TorusKnotGeometry, sometimes an icosahedron. OrbitControls, so the visitor can spin your hero upside down. One DirectionalLight at intensity 1, placed near the camera, which is the one position that erases every shadow describing the form. A MeshStandard surface in grey. A requestAnimationFrame loop adding a fixed number to rotation.y forever.
That is not a failure of the model. It is an accurate reading of a brief that specified nothing, and the three models agree because they are agreeing about what a default is rather than about what you want.
The instinct at this point is to write more. More usually means more adjectives: cinematic lighting, premium feel, buttery smooth, award winning. None of those words maps to an argument in an API, so none of them changes a single line of the output. The field of view stays at 50 because nothing said otherwise.
- Cinematic does not set a focal length.
- Premium does not set a roughness.
- Smooth does not set a duration or a curve.
- Modern does not set a palette, and reliably produces the blue and purple gradient that readers have learned to recognise on sight.
Words that describe a decision are not the same as words that make one. The seven fields below make them, one field at a time, and a model will honour every single one because each names a thing it already knows how to set.
The seven fields#
Each field closes one decision. Leave a field out and the default fills it, which is the entire mechanism behind every generated page looking related to every other.
| Field | Example value | What breaks without it |
|---|---|---|
| Subject | one lathed vessel, 2.1 units tall, shoulder at 0.7 height | a torus knot, because it needs no model file |
| Camera | 35 degrees at (0, 0.45, 5.2), target (0, 0.15, 0) | 50 degrees dead centre, no depth compression |
| Light | key 2.6 warm at (3.2, 4.8, 2.1), fill 0.35 cool, RoomEnvironment | one white light by the camera, nothing to reflect |
| Material | roughness: 0.42, metalness: 0, clearcoat: 0.35 | MeshStandardMaterial in grey plastic |
| Motion | camera z 5.2 to 3.1, pinned +=220%, scrub: 1 | an endless spin driven by a clock |
| Ceiling | under 150k triangles, DPR clamped to 2, loop paused off screen | uncapped pixel ratio and a hot phone |
| Refusals | no orbit controls, no gradient background, no bloom | all three arrive, because all three are defaults |
The refusal row is the one nobody writes and the one that changes the result most. A brief stating what the scene will not do is the only kind that produces two unrelated pages from the same starting point, because the positive half of a brief converges far faster than the negative half. Ask ten people for a premium three dimensional hero and you get ten similar answers. Ask them what their scene refuses and the answers scatter immediately.
Six of the seven rows are values you can decide in a minute each. The seventh, the subject, is the only one that needs any taste, and even that reduces to a silhouette and a height.
Two prompts, one model#
Here is the specified version, complete, at ninety-one words. Nothing in it is clever and nothing in it is long.
Build one hero section, full viewport, canvas only.
Subject: a single lathed ceramic vessel, 2.1 units tall, shoulder at 0.7 height.
Camera: PerspectiveCamera(35), position (0, 0.45, 5.2), target (0, 0.15, 0).
No OrbitControls.
Light: DirectionalLight 2.6 at (3.2, 4.8, 2.1), AmbientLight #8FA6C4 at 0.35,
RoomEnvironment via PMREMGenerator, envMapIntensity 1.25.
Material: MeshPhysicalMaterial, #E8E2D6, roughness 0.42, metalness 0,
clearcoat 0.35, clearcoatRoughness 0.28.
Ground #0B0B0F flat, FogExp2 0.085. ACESFilmic, exposure 1.12, DPR max 2.
Motion: ScrollTrigger pin +=220%, scrub 1, camera z 5.2 to 3.1, ease none.
Refuse: gradients, bloom, orbit, spinning.Those ninety-one words return a scene that looks like a decision. The eleven word version returns a scene that looks like the average of every scene. The gap between them is not phrasing, and no amount of rewriting the adjectives in the short version will close it, which is the whole case against prompting harder.
Two details in the block are worth pointing at, because almost every generated scene gets both wrong.
ease: "none"sits besidescrub: 1deliberately. A scrubbed timeline is already being eased by the reader's own scroll velocity, and a curve on top of that makes the camera feel like it is lagging behind the wheel.FogExp2at0.085against a flat#0B0B0Fground is what puts the object in a space. Without it the form floats, and the usual patch for a floating form is a background gradient, which is the tell readers report most often.
The performance ceiling belongs in the brief#
Six numbers keep a three dimensional hero fast, and a model will respect all six when they are stated. None of them are advanced.
- Triangles visible at once under
150,000. A hero object rarely needs more, and a casually subdivided sphere can quietly carry ten times that. - Textures at
2048pxon the long edge, in KTX2 once there are more than two. - Device pixel ratio clamped with
Math.min(window.devicePixelRatio, 2). This one line is the largest frame rate win available on a phone. - The render loop stopped by an
IntersectionObserverwhen the canvas leaves the viewport, rather than running behind three screens of text. - Explicit canvas dimensions and an
aspect-ratioon the wrapper, so nothing shifts while the scene loads. - The glb compressed with Draco or meshopt and fetched after first paint, with the hero sentence readable before the canvas arrives.
State the ceiling and the model builds under it. Leave it out and you get a scene that holds sixty frames on the machine it was written on, which is the one machine whose opinion does not matter.
The last bullet is a design decision disguised as a loading strategy. If the headline waits for the scene, the first thing a visitor meets is a blank canvas with a spinner in it, and the sentence that was supposed to sell the product arrives second, to a reader who has already decided how much to trust the page. Text first, scene second, on every project, with no exceptions worth making.
The verdict on prompt length#
Ninety words of values beat four hundred words of atmosphere, and the reason is mechanical rather than stylistic. A model resolves every unstated decision towards the centre of its training distribution. Values remove the decision. Adjectives describe the decision without making it, which is why a prompt can be long, vivid and completely inert.
So the useful test on any three dimensional prompt is to count the numbers in it. Fewer than six and the output will be a default scene wearing your subject matter. Six or more, each attached to a named API argument, and the scene is yours. That count is a better predictor of the result than the word count, the tone or the model you send it to.

Writing that brief per project is the slow part, and the reason a paid prompt is not the same thing as a free component. Thanor is a subscription library of these briefs: full art direction per design, in exact values, covering three dimensional scenes, animated backgrounds, page sections and app interfaces. $89 for three months, $189 a year, $299 once, with free accounts opening the designs marked free.
Browse the Thanor library and count the values in one design's direction. For the scene anatomy behind these fields rather than the prompt itself, there is the Claude walkthrough, and if you want to see what the library actually hands over, start with a free one.
Questions this raises
What should a 3D website AI prompt include?
Seven fields: the subject and its silhouette, the camera, the light, the material, the motion and its trigger, the performance ceiling, and the refusal list. Six of the seven are numbers or named constants, which is why the whole brief fits on a screen.
Does a longer prompt give a better 3D scene?
No. Length is not the variable, specificity is. Three hundred words of adjectives leave every technical decision open, and a model fills an open decision with the most common answer it has seen. Ninety words of values close them all.
Will the same prompt work in Claude, ChatGPT and Gemini?
Yes, when it is written as values. Each of them knows the Three.js API, so a brief naming a field of view, a roughness and a scroll distance produces the same scene in all three. What varies between models is the code around the scene, not the scene.
Do I need a 3D model file to start?
Not for the first version. A lathe or an extrude geometry built in code covers a surprising number of hero objects, and a directed material on a simple form beats a detailed model under flat light every time.
Part of
Build it with AI
Three dimensional scenes, scroll driven motion and animated pages, built by handing a model a specification instead of an adjective. What to write, and what comes back.
