
Claude writes a working Three.js scene on the first pass, reliably. What it cannot do is guess your focal length, so it takes the library default: a 50 degree field of view, MeshStandardMaterial at roughness: 1, one white DirectionalLight at intensity 1, OrbitControls, and a mesh rotating at a constant speed forever.
That page is technically three dimensional and visually nothing. The object reads as grey plastic because nothing told the material what it is made of, and the highlights blow out to flat white because nothing told the renderer how to map them.
Write the scene as values and it lands in one pass. About a dozen numbers do the whole job, and every one of them is in this article.
What comes back when a brief says 3D#
Ask for a three dimensional hero with no numbers in the request and the output is consistent enough to predict line by line.
- A
TorusKnotGeometryor anIcosahedronGeometry, because those are the two primitives that look like something without a model file. OrbitControls, which hands the visitor a turntable and hands you a hero they can spin upside down.- One
DirectionalLight, usually near the camera, which is the one position that removes every shadow that would describe the form. MeshStandardMaterialwith a colour and nothing else, so the surface has no roughness story, no clearcoat and no environment to reflect.- A
requestAnimationFrameloop adding a fixed amount torotation.yon every frame, which is motion tied to a clock rather than to a reader. - No tone mapping, no output colour space set, no device pixel ratio clamp.
Each of those is a reasonable library default and a poor design decision, and they share one cause: the brief left the decision open. A model fills an open decision with the most common answer in its training data, which is exactly what a default is.
The cure is not a longer request. A paragraph saying cinematic, premium and modern leaves every one of the six items above unspecified, so the six come back unchanged. Adjectives do not close decisions, and an open decision is where a default lives.
The cure is naming values, and there are fewer of them than people expect. Twelve settings cover a hero scene end to end, eleven of the twelve are a single number or a single constant, and none of them require you to know how a renderer works. You only have to state the answer instead of leaving the question open, which is the same discipline that makes a specified brief beat a longer one.
The scene, written as values#
Twelve settings separate a default scene from a directed one. Each maps to a single argument in an API Claude already knows, so a brief naming them gets honoured rather than approximated.
| Setting | Value to name | What the default does instead |
|---|---|---|
| Camera | PerspectiveCamera(35, aspect, 0.1, 60) | 50 degrees, which fattens the object and flattens the light |
| Camera position | (0, 0.45, 5.2) looking at (0, 0.15, 0) | dead centre at eye height, the one angle with no drama in it |
| Pixel ratio | Math.min(window.devicePixelRatio, 2) | uncapped, so a phone renders nine times the pixels it needs |
| Tone mapping | ACESFilmicToneMapping, exposure 1.12 | none, so every specular highlight clips to #FFFFFF |
| Colour space | outputColorSpace = SRGBColorSpace | right since r152, if the textures are tagged too |
| Environment | RoomEnvironment through PMREMGenerator, envMapIntensity: 1.25 | a black void, which no real surface has ever reflected |
| Key light | DirectionalLight(0xffffff, 2.6) at (3.2, 4.8, 2.1) | intensity 1 beside the camera |
| Fill | AmbientLight(0x8FA6C4, 0.35), cool against the warm key | nothing, so the shadow side goes to pure black |
| Shadows | PCFSoftShadowMap, mapSize 2048, normalBias 0.02 | off, or on with acne across every curved face |
| Fog | FogExp2(0x0B0B0F, 0.085) | none, so the object floats with no depth cue |
| Controls | none, or enableZoom: false and maxPolarAngle: 1.45 | full orbit, including upside down |
| Background | one flat value, #0B0B0F | a gradient, standing in for a decision nobody made |
Two rows outweigh the rest. A field of view of 35 compresses depth the way a portrait lens does, which is what makes a render read as photographed rather than rendered. With no environment map, reflection has nothing to describe, so every surface reads as paint.
The material carries the look#
A surface is four numbers, and roughness is the one that decides everything. Below 0.1 you get a mirror, around 0.3 you get lacquer, near 0.45 you get fired ceramic, past 0.7 you get chalk. Naming a band rather than a mood is the entire difference between a directed material and a default one.
const ceramic = new THREE.MeshPhysicalMaterial({
color: new THREE.Color("#E8E2D6"),
roughness: 0.42, metalness: 0,
clearcoat: 0.35, clearcoatRoughness: 0.28,
sheen: 0.45, sheenColor: new THREE.Color("#FFD2A8"),
envMapIntensity: 1.25,
});
const glass = new THREE.MeshPhysicalMaterial({
transmission: 1, thickness: 1.4, ior: 1.48,
roughness: 0.08, metalness: 0,
attenuationColor: new THREE.Color("#7FB2C9"), attenuationDistance: 0.9,
});
The glass recipe is where most generated scenes go wrong, because transmission without thickness produces a soap bubble. Thickness is what gives refraction something to bend through, and attenuationColor with attenuationDistance is what makes a thick edge read as tinted while a thin one stays clear. That pair is the difference between glass and a transparent grey.
Name one material and one accent, never three of each. A scene with a single directed surface and a single light temperature contrast looks composed; a scene with chrome, glass and ceramic in it looks like a demo of the material system, which is the same failure as a page that shows off every effect at once.
Bind the camera to the scroll, not to a clock#
A mesh that spins forever is the clearest tell that nobody directed the scene. Motion that responds to the reader is a different category of thing, and with GSAP it is one timeline.
gsap.registerPlugin(ScrollTrigger);
gsap.timeline({ scrollTrigger: {
trigger: "#stage", start: "top top", end: "+=220%",
pin: true, scrub: 1, anticipatePin: 1, invalidateOnRefresh: true,
}})
.to(camera.position, { z: 3.1, y: 0.22, duration: 1, ease: "none" })
.to(model.rotation, { y: 0.62, duration: 1, ease: "none" }, "<")
.to(camera.position, { x: 1.35, duration: 1, ease: "none" });Four of those values are decisions, not boilerplate. end: "+=220%" sets the scroll budget: the stage stays pinned for two and a fifth viewport heights, which is enough for two camera moves and not enough to feel like a hostage situation. scrub: 1 gives the timeline a one second catch-up, so the camera glides instead of snapping to every wheel tick. ease: "none" is mandatory under a scrub, because the reader's scroll is already the easing curve and a second one fights it. And anticipatePin: 1 removes the single frame jump that makes a pinned section feel broken on a trackpad.
For the render loop itself, damp towards a target rather than assigning to it: THREE.MathUtils.damp(current, target, 4, delta) at a lambda of 4 settles in roughly 300ms and is frame rate independent, which matters the moment someone opens the page on a 120Hz phone.
Wrap the whole thing in gsap.matchMedia() keyed to (prefers-reduced-motion: no-preference) and give the reduced branch the final camera position with no animation. That is two extra lines and it is the difference between a scene that respects an accessibility setting and one that ignores it.
The verdict, and the block to paste#
Hand Claude the values and the default scene never appears. Hand it adjectives and you get the torus knot, every time, because the model is not failing at design: it is filling gaps in the brief with the most average answer available.
Copy this into the prompt, change the numbers to your own, and the output is directed rather than generic:
- Camera
35degrees, at(0, 0.45, 5.2), target(0, 0.15, 0), no orbit controls. ACESFilmicToneMappingat exposure1.12, DPR clamped to2.RoomEnvironmentthroughPMREMGenerator,envMapIntensity: 1.25.- Key
DirectionalLight2.6at(3.2, 4.8, 2.1), coolAmbientLight0.35. Soft shadows,normalBias 0.02. - One material:
roughness: 0.42,metalness: 0,clearcoat: 0.35. - Ground
#0B0B0F,FogExp2at0.085, no gradient anywhere. - Camera on
ScrollTrigger, pinned+=220%,scrub: 1,ease: "none", reduced motion branch holds the end frame.
That is the shape of every brief worth pasting, and writing one per project is the part that takes longer than the build. Thanor is a subscription library of exactly these briefs: full art direction per design, written as values, for three dimensional scenes, animated backgrounds, page sections and app interfaces. $89 for three months, $189 a year, $299 once, and free accounts open the designs marked free.
Browse the Thanor library and read one scene's direction end to end, then compare it to the brief you would have written yourself. If you want the prompt itself rather than the scene anatomy, the companion piece is a 3D website AI prompt broken into its seven fields, and the reason defaults converge at all is set out in the pillar.
Questions this raises
Can Claude actually build a 3D website?
Yes, and in one pass. It writes the Three.js or React Three Fiber setup, the render loop, the resize handler and the scroll binding without stalling. What it cannot do is invent your focal length, roughness value or scroll distance, so it uses the library defaults unless the brief names them.
Which 3D library should the prompt name?
Three.js directly for a single hero object, because the whole scene is about eighty lines and nothing else needs to know about it. React Three Fiber with drei when the page is already React and the scene has to react to component state.
How heavy can a hero scene be?
Keep the visible triangle count under 150,000, textures at 2048px on the long edge, and the glb compressed with Draco or meshopt. Clamp the device pixel ratio to 2 and a mid range phone holds 60fps.
Does a 3D hero slow the page down?
Not when the render loop is paused off screen by an IntersectionObserver, the canvas has explicit width and height so nothing shifts, and the model loads after first paint. The cost is bandwidth, and a compressed glb of one object is smaller than most hero videos.
Part of
Build it with AI
Three dimensional scenes, scroll driven motion and animated pages, built by handing a model a specification instead of an adjective. What to write, and what comes back.
