A trilingual site, live on its own domain, designed and built by me and an AI, from the first commit to the text of every case. Eighteen working sessions, all recorded as they happened.
I started out assuming the new skill would be writing good prompts. It was not.
It read the colour of the pixel, and got half of it right
I needed the design system palette and all I had was a screenshot. The labels carrying the codes were text too small to read: on the image, smudges. Instead of trying to decipher them, the AI read the colour of the pixel in each swatch and returned the whole scale.
One of the steps came out #3381ff, exactly the value I already knew. I took
the match as proof that the method was right.
It was not. Checked later against the real source, the brand colour was exact across all twenty steps. The neutral scale was not: it runs to 1100, and I had recorded it stopping at 1000, with a step that does not exist shifting everything else along.
The method confirms a ramp whose structure you already know. It does not discover the structure. I got it right where I already knew the answer, and that is exactly where I thought I had proved something.
The label is illegible at the source and the code comes out exact on the other side, because what was read was not the text: it was the colour of the pixel. This scale came out right across all twenty steps. The neutral one did not.
Twelve good-looking images, rejected by one measurement
I wanted the cat in the mark to walk across the screen. I asked the AI for the poses, and they came out beautiful.
Then I measured. In a real walk the paw travels 35% to 45% of the body between one frame and the next. Across the twelve generated images, travel sat between 5% and 16%. At the size the cat actually appears on the site, that difference is three pixels: they were twelve copies of the same image.
Nothing about them looked wrong, and that is the point. What separated "pretty" from "works" was writing the acceptance criterion as a number before ordering the material.
None of these questions has a technical answer
Choosing the problem. Which cases go in, what each one has to prove, what stays out.
Seeing what is wrong on screen. A case title cut mid-word. A cover that did not talk to its neighbour. A claim about parking that only held for the example photo. None of it breaks the build.
Answering what is in no file anywhere. One of the cases looked like work done for a real company, and no reading of the material would settle it: only I knew the brief had come from a bootcamp. That answer became the spine of the whole case.
A defect no test catches: the period never gave up space and swallowed the case name. On one of the measured screens, one pixel was left for the title.
The part that decides did not get smaller
One work case is still unwritten, and this one covers a project that is still happening. The log stays open.
What can already be said: execution really did get faster. The part that decides whether the result is any good is still mine, and it did not get smaller. It got more visible.
