Draw the page first, then let the agent build it
2026-08-18
Complex layouts were slow to specify in words: photos, captions, thin rules and spacing take paragraphs to describe and still come out wrong. A new skill, layout-from-image, is now installed on every agent on the platform, and it takes the opposite route. You give the agent a picture of the page — drawn by any image model — and it does the rest itself: it cuts the photographs out of the mock-up, writes the HTML, opens the result in a browser, compares its own render against your picture and keeps fixing the differences until they close. Nothing to switch on: the agent reaches for the skill whenever a task contains an image of a page. The comparison is the part that makes it work — the agent measures how far its render is from the mock-up and looks at the overlay, so corrections stop being guesswork.