In November and December 2025, inspired by Simon Willison’s pelican-riding-a-bicycle benchmark, I had some then-current LLMs create SVGs from thirty similar prompts, such as “Generate an SVG of an octopus operating a pipe organ.” Simon mentioned that experiment on his blog [1].
Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.
Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.
[1] https://simonwillison.net/2025/Nov/25/
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
[1] https://dorrit.pairsys.ai/
reply