2026-07-22 · ← Radar
The pelican benchmark shows how easily AI metrics become folklore
Dylan Castillo turned the pelicanmaxxing joke into a small but useful experiment: 8 animals, 6 vehicles, 48 prompts, 3 runs per prompt and 7 models. Simon Willison highlighted it because it moves an informal benchmark closer to measurement.
The pelican got 144 trials instead of one gallery impression
The question was whether AI labs had started tuning image models for Willison’s well known test of a pelican riding a bicycle. Castillo compared not just pelicans and bicycles, but other animals and vehicles too.
He tested GPT-5.6 Terra, Claude Sonnet 5, Gemini 3.5 Flash, Grok 4.5, Qwen3.7-Max, GLM-5.2 and DeepSeek V4 Pro. GPT-5.6 Luna and Gemini 3.1 Flash-Lite helped evaluate the outputs.
The result is cautious: in this test set, pelicans, bicycles and their combination did not look suspiciously better than comparable cases. GLM-5.2 produced the most interesting individual sample, but the effect was small and not significant.
A meme test becomes a lesson in what evals actually measure
For developers and product teams, the point is larger than the bird. Once a benchmark becomes a meme, it starts measuring model capability, market attention, training data and the vendor reflex to optimize for visible tests.
Castillo’s method is the healthier pattern: do not trust one iconic prompt, expand the case space and ask whether one cell outperforms what its parts predict. That is the difference between a fun demo and an eval you can use to make a decision.
A careful mini benchmark is still not an audit of the labs
The limit is clear. 48 prompts and 7 models produce a better signal than casual spot checks, but they do not prove that no lab has ever seen a similar test during training. Models can also fail in visual ways that human or LLM assisted grading may miss.
So the honest reading is negative and narrow: this dataset does not show a suspicious boost. That is less dramatic than a story about secret tuning, but much more useful.
The next signal will come from benchmarks that are harder to memorize
The thing to watch is whether the community builds more parametric tests like this. Fewer famous images, more prompt grids, more hidden variants and regular refreshes.
If an eval fits in one viral screenshot, it eventually becomes part of the game. Good measurement has to be more boring than the meme, or models and people will learn to play it.
Lilith's verdict
The bicycle pelican is a useful canary, but a terrible judge. Once marketing knows the benchmark, the cage needs more birds.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗