2026-10-02 · ← News
16,200 questions drew the map hidden inside a language model
A simple evaluation asked a model 16,200 times whether a coordinate was on land or water, then assembled the answers into a world map. The result vividly exposes geographic knowledge in the weights, but without a score, baseline and full method it is better treated as an X-ray than a leaderboard.
The image could not be loaded.
A simple evaluation asked a model 16,200 times whether a coordinate was on land or water, then assembled the answers into a world map. The result vividly exposes geographic knowledge in the weights, but without a score, baseline and full method it is better treated as an X-ray than a leaderboard.
Text coordinates turned into the outline of continents
Andrej Karpathy highlighted an experiment that gives an LLM latitude and longitude as text and requests the binary answer “Land” or “Water.” After 16,200 queries, the answers are plotted as an image. Recognizable continents appear in the shared result, showing that the model connects numeric coordinates with the coarse shape of Earth.
The public post does not provide the complete prompt, point spacing, temperature, repeated runs or aggregate accuracy. It points to results for Claude models, but the short post itself contains no table suitable for a reliable model comparison. What we can see is a striking image rather than a fully specified benchmark.
The map in the weights came from the internet's textual traces
The model does not open a mapping service during this test. It answers from relationships learned during training, where coordinates appear beside place names, map records, weather reports, travel writing and other geographic text. The experiment turns dispersed statistical memory into an image a person can read in a second.
Earlier GeoLLM research showed that language models contain substantial geographic information. It also found that bare coordinates can be a weak interface for harder predictions and that context from OpenStreetMap improves performance significantly. The newer GPSBench, with 57,800 examples across 17 tasks, describes a related boundary: models handle coarse geography better than exact city localization and complex coordinate calculations.
A beautiful map can hide an easy baseline and regional errors
An outline of the continents does not reveal how many points were classified correctly. The sampling geometry, share of water, grid resolution and policy for coastal cells can all reshape the result. Without simple baselines, such as always answering “Water,” and without regional metrics, we cannot tell whether a model knows smaller islands and coastlines or merely sketches a rough silhouette.
The test also does not measure navigation or geometric reasoning. A model may recall that a pair of numbers is associated with land while failing to calculate distance, direction or the location of underrepresented regions. A compelling visualization can smooth over those distinctions.
Open data and an error map can turn the demo into a benchmark
The next step is to publish all 16,200 points, the exact prompt, each model's answers and the ground truth. Alongside overall accuracy, the evaluation should report balanced accuracy for land and water, performance by distance from the coast and a regional error map.
Stability will be the most revealing test. If the continents dissolve after a small change in coordinate format or wording, the experiment found a brittle association. If the pattern survives different formats, models and repeated runs, it becomes a cheap evaluation of how much of the world an LLM actually carries in its weights.
Lilith's verdict
The model received 16,200 pins and placed them from memory until continents appeared on the globe. Now we need to rotate it and count how many pins landed in the sea.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗