Lilith.
⌕
Editorial illustration: Qwen without reasoning solved only 23.57% of sums written as words
Lilith illustration · editorial remix

Qwen3.8 27B received a simple instruction: add two positive integers and return the result using English words only. Simon Willison ran the quantized Qwen3.8-27B-Q4_K_M model locally on a DGX Spark and tested 5,070 cases with reasoning disabled. It answered 1,195 correctly, or 23.57%.

The model preserved the format while losing the numbers

It still complied with the requested format in 96.17% of cases. Accuracy reached 97.04% for operands with one to three digits, then fell to 6.44% for operands with ten to thirteen digits. This does not look like random instruction failure. The model usually knew what the answer should look like, but failed to calculate it.

Willison was revisiting an older experiment by Colin Frasier with GPT-4o. Running locally gave him control over the setup and allowed him to publish the full grid of results.

Reasoning changed both the score and the cost of the test

With medium reasoning enabled, Qwen answered 167 of 169 cases correctly. That run used only one attempt for each pair of operand lengths, rather than the 30 attempts per cell used without reasoning, because each answer took much longer. The figures show a strong difference, but they are not two equally sized samples.

For people building evals, the design of the task matters more than a model leaderboard. Addition is trivial to verify with code, but forbidding a calculator forces the model to preserve intermediate state while generating text. The test exposes a gap between instruction compliance and reliable computation that aggregate benchmarks can hide.

Two misses in a small sample do not establish reliability

A score of 167 out of 169 looks compelling, but the reasoning run has thirty times fewer attempts per cell than the baseline. Willison himself notes that another run would probably place the errors elsewhere. The experiment also does not report latency or energy use, even though slower responses caused the sample to shrink.

This is one quantized model on one machine performing one narrow task. It cannot establish a general mathematical advantage for reasoning or a direct win over GPT-4o, because the two experiments did not use identical conditions.

Repeated runs will show whether reasoning sustains the gain

The useful next step is to repeat all 169 combinations several times, publish latency and separate arithmetic errors from errors in rendering numbers as words. Only then can we tell whether reasoning reliably repairs the calculation or merely raises the odds of a correct answer.

For production systems, the safer design remains simple: send arithmetic to a calculator and let the model phrase the result. This experiment shows that a polished English number can still be an error delivered with excellent diction.

Lilith's verdict

Without reasoning, the model delivered a correctly labeled box with the wrong sum inside. In production, arithmetic belongs to a calculator, not to a language model with confident diction.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗