Qwen3.8 27B addition in words
Simon Willison locally retests Qwen3.8 27B on summing numbers and answering in words.
Simon Willison reproduced an experiment, originally run by Colin Frasier on GPT-4o, that asks a model to add numbers and return the sum in words. He ran it locally on an NVIDIA DGX Spark against the quantized GGUF Qwen3.8-27B-Q4_K_M, using a Codex Remote session driven by GPT-6 Astra. Willison noted the earlier GPT-4o run made many arithmetic errors, which he took as evidence it was not using a calculator. The post refers to a 30-run result, but the excerpt ends before any scores are given.
- Colin Frasier previously tested GPT-4o on addition answered in words.
- Willison repeated the test locally on an NVIDIA DGX Spark.
- The model under test was Qwen3.8-27B-Q4_K_M.gguf.
- GPT-6 Astra in Codex Remote was used to run the experiment.
- The excerpt ends before reporting the 30-run scores.
Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment. I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf . Here's the result for a run of 30…
This source does not provide full text. Read it at simonwillison.net.