Mistral Large 4
Simon Willison compares Mistral Large 4 with Claude, GPT, and Gemini on a whimsical SVG prompt.
Simon Willison commented on a Hacker News thread about Mistral Large 4, responding to a claim that standard benchmarks are saturated. Using his llm CLI at default reasoning levels, he asked Claude Opus 5.5, GPT-6.1 Sol, Gemini 3.8 Flash, and mistral/mistral-large-4 to generate the same SVG of an armadillo in fishnet tights jaywalking on Mars. The post is a light model comparison and does not include release specs, parameter counts, or benchmark scores.
- Comment is about the Mistral Large 4 model on Hacker News
- Compared with Claude Opus 5.5, GPT-6.1 Sol, and Gemini 3.8 Flash
- Each model is asked to generate the same whimsical SVG
- Post argues ordinary benchmarks are already saturated
My comment on Mistral Large 4 — Hacker News. wren6991 : The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars. OK well I couldn't resist this one: llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars' Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
This source does not provide full text. Read it at simonwillison.net.