Qwen3.6-35B-A3B beats Claude Opus 4.7 on pelican test
A new comparison of model outputs found that Qwen3.6-35B-A3B running locally on a MacBook Pro M5 produced a better pelican illustration than Anthropic’s Claude Opus 4.7. The test also included a flamingo-on-a-unicycle SVG, which again went to Qwen, according to the article’s author.
A local Qwen model running on a MacBook Pro has outperformed Claude Opus 4.7 on a small but telling visual test. Simon Willison reported that Qwen3.6-35B-A3B, run on his laptop, produced a better pelican illustration than Anthropic’s flagship model in a side-by-side comparison.
The comparison was not a formal benchmark, but it was designed to test a task that many text-only evaluations miss: whether a model can generate usable, appealing SVG graphics. SVG, or Scalable Vector Graphics, is a text-based image format that is often used for icons, diagrams, and simple illustrations because it stays sharp at any size.
Willison said the pelican test gave the edge to Qwen3.6-35B-A3B. He then ran a second prompt, asking for a flamingo on a unicycle in SVG form, and Qwen again came out ahead in his judgment. The result is notable because the Qwen model was not being hosted on a large cloud service, but locally on his own MacBook Pro M5.
Qwen3.6-35B-A3B is part of Alibaba’s Qwen family of models. The name indicates a 35-billion-parameter model with a three-billion-parameter active component, a structure that aims to balance capability with lower compute cost at inference time. In practical terms, that means a model can be smaller to run for a given request while still drawing on a much larger pool of learned weights.
Running a model locally matters for anyone who wants more control over latency, privacy, and cost. It also changes the way people evaluate models, because the hardware environment becomes part of the story. A model that looks expensive or heavyweight in the cloud may be surprisingly capable when tuned for a laptop-class system.
The comparison also fits a broader pattern in current AI tooling, where text generation, coding help, image generation, and structured outputs are increasingly judged task by task rather than by a single leaderboard. A model can be excellent at one narrow job and mediocre at another, which is why many developers now test with the exact use case they care about instead of relying only on general-purpose benchmarks.
Willison’s example is especially relevant because it uses a creative output that is easy for humans to inspect. That makes the difference visible without specialized evaluation software, and it highlights a common problem in model selection: benchmark scores do not always predict which system will produce the better practical result for a specific prompt.
The setup also hints at the growing role of local inference tools such as LM Studio, which can make it easier to run open models on personal hardware. As more capable models become available in formats that fit consumer machines, the gap between cloud-first AI and local AI keeps narrowing in areas where the prompt, the format, and the renderer all matter.
In this case, the laptop-generated pelican was the better bird.