Moonshot AI has introduced PerceptionBench, a benchmark that tests the visual perception of multimodal language models in isolation from logical reasoning – and the results are sobering: no leading AI model breaks the 60-percent accuracy barrier on basic vision tasks.
Key Facts
- PerceptionBench breaks visual perception into ten atomic sub-competencies instead of mixing perception, knowledge, and reasoning in a single task
- Top models reach a maximum of 59.7 percent accuracy: GPT-5.6 Sol leads narrowly ahead of Kimi K3 (58.5 %) and Claude Fable 5 (57.2 %)
- Error analysis reveals: many supposed logic errors originate at the image-reading stage, not during reasoning
- From over 17,000 verified questions, 3,000 tasks were published, based on real model failure patterns
- Open-source models lag significantly behind, with GLM-4.6V achieving only 32.5 %
How the Benchmark Works
The team behind the Chinese AI assistant Kimi developed PerceptionBench differently from conventional tests: instead of defining categories upfront, the authors analyzed real error patterns from 42 existing open-source benchmarks. These overlap minimally – each covers only a different subset of visual weaknesses.
This yielded ten capability areas: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Each question is answerable by looking alone – without external knowledge or reasoning.
The tasks appear trivial at first glance: determining the clock position of a symbol on a dial, counting flowers in a red box, or deciding which of two pencil holders shows a gray-pink combination. Yet here the limits become apparent.
Results: Frontier Models Fall Short of 60 Percent
| Model | Accuracy |
|---|---|
| GPT-5.6 Sol | 59.7 % |
| Kimi K3 | 58.5 % |
| Claude Fable 5 | 57.2 % |
| Gemini 3.1 Pro | 56.2 % |
| GPT-5.5 | 55.8 % |
| Qwen3.5-397B-A17B | 47.5 % |
| GLM-4.6V | 32.5 % |
Between the five best systems lie only roughly four percentage points – suggesting that vision is a shared weakness across all frontier models. Open-source models drop off significantly.
What This Means
The study confirms what recent research has hinted at: even the best models fail at basic vision tasks. This has practical consequences. When a model misreads a table or confuses objects, it is not a reasoning error – the error originates at the image-reading stage itself.
For organizations deploying multimodal AI in practice – whether in document processing, quality control, or medical image analysis – this signals caution: do not assume models see what you see. Validate visual outputs critically, especially in safety-critical applications. Accuracy on "simple" visual tasks is significantly lower than on language tasks – an important factor when selecting and calibrating systems.
Sources
Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.




