DataMultimodal AIBenchmarkVisual Perception

PerceptionBench: Multimodal AI Models See Far Worse Than Expected

A new benchmark from Moonshot AI shows: no frontier model reaches 60 percent accuracy on pure visual perception tasks. Many supposed reasoning errors originate at the image-reading stage.

No frontier model reaches 60 % accuracy on visual perception

PerceptionBench: Multimodal AI Models See Far Worse Than Expected

Moonshot AI has introduced PerceptionBench, a benchmark that tests the visual perception of multimodal language models in isolation from logical reasoning – and the results are sobering: no leading AI model breaks the 60-percent accuracy barrier on basic vision tasks.

Key Facts

  • PerceptionBench breaks visual perception into ten atomic sub-competencies instead of mixing perception, knowledge, and reasoning in a single task
  • Top models reach a maximum of 59.7 percent accuracy: GPT-5.6 Sol leads narrowly ahead of Kimi K3 (58.5 %) and Claude Fable 5 (57.2 %)
  • Error analysis reveals: many supposed logic errors originate at the image-reading stage, not during reasoning
  • From over 17,000 verified questions, 3,000 tasks were published, based on real model failure patterns
  • Open-source models lag significantly behind, with GLM-4.6V achieving only 32.5 %

How the Benchmark Works

The team behind the Chinese AI assistant Kimi developed PerceptionBench differently from conventional tests: instead of defining categories upfront, the authors analyzed real error patterns from 42 existing open-source benchmarks. These overlap minimally – each covers only a different subset of visual weaknesses.

This yielded ten capability areas: Visual Relation, Counting, Attribute, Depth & 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Each question is answerable by looking alone – without external knowledge or reasoning.

The tasks appear trivial at first glance: determining the clock position of a symbol on a dial, counting flowers in a red box, or deciding which of two pencil holders shows a gray-pink combination. Yet here the limits become apparent.

Results: Frontier Models Fall Short of 60 Percent

Model Accuracy
GPT-5.6 Sol 59.7 %
Kimi K3 58.5 %
Claude Fable 5 57.2 %
Gemini 3.1 Pro 56.2 %
GPT-5.5 55.8 %
Qwen3.5-397B-A17B 47.5 %
GLM-4.6V 32.5 %

Between the five best systems lie only roughly four percentage points – suggesting that vision is a shared weakness across all frontier models. Open-source models drop off significantly.

What This Means

The study confirms what recent research has hinted at: even the best models fail at basic vision tasks. This has practical consequences. When a model misreads a table or confuses objects, it is not a reasoning error – the error originates at the image-reading stage itself.

For organizations deploying multimodal AI in practice – whether in document processing, quality control, or medical image analysis – this signals caution: do not assume models see what you see. Validate visual outputs critically, especially in safety-critical applications. Accuracy on "simple" visual tasks is significantly lower than on language tasks – an important factor when selecting and calibrating systems.

Sources

Editorially owned by Ideal Syka. Sources and method: Newsroom & method. Tips and corrections: ai@i6eal.de.

Share
← All articles

All analyses are based on i6eal's own measurements or on clearly labelled sources. Figures are snapshots and may change; corrections are disclosed transparently.