[{"data":1,"prerenderedAt":30},["ShallowReactive",2],{"nr-en-perception-bench-multimodal-ai-vision-limits":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":22,"trackerLabel":23,"headlineStat":24,"image":25,"ogImage":26,"imageAlt":5,"csv":10,"minutes":27,"words":28,"html":29},"perception-bench-multimodal-ai-vision-limits","PerceptionBench: Multimodal AI Models See Far Worse Than Expected","A new benchmark from Moonshot AI shows: no frontier model reaches 60 percent accuracy on pure visual perception tasks. Many supposed reasoning errors originate at the image-reading stage.","2026-08-15","09:02","2026-08-15T09:02:00+02:00","","August 15, 2026","daten","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"Multimodal AI","Benchmark","Visual Perception","LLM Evaluation","Moonshot AI","\u002Fstand-der-ki","AI Progress","No frontier model reaches 60 % accuracy on visual perception","\u002Fnewsroom\u002Fimg\u002Fperception-bench-multimodal-ai-vision-limits.webp","\u002Fog-nr\u002Fperception-bench-multimodal-ai-vision-limits.en.png",2,480,"\u003Cp>Moonshot AI has introduced \u003Cstrong>PerceptionBench\u003C\u002Fstrong>, a benchmark that tests the visual perception of multimodal language models in isolation from logical reasoning – and the results are sobering: no leading AI model breaks the 60-percent accuracy barrier on basic vision tasks.\u003C\u002Fp>\n\u003Ch2>Key Facts\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>PerceptionBench\u003C\u002Fstrong> breaks visual perception into ten atomic sub-competencies instead of mixing perception, knowledge, and reasoning in a single task\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Top models reach a maximum of 59.7 percent accuracy\u003C\u002Fstrong>: GPT-5.6 Sol leads narrowly ahead of Kimi K3 (58.5 %) and Claude Fable 5 (57.2 %)\u003C\u002Fli>\n\u003Cli>Error analysis reveals: many supposed \u003Cstrong>logic errors originate at the image-reading stage\u003C\u002Fstrong>, not during reasoning\u003C\u002Fli>\n\u003Cli>From over \u003Cstrong>17,000 verified questions\u003C\u002Fstrong>, 3,000 tasks were published, based on real model failure patterns\u003C\u002Fli>\n\u003Cli>Open-source models lag significantly behind, with GLM-4.6V achieving only 32.5 %\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>How the Benchmark Works\u003C\u002Fh2>\n\u003Cp>The team behind the Chinese AI assistant Kimi developed PerceptionBench differently from conventional tests: instead of defining categories upfront, the authors analyzed real error patterns from 42 existing open-source benchmarks. These overlap minimally – each covers only a different subset of visual weaknesses.\u003C\u002Fp>\n\u003Cp>This yielded ten \u003Cstrong>capability areas\u003C\u002Fstrong>: Visual Relation, Counting, Attribute, Depth &amp; 3D, Localization, Comparison, Fine-grained Recognition, Context Integration, OCR, and Hallucination. Each question is answerable by looking alone – without external knowledge or reasoning.\u003C\u002Fp>\n\u003Cp>The tasks appear trivial at first glance: determining the clock position of a symbol on a dial, counting flowers in a red box, or deciding which of two pencil holders shows a gray-pink combination. Yet here the limits become apparent.\u003C\u002Fp>\n\u003Ch2>Results: Frontier Models Fall Short of 60 Percent\u003C\u002Fh2>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Model\u003C\u002Fth>\n\u003Cth>Accuracy\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>GPT-5.6 Sol\u003C\u002Ftd>\n\u003Ctd>59.7 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Kimi K3\u003C\u002Ftd>\n\u003Ctd>58.5 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Claude Fable 5\u003C\u002Ftd>\n\u003Ctd>57.2 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Gemini 3.1 Pro\u003C\u002Ftd>\n\u003Ctd>56.2 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>GPT-5.5\u003C\u002Ftd>\n\u003Ctd>55.8 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Qwen3.5-397B-A17B\u003C\u002Ftd>\n\u003Ctd>47.5 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>GLM-4.6V\u003C\u002Ftd>\n\u003Ctd>32.5 %\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Cp>Between the five best systems lie only roughly four percentage points – suggesting that vision is a shared weakness across all frontier models. Open-source models drop off significantly.\u003C\u002Fp>\n\u003Ch2>What This Means\u003C\u002Fh2>\n\u003Cp>The study confirms what recent research has hinted at: \u003Cstrong>even the best models fail at basic vision tasks\u003C\u002Fstrong>. This has practical consequences. When a model misreads a table or confuses objects, it is not a reasoning error – the error originates at the image-reading stage itself.\u003C\u002Fp>\n\u003Cp>For organizations deploying multimodal AI in practice – whether in document processing, quality control, or medical image analysis – this signals caution: do not assume models see what you see. Validate visual outputs critically, especially in safety-critical applications. Accuracy on &quot;simple&quot; visual tasks is significantly lower than on language tasks – an important factor when selecting and calibrating systems.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.de\u002Fneuer-benchmark-bestaetigt-ki-modelle-sehen-weiterhin-schlecht\u002F\">The Decoder (DE)\u003C\u002Fa>\u003C\u002Fli>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fthe-decoder.com\u002Fnew-benchmark-confirms-ai-models-still-perform-poorly-at-visual-perception\u002F\">The Decoder\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1786777672891]