Bullish

PerceptionBench Open-Sourced; GPT-5.6-Sol Leads with 59.7% Accuracy, No Model Surpasses 60%

2026-08-02 23:17:19

Kimi releases PerceptionBench benchmark testing 10 atomic visual skills. GPT-5.6-Sol tops 16 models at 59.7%, yet none exceed 60% accuracy, highlighting persistent hallucination issues.

Woofun AI reports that Kimi has open-sourced PerceptionBench, a multimodal visual perception evaluation benchmark. The dataset comprises 3,000 manually verified questions derived from failure cases across 42 existing sets, isolating 10 atomic capabilities including visual relationships, counting, and hallucination detection without requiring reasoning. Among 16 leading multimodal models evaluated, no system achieved an overall accuracy above 60%. GPT-5.6-Sol secured the top position with 59.7% accuracy, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%. The findings indicate that visual hallucination remains the most significant weakness across all tested architectures.

WOOFUN AI

Impact Assessment · Quick Read

The release of PerceptionBench provides a granular view of current multimodal limitations, specifically in atomic visual tasks. With top models failing to breach the 60% accuracy threshold, it suggests that fundamental perception gaps persist despite advances in general reasoning. This benchmark may drive future model iterations to prioritize robustness against hallucinations over broader capability expansion.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions