PerceptionBench Open-Sourced; GPT-5.6-Sol Leads with 59.7% Accuracy, No Model Surpasses 60%
Kimi releases PerceptionBench benchmark testing 10 atomic visual skills. GPT-5.6-Sol tops 16 models at 59.7%, yet none exceed 60% accuracy, highlighting persistent hallucination issues.
Woofun AI reports that Kimi has open-sourced PerceptionBench, a multimodal visual perception evaluation benchmark. The dataset comprises 3,000 manually verified questions derived from failure cases across 42 existing sets, isolating 10 atomic capabilities including visual relationships, counting, and hallucination detection without requiring reasoning. Among 16 leading multimodal models evaluated, no system achieved an overall accuracy above 60%. GPT-5.6-Sol secured the top position with 59.7% accuracy, followed by Kimi K3 at 58.5%, Claude-Fable-5 at 57.2%, Gemini-3.1-Pro at 56.2%, and GPT-5.5 at 55.8%. The findings indicate that visual hallucination remains the most significant weakness across all tested architectures.
Comments
No comments yet.