Qwen3.8-Max Matches Opus5 Recall Rate at Half the Cost in Security Audit
Qwen3.8-Max achieved 81.3% recall on 32 vulnerabilities, tying Opus5 but at $821 cost. Inconsistency noted across runs, though DeepSeek V4 Flash remains cheaper.
Woofun AI data shows that Aikido evaluated seven models, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, GPT-5.6 Sol, Luna, and Terra, against 32 known vulnerabilities using Agent testing. Each model underwent three audit runs. Qwen3.8-Max detected 26 vulnerabilities, yielding an 81.3% recall rate equal to Opus 5, with an F1 score of 83.2%.
However, consistency varied: only 10 vulnerabilities were identified in all three runs for Qwen, compared to 19 for Opus 5 and Sol. Total costs for Qwen reached approximately $821, half that of Opus 5 and Sol, yet five times higher than DeepSeek V4 Flash.
Comments
No comments yet.