Bullish

Claude Code Costs 380% More Than Hermes Agent on Kimi K3 Tasks

2026-08-01 14:48:36

Composio benchmark reveals Claude Code is 3.8x costlier and 2.2x slower than Hermes using Kimi K3, despite similar task completion rates across six agent frameworks.

Woofun AI reports that AI infrastructure firm Composio conducted a comparative analysis of six agent frameworks utilizing the Kimi K3 model to execute 26 identical tasks. The evaluation measured performance metrics including task completion volume, average cost per task, and median execution time.

Kimi Code achieved the highest completion rate with 21 tasks, while Hermes completed 20. Pi Agent and Claude Code both finished 19 tasks, followed by OpenCode with 18 and Codex with 17. Cost analysis indicates Hermes and Pi Agent incurred the lowest average expenses at $0.39 and $0.40 per task respectively. In contrast, Claude Code averaged $1.47 per task, representing a 3.8-fold increase over Hermes. Regarding speed, Pi Agent recorded the shortest median completion time of 161.7 seconds, whereas Claude Code was the slowest at 347.6 seconds, exceeding the fastest framework by more than double.

WOOFUN AI

Impact Assessment · Quick Read

This benchmark highlights significant efficiency disparities among agent frameworks even when controlling for the underlying LLM. The substantial cost and latency differences suggest that framework optimization is a critical variable in AI agent economics. Developers may need to prioritize framework selection over model choice to optimize for cost-performance ratios in production environments.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions