All posts

Qwen3.8 Max vs Kimi K3 vs Claude Opus: benchmark scores and costs

Huma ShaziaAugust 8, 2026 at 10:32 PM6 min read
Qwen3.8 Max vs Kimi K3 vs Claude Opus: benchmark scores and costs

Key Takeaways

Qwen3.8 Max vs Kimi K3 vs Claude Opus: benchmark scores and costs
Source: The Decoder
  • Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8
  • Kimi K3 outperforms both at 57 points while costing 25% less per task
  • Qwen3.8 Max's hallucination rate jumped from 23% to 40% compared to the previous version

Alibaba's Qwen3.8 Max just matched Claude Opus 4.8 on the Artificial Analysis Intelligence Index, but Moonshot AI's Kimi K3 beats them both while running 25% cheaper. For teams evaluating LLM APIs, the latest benchmark data forces a harder question: when does raw capability matter more than cost efficiency?

$1.14 vs $0.86
Cost per task: Qwen3.8 Max vs Kimi K3 on the Intelligence Index
Advertisements

How do the benchmark scores compare?

Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max's 46. That puts it level with Claude Opus 4.8 and ahead of GLM-5.2 at 51. Kimi K3 sits one point higher at 57.

The picture shifts on GDPval-AA, a benchmark focused on work-related tasks. Here Qwen3.8 Max leaps 468 Elo points to 1,739, passing Kimi K3 at 1,685. Only Claude Opus 5 scores higher with 1,852.

Qwen3.8 Max benchmark scores on the Artificial Analysis Intelligence Index compared to Claude Opus and Kimi K3
Image (Source: The Decoder)

Two different benchmarks, two different winners. Teams building agentic workflows may weight GDPval-AA more heavily. Teams prioritizing general reasoning might favor the Intelligence Index.

What does each model actually cost?

Alibaba cut Qwen's token prices. Input dropped from $2.50 to $2.00 per million tokens. Output fell from $7.50 to $6.00. Cache hits went from $0.50 to $0.25.

Those cuts don't translate to lower task costs. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends full conversation history at each step. A single task on the Intelligence Index now costs $1.14, more than double Qwen3.7 Max's $0.53.

ModelIntelligence Index ScoreGDPval-AA (Elo)Cost per TaskSteps per Task
Kimi K3571,685$0.86
Qwen3.8 Max561,739$1.1464
Claude Opus 4.856
GLM-5.251$0.57
Qwen3.7 Max461,271$0.5314
Claude Opus 51,852

Kimi K3 delivers the best price-performance ratio: one point higher than Qwen3.8 Max at $0.86 per task. GLM-5.2 is the budget option at $0.57, though it trails by 5 points on the Intelligence Index.

Where does Qwen3.8 Max regress?

Higher benchmark scores came with tradeoffs. AA-LCR, which tests whether a model correctly synthesizes information from long texts, dropped 2 points. More concerning is AA-Omniscience, which fell 10 points.

AA-Omniscience measures whether a model answers knowledge questions correctly or admits it doesn't know. Accuracy stayed around 31%, but the hallucination rate jumped from 23% to 40%. Qwen3.8 Max guesses far more often instead of declining to answer.

For production systems where reliability matters more than capability, that hallucination spike is disqualifying. A model that confidently invents answers breaks trust faster than one that says "I don't know."

ℹ️

Logicity's Take

The benchmark race is hitting diminishing returns. Qwen3.8 Max gains 10 points over its predecessor but costs twice as much per task and hallucinates nearly twice as often. Kimi K3 looks like the pragmatic choice for teams who need strong general reasoning without budget blowouts. But if your use case involves complex multi-step work tasks, Qwen3.8 Max's GDPval-AA lead is real. Know which benchmark matches your workload before picking a winner.

Advertisements

Which model fits which use case?

The data points to three distinct positioning plays.

Kimi K3 suits teams running high-volume inference where cost per call compounds fast. One point of benchmark difference rarely shows up in production, but 25% cost savings does.

Qwen3.8 Max makes sense for complex agentic workflows where GDPval-AA performance matters. If your agents are executing multi-step work tasks and accuracy on those tasks drives business value, the 54-point Elo gap over Kimi K3 is significant.

Claude Opus 5 remains the ceiling on work-task benchmarks at 1,852 Elo. Teams building enterprise products where they can pass API costs to customers may still prefer the Claude stack for its reliability reputation, even without published task costs.

The hidden cost of thoroughness

Qwen3.8 Max's 64-step approach reveals a strategic bet: Alibaba is optimizing for capability scores, not inference efficiency. The model works more thoroughly but runs slower. For latency-sensitive applications, that's a problem.

The 15x increase in input tokens also means context window costs scale poorly. Long-running conversations or agents that maintain extensive state will burn through token budgets faster than the headline price cuts suggest.

Moonshot AI appears to have found a better balance. Kimi K3 achieves higher Intelligence Index scores with fewer steps, suggesting architectural differences in how the models approach reasoning tasks.

Frequently Asked Questions

Is Qwen3.8 Max better than Claude Opus 4.8?

They score identically at 56 on the Artificial Analysis Intelligence Index. Qwen3.8 Max costs more per task but has published pricing, while Anthropic hasn't disclosed equivalent task costs.

Why did Qwen3.8 Max's hallucination rate increase?

The model guesses more often instead of admitting it doesn't know. Accuracy stayed at 31%, but wrong confident answers jumped from 23% to 40%.

Which model has the best price-performance ratio?

Kimi K3 scores highest (57) at the lowest task cost ($0.86). GLM-5.2 is cheaper at $0.57 but scores 6 points lower.

ℹ️

Need Help Implementing This?

Evaluating LLM providers for your product? Our team helps AI builders navigate model selection, cost optimization, and production deployment. Reach out at logicity.in/consulting.

Source: The Decoder / Maximilian Schreiner

H

Huma Shazia

Senior AI & Tech Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.