Key Takeaways

- Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, matching Claude Opus 4.8
- Kimi K3 outperforms both at 57 points while costing 25% less per task
- Qwen3.8 Max's hallucination rate jumped from 23% to 40% compared to the previous version
Alibaba's Qwen3.8 Max just matched Claude Opus 4.8 on the Artificial Analysis Intelligence Index, but Moonshot AI's Kimi K3 beats them both while running 25% cheaper. For teams evaluating LLM APIs, the latest benchmark data forces a harder question: when does raw capability matter more than cost efficiency?
How do the benchmark scores compare?
Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump from Qwen3.7 Max's 46. That puts it level with Claude Opus 4.8 and ahead of GLM-5.2 at 51. Kimi K3 sits one point higher at 57.
The picture shifts on GDPval-AA, a benchmark focused on work-related tasks. Here Qwen3.8 Max leaps 468 Elo points to 1,739, passing Kimi K3 at 1,685. Only Claude Opus 5 scores higher with 1,852.

Two different benchmarks, two different winners. Teams building agentic workflows may weight GDPval-AA more heavily. Teams prioritizing general reasoning might favor the Intelligence Index.
What does each model actually cost?
Alibaba cut Qwen's token prices. Input dropped from $2.50 to $2.00 per million tokens. Output fell from $7.50 to $6.00. Cache hits went from $0.50 to $0.25.
Those cuts don't translate to lower task costs. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends full conversation history at each step. A single task on the Intelligence Index now costs $1.14, more than double Qwen3.7 Max's $0.53.
| Model | Intelligence Index Score | GDPval-AA (Elo) | Cost per Task | Steps per Task |
|---|---|---|---|---|
| Kimi K3 | 57 | 1,685 | $0.86 | — |
| Qwen3.8 Max | 56 | 1,739 | $1.14 | 64 |
| Claude Opus 4.8 | 56 | — | — | — |
| GLM-5.2 | 51 | — | $0.57 | — |
| Qwen3.7 Max | 46 | 1,271 | $0.53 | 14 |
| Claude Opus 5 | — | 1,852 | — | — |
Kimi K3 delivers the best price-performance ratio: one point higher than Qwen3.8 Max at $0.86 per task. GLM-5.2 is the budget option at $0.57, though it trails by 5 points on the Intelligence Index.
Where does Qwen3.8 Max regress?
Higher benchmark scores came with tradeoffs. AA-LCR, which tests whether a model correctly synthesizes information from long texts, dropped 2 points. More concerning is AA-Omniscience, which fell 10 points.
AA-Omniscience measures whether a model answers knowledge questions correctly or admits it doesn't know. Accuracy stayed around 31%, but the hallucination rate jumped from 23% to 40%. Qwen3.8 Max guesses far more often instead of declining to answer.
For production systems where reliability matters more than capability, that hallucination spike is disqualifying. A model that confidently invents answers breaks trust faster than one that says "I don't know."
Logicity's Take
The benchmark race is hitting diminishing returns. Qwen3.8 Max gains 10 points over its predecessor but costs twice as much per task and hallucinates nearly twice as often. Kimi K3 looks like the pragmatic choice for teams who need strong general reasoning without budget blowouts. But if your use case involves complex multi-step work tasks, Qwen3.8 Max's GDPval-AA lead is real. Know which benchmark matches your workload before picking a winner.
Which model fits which use case?
The data points to three distinct positioning plays.
Kimi K3 suits teams running high-volume inference where cost per call compounds fast. One point of benchmark difference rarely shows up in production, but 25% cost savings does.
Qwen3.8 Max makes sense for complex agentic workflows where GDPval-AA performance matters. If your agents are executing multi-step work tasks and accuracy on those tasks drives business value, the 54-point Elo gap over Kimi K3 is significant.
Claude Opus 5 remains the ceiling on work-task benchmarks at 1,852 Elo. Teams building enterprise products where they can pass API costs to customers may still prefer the Claude stack for its reliability reputation, even without published task costs.
The hidden cost of thoroughness
Qwen3.8 Max's 64-step approach reveals a strategic bet: Alibaba is optimizing for capability scores, not inference efficiency. The model works more thoroughly but runs slower. For latency-sensitive applications, that's a problem.
The 15x increase in input tokens also means context window costs scale poorly. Long-running conversations or agents that maintain extensive state will burn through token budgets faster than the headline price cuts suggest.
Moonshot AI appears to have found a better balance. Kimi K3 achieves higher Intelligence Index scores with fewer steps, suggesting architectural differences in how the models approach reasoning tasks.
Frequently Asked Questions
Is Qwen3.8 Max better than Claude Opus 4.8?
They score identically at 56 on the Artificial Analysis Intelligence Index. Qwen3.8 Max costs more per task but has published pricing, while Anthropic hasn't disclosed equivalent task costs.
Why did Qwen3.8 Max's hallucination rate increase?
The model guesses more often instead of admitting it doesn't know. Accuracy stayed at 31%, but wrong confident answers jumped from 23% to 40%.
Which model has the best price-performance ratio?
Kimi K3 scores highest (57) at the lowest task cost ($0.86). GLM-5.2 is cheaper at $0.57 but scores 6 points lower.
Need Help Implementing This?
Evaluating LLM providers for your product? Our team helps AI builders navigate model selection, cost optimization, and production deployment. Reach out at logicity.in/consulting.
Source: The Decoder / Maximilian Schreiner
Huma Shazia
Senior AI & Tech Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




