Qwen3.8 Max追平Claude Opus 4.8,但Kimi K3更优且成本低25%
原标题:Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less
AI 摘要
阿里巴巴的Qwen3.8 Max在Artificial Analysis Intelligence Index上得分56,追平Claude Opus 4.8,但落后于Kimi K3(57分),且Kimi K3成本低25%。在GDPval-AA基准上,Qwen3.8 Max得分超过Kimi K3,但需要更多步骤和输入token,导致成本更高。此外,该模型在AA-LCR和AA-Omniscience基准上出现性能倒退,幻觉率从23%升至40%。
正文节选
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper. On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (