Cerebras 以最高 750 token/秒服务 GPT-5.6 Sol
原标题:How Cerebras serves GPT-5.6 Sol at up to 750 tokens per second
AI 摘要
Cerebras 宣布为 OpenAI 的 GPT-5.6 Sol 提供 Ultrafast 模式,输出速度最高达每秒 750 个 token。该服务运行在 Cerebras WSE-3 晶圆级芯片上,模型架构、权重、精度、上下文配置和推理设置与标准 OpenAI 端点完全一致,未做蒸馏或量化。Cerebras 称在 GDP-Val 任务上完成速度比标准端点快 5.6 倍,在 Humanity's Last Exam 上快 6.9 倍,主要面向开发者、知识工作者和计算机使用自动化场景。
正文节选
For the last two years, AI models have become dramatically more capable. They can reason longer, write production-ready code, operate computers, use browsers, and produce professional work across science, finance, mathematics, physics, and engineering. But one part of the experience hasnʼt changed: waiting. OpenAI is previewing a new Ultrafast mode for GPT‑5.6 Sol running on Cerebras, at up to 750 output tokens per second. This speed changes how we use AI:we can now stay in the flow, and collabo