Qwen3.8-Omni-Flash 以更低价格对标 Gemini Flash 多模态能力
原标题:Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks
AI 摘要
Qwen 发布首个面向 AI 智能体的多模态模型 Qwen3.8-Omni-Flash,可同时处理音频与视频并自主调用工具完成视频剪辑、短视频翻译、电影摘要等任务,上下文窗口达一百万 token。官方称其在音视频任务上接近 Gemini 3.8 Flash 的水平,但 API 定价更低:输入每百万 token 0.15 美元、输出 0.47 美元,音频输入每小时不足 0.01 美元。该模型通过 Qwen Studio、Qwen Cloud 和 API 提供,并配套开源 Qwen-MM-Plugins 与 Qwen-Live Harness。
正文节选
Qwen3.8-Omni-Flash undercuts Google's Gemini Flash pricing while matching its multimodal benchmarks Qwen3.8-Omni-Flash is Qwen's first multimodal model built for AI agents. It processes audio and video together, draws conclusions, and uses tools on its own to edit vlogs, translate short videos, or summarize movies. The context window spans one million tokens. On audio-video tasks, Qwen says it comes close to matching Gemini 3.8 Flash. API pricing sits at $0.15 per million input tokens and $0.47