Google DeepMind 发布 Gemma 4 12B:无编码器多模态模型,面向笔记本电脑
原标题:Introducing Gemma 4 12B: a unified, encoder-free multimodal model
AI 摘要
Google DeepMind 发布了 Gemma 4 12B,这是一款面向笔记本电脑的无编码器多模态模型,支持原生音频输入,性能接近其 26B MoE 模型,但内存占用更小。该模型采用统一架构,视觉和音频输入直接进入 LLM 主干,无需传统编码器,并支持多 token 预测以降低延迟。模型以 Apache 2.0 许可开源,可在 16GB 内存的设备上本地运行,并已集成到多个开发工具中。
正文节选
Introducing Gemma 4 12B: a unified, encoder-free multimodal model Today, we are introducing Gemma 4 12B, our latest model designed to bring agentic multimodal intelligence directly to laptops. Bridging the gap between our edge-friendly E4B and our more advanced 26B Mixture of Experts (MoE), Gemma 4 12B packages powerful capabilities inside a reduced memory footprint. It is also our first mid-sized model to feature native audio inputs. Thanks to the developer community, Gemma 4 models have now cr