KVAE:多模态生成模型的新一代 Tokenizer 系列
原标题:KVAE: Family of Tokenizers for Multimodal Generative Models
AI 摘要
Kandinsky Lab 发布了 KVAE 系列 tokenizer,涵盖音频、图像和视频,专为文本条件下的潜在扩散模型生成设计。KVAE-Audio 支持 48 kHz 全频带,KVAE-3D 提供两种视频压缩方案,KVAE-2D 实现 8 倍图像压缩。实验表明,KVAE 在重建和生成指标上达到或超越 Wan-2.2、HunyuanVideo-1.5、FLUX.2 等开源 tokenizer。代码和权重已公开。
正文节选
KVAE: Family of Tokenizers for Multimodal Generative Models Abstract Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-condi