返回全部动态
扩散语言模型用于无损文本压缩:突破吞吐瓶颈
原标题:Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
AI 摘要
本文首次将扩散语言模型(DLMs)引入无损文本压缩领域,以替代自回归LLM的逐符号生成瓶颈。作者设计了符号提交调度和初始上下文策略,解决DLM在压缩中的算法挑战。在enwik8基准上的实验表明,该框架在压缩率和吞吐量权衡上优于现有神经压缩方法,并具有进一步改进的潜力。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression Abstract We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data — including plain text, source code, and structured formats such as XML — and by recent advances in neural language model-based compression. In particular, recent LLM-based approaches, whether built on symbol-ranking pipelines or paired with a statistical compressor, have demonstrated
发布时间:2026-08-13 12:00
抓取时间:2026-08-13 12:05
来源机构:arXiv