返回全部动态

Together AI 发布 YAQA 量化算法,显著降低模型输出偏差

原标题:Model-Preserving Adaptive Rounding with YAQA

Together AI Blog一手来源研究质量 85

AI 摘要

Together AI 发布了新的后训练量化方法 YAQA,通过直接最小化 KL 散度来保留原始模型输出,与现有量化器(如 QTIP)兼容,可将 KL 散度降低超过 30%,并在下游任务上取得最优性能。该方法利用 Fisher 信息矩阵和 Kronecker 分解来近似 Hessian,从而在保持模型质量的同时实现高效量化。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We’re excited to announce YAQA (Yet Another Quantization Algorithm), a new weight-only LLM post-training quantization method that quantizes models to directly preserve the original model’s outputs. YAQA is quantizer-agnostic, meaning that it can be used with both hardware-accelerated datatypes and special memory-bound quantizers like QTIP. Across a wide range of models and quantizers, YAQA consistently reduces the KL divergence to the original model by >30% over existing rounding algorithms, res


发布时间:—
抓取时间:2026-08-03 01:13
来源机构:Together AI
阅读原文together.ai