返回全部动态

大语言模型的概率结构

原标题:The Probabilistic Structure of Large Language Models

arXiv cs.LG一手来源研究质量 81

AI 摘要

这篇 arXiv 论文从概率论视角统一阐述大语言模型,将 LLM 描述为 token 序列空间上的概率测度,通过自回归条件分布指定,训练被形式化为最大似然估计并用随机梯度方法求解,文本生成则视为该随机过程的序贯模拟。论文还分析了 KL 散度不对称性在文本生成中的作用,并将其与幻觉现象以及统计合理性与真实性的区分联系起来。作为同一视角的补充,论文讨论了基于 score 函数的扩散模型,将生成视为逆向时间随机过程模拟。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The Probabilistic Structure of Large Language Models Abstract This paper presents a probabilistic perspective on large language models (LLMs), developed with the aim of bringing together, in a single self-contained account, tools that are usually treated separately across the literature. LLMs are described through probability measures on the set of sequences of tokens, specified via their autoregressive conditional distributions. Training is formulated as a maximum-likelihood estimation problem,


发布时间:2026-09-23 12:00
抓取时间:2026-09-23 12:14
来源机构:arXiv
阅读原文arxiv.org