返回全部动态

通过采样引导和扩展LLM的配方

原标题:Recipes for Steering and Scaling LLMs via Sampling

arXiv cs.CL一手来源研究质量 83

AI 摘要

本文提出了一种通过采样来引导和扩展自回归大语言模型(LLM)的灵活且理论基础的框架,涵盖两种算法:基于序贯蒙特卡洛(SMC)和副本交换(RE)的方法,用于从幂化、乘积或倾斜的基础分布中采样。实验表明,这些方法在数学推理任务上无需外部监督或奖励模型即可提升生成质量,且扩展性优于Best-of-N和标准MCMC基线。该框架为LLM的概率推断提供了系统化的采样方案。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

\ul Recipes for Steering and Scaling LLMs via Sampling Abstract Large Language Models (LLMs) are probabilistic models, typically defined by an autoregressive factorization. While recent work has begun to study richer target distributions beyond the base model, the sampling strategies remain highly inefficient. In this paper, we present a flexible and theoretically grounded framework for steering and scaling autoregressive LLMs with sampling. Within this framework, we describe two algorithms—one


发布时间:2026-08-28 12:00
抓取时间:2026-08-28 18:10
来源机构:arXiv
阅读原文arxiv.org