返回全部动态

Together AI 平台支持 DPO:技术深度解析

原标题:Direct Preference Optimization: A Technical Deep Dive

Together AI Blog一手来源产品发布质量 72

AI 摘要

Together AI 宣布其微调平台现已支持直接偏好优化(DPO),这是一种无需强化学习即可将语言模型与人类偏好对齐的技术。文章详细介绍了 DPO 的原理、与 RLHF 的对比,并推荐先进行监督微调(SFT)再进行 DPO 的两阶段训练方法,以提升模型的有用性、真实性和无害性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We're excited to announce that the Together Fine-Tuning Platform now supports Direct Preference Optimization (DPO)! This technique allows developers to align language models with human preferences creating more helpful, accurate, and tailored AI assistants. In this deep-dive blogpost, we provide details of what DPO is, how it works, when to use it and code examples. If you'd like to jump straight into code have a look at our code notebook. Tuning LLMs on Preference Data Modern language model dev


发布时间:—
抓取时间:2026-08-03 01:13
来源机构:Together AI
阅读原文together.ai