返回全部动态

研究者从专有LLM API中窃取推理痕迹

原标题:Stealing Reasoning Traces from Proprietary LLM APIs

Simon Willison's Weblog研究质量 79

AI 摘要

研究者发现Anthropic、OpenAI和Google的专有LLM API返回的加密推理链可被重放至同系列较弱模型,通过越狱恢复明文推理内容。该漏洞已被修复,但论文附录展示了提取的推理细节,揭示了专有模型的思维链。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

11th August 2026 - Link Blog Stealing Reasoning Traces from Proprietary LLM APIs (via) A vanity domain name (stolen-thoughts.com) for a neat paper: Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced by a frontier model, replay it into a weaker sibling, jailbreak the weaker model, and recover the stronger model’s hidden reasoning in plaintext You can see an example of these encrypted b


发布时间:2026-08-12 06:40
抓取时间:2026-08-12 06:56
来源机构:Simon Willison
阅读原文simonwillison.net