返回全部动态

普通打字错误如何破坏LLM激活探针

原标题:Latent Undertow: How Ordinary Typos Break Probes

arXiv cs.CL一手来源研究质量 78

AI 摘要

该论文研究普通打字错误(typo)如何破坏基于激活的LLM安全探针。作者发现,扰动token处的激活向量会发生大幅旋转,并在下游token迅速衰减,导致单位置提示注入探针的TPR@FPR1%显著下降,且无法仅靠阈值重校准恢复。为此提出KV-cache fork方法,在用户消息后附加固定后缀让探针读取扰动下游的token,将差距缩小约一个数量级,优于扰动增强训练。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Latent Undertow: How Ordinary Typos Break Probes Abstract LLMs handle ordinary typing variation fluently: a typo or missing punctuation leaves both user intent and the model’s response substantively unchanged. Yet probes that detect malicious prompts by reading the model’s hidden states tell a different story: the same edit rotates the readout vector by – at the perturbed token, decaying below within downstream tokens. Stacking common typos per message cuts a single-position prompt-injection pro


发布时间:2026-09-16 12:00
抓取时间:2026-09-16 12:10
来源机构:arXiv
阅读原文arxiv.org