AI检测在学术诚信中的失效:误报与规避问题研究
原标题:Why AI Detection Fails for Academic Integrity
AI 摘要
该研究通过对照实验发现,商业AI检测器无法区分AI辅助编辑与完全由LLM生成的文本,导致合规使用AI辅助写作的学生面临更高的处罚风险,而使用人类化工具规避检测的作弊行为却几乎无法被识别。研究分析了2013-2015年与2023-2025年四个领域的642篇摘要,发现未修改的2023-2025年原文被误报率为9-15%,非STEM领域更高,而经Undetectable AI人类化处理后,AI生成文本的检测率降至4%以下。研究结论认为检测器分数不应作为学术不端的独立证据。
正文节选
Why AI Detection Fails for Academic Integrity Abstract. Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at . Light refine (abstract only) edits, a proxy for guideline-compliant AI assistance, are flagged at 64 to 80% (Pangra