返回全部动态

身份还是提示噪声?LLM 代码生成的校准不变性审计

原标题:Identity or Prompt Noise? A Calibrated Invariance Audit of LLM Code Generation

arXiv cs.SE一手来源研究质量 79

AI 摘要

该研究对七个模型在 HumanEval+ 和 MBPP+ 上的 3073 万次 Python 代码生成进行了校准不变性审计,考察模型分配性别、国家和职业人设是否影响代码输出。经任务内随机化和 FDR 校正后,职业是最一致的结构性轴:CodeBLEU 离散度在 10/14 个模型-基准单元中超过可交换性零假设,但中位超额仅 0.141 分,且通过率离散度均不显著。作者认为证据支持职业条件下代码形式存在微小可复现变化,而非对特定身份的稳定劣势或已证实的下游危害。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Identity or Prompt Noise? A Calibrated Invariance Audit of LLM Code Generation Abstract. Identity cues are irrelevant to a fixed programming specification, but raw counterfactual differences can arise from unequal samples and prompt wording. We audit 30.73 million executed Python generations from seven checkpoints on HumanEval+ and MBPP+, supplemented by an exploratory 550B slice, under model-assigned gender, country, and occupation personas. Within-task randomization and false-discovery-rate co


发布时间:2026-09-22 12:00
抓取时间:2026-09-22 12:56
来源机构:arXiv
阅读原文arxiv.org