代码的 Token 签名:比较不同大语言模型的编码行为
原标题:Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models
AI 摘要
Visa Research 提出 CLIC(Code Learning for Identification and Comparison),通过统计 LLM 生成代码中的 token 频率,将每个代码样本表示为特征向量,并训练可解释决策树来区分两个 LLM 的代码集合。研究定义了两个新指标:鲁棒性(逐步移除最具区分度 token 后两模型是否仍可区分)和集中度(差异由少数主导 token 还是大量 token 驱动)。在 22 个 Kaggle ML 任务上比较 10 个 LLM 的案例研究,为模型选择和提示工程提供可操作见解。
正文节选
0 \vgtccategoryResearch \vgtcpapertypeAnalytics & Decisions \authorfooterJ. Wang, Y. Chen, M. Pan, U. Saini, Y. Cai are with Visa Research. E-mail: {junpenwa, yuzchen, menpan, udasaini, yicai}@visa.com \teaser \tl_set:Ne\reqboxonereqboxone \tl_set:Ne\reqboxtworeqboxtwo \tl_set:Ne\reqboxthreereqboxthree \tl_set:Ne\sharpcascadingsharpcascading \tl_set:Ne\deepdistributeddeepdistributed \tl_set:Ne\fragilediffusefragilediffuse \tl_set:Ne\lexicalbrittlelexicalbrittle \tl_set:Ne\sumllmboxsumllmbox \tl_