返回全部动态

现成LLM需求质量评估基准:性能、误报与漏检分析

原标题:Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses

arXiv cs.SE一手来源研究质量 83

AI 摘要

弗吉尼亚理工大学和亚利桑那大学的研究人员首次对现成LLM在需求质量评估中的表现进行了基准测试,基于INCOSE标准评估了OpenAI和Anthropic两大家族共十个模型。结果显示,最佳Anthropic模型仅能检测出47%的专家识别问题,同时产生11%的误报,且对必要性和正确性问题的漏检严重。研究还发现模型代际进步非单调,温度采样影响有限,表明现成LLM尚不能作为自主评估器,更适合作为人在回路的决策支持工具。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

[Page 1] Two Truths and A Lie? Benchmarking off-the-shelf LLMs for Requirements Quality Assessment – Performance, False Alarms, and Misses Jannatul Shefa1, Alejandro Salado2, Paul Wach2, and Taylan G. Topcu1 1Grado Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg, Virginia, USA 2Department of Systems and Industrial Engineering, University of Arizona, Tucson, Arizona, USA Correspondence: Taylan G. Topcu, PhD (ttopcu@vt.ed


发布时间:2026-09-04 12:00
抓取时间:2026-09-04 13:45
来源机构:arXiv
阅读原文arxiv.org