LLMs能否真正理解题目难度?对自动生成题目的启示
原标题:Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
AI 摘要
该研究探讨了大型语言模型(LLMs)在预测大规模阅读与写作测试中题目难度水平的表现。研究发现,零样本GPT-4.1(温度设为0)的预测准确率最高,二次加权kappa(QWK)为0.578,但仍低于ConvBERT(QWK=0.625)。所有LLMs在标记难题时表现不佳,GPT-5.4倾向于低估难度,且嵌入分析表明仅凭语义信息不足以预测难度。研究建议在利用LLMs生成特定难度题目时应谨慎。
正文节选
Computer Science > Computation and Language Title:Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs View PDF Abstract:The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study investigated variou