部分可观测目标下的CBCT临床报告生成推理
原标题:Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation
AI 摘要
该研究针对锥形束CT(CBCT)颌面部报告生成任务,提出一种复合评分目标:80%权重依赖大语言模型对事实蕴含的判断(RadFact),20%权重为词汇重叠指标(BLEU-4和METEOR),但开发阶段仅可见词汇部分。作者用纯Python复现评分器并构建离线蕴含代理模型,在622例公开数据集上,按可见词汇排名选择的报告得分0.2909,而按复合目标选择的报告得分0.4122,说明追求n-gram重叠会降低蕴含精度。系统最终在50例未见中心测试集上达到METEOR 0.3542,并开源了数据集与代码。
正文节选
Clinical Reasoning Under a Partially Observed Objective in Cone Beam CT Report Generation Abstract Maxillofacial report generation from cone beam computed tomography is scored here by a composite objective placing 80% of its weight on a large language model judgement of factual entailment and 20% on lexical overlap, of which only the lexical fifth is visible during development. The grader’s BLEU-4 and METEOR routines are reproduced in pure Python and match the reference to machine precision, and