大语言模型在结直肠息肉光学诊断中的性能评估
原标题:Performance of large language models in the optical diagnosis of colorectal polyps
AI 摘要
一项发表在 arXiv 上的研究评估了多种多模态大语言模型(MLLMs)在结直肠息肉光学诊断中的表现。研究使用 PRIME 数据集,比较了 Claude Opus 4、Gemini 2.5 Pro、GPT-o3、GPT-4o 和 GPT-5 的分类准确性。结果显示,所有模型在区分肿瘤性与非肿瘤性息肉方面 F1 分数均高于 0.9,但 Gemini 2.5 Pro 在区分浸润性与非浸润性息肉及低级别与高级别腺瘤方面表现最佳。Claude Opus 4 和 GPT-5 在巴黎分类中正确率最高,但整体敏感性和特异性未达到 ESGE 标准,需进一步前瞻性试验和人工介入流程。
正文节选
Computer Science > Computer Vision and Pattern Recognition Title:Performance of large language models in the optical diagnosis of colorectal polyps View PDF Abstract:Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology.