AI聊天机器人检索医学研究的能力评估:模型、用户角色和样本量的影响
原标题:Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
AI 摘要
一项研究评估了Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5在医学问题中检索临床研究的能力,发现平均仅能检索到Cochrane综述中39.2%的纳入研究,且存在对大规模试验的偏见。模型和用户角色显著影响检索性能,其中ChatGPT表现最佳,研究者角色检索效果优于临床医生和患者角色。
正文节选
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions Abstract Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of the retrieved studies and the factors driving their selection, particularly for newer models with stronger reasoning capabilities.