SNAIL:从文献中自动识别生物信息学软件名称的混合框架
原标题:Automatic bioinformatic software named entity recognition from literature
AI 摘要
堪萨斯大学等机构的研究者提出了SNAIL,一个混合命名实体识别框架,用于从生物医学文献中自动识别生物信息学软件和数据库名称。SNAIL结合词汇与语义建模,利用SciBERT等Transformer模型和token掩码策略,并通过引用提示和LLM辅助蒸馏构建训练语料。在多个基准和真实文献上,SNAIL显著优于bioNerDS2及ChatGPT、Gemini等通用大模型,并能揭示期刊层面的工具使用偏好。该研究为生物信息学资源识别提供了准确可扩展的解决方案,支持大规模文献元分析。
正文节选
[Page 1] Title • Automatic bioinformatic software named entity recognition from literature Authors Hao Xuan1, Rithvij Pasupuleti1, Ben Liu1, Haishuo Sun1, Jun Zhang2,3, Zijun Yao1,* and Cuncong Zhong1,4,5,* Affiliations 1Department of Electrical Engineering and Computer Science, University of Kansas, Lawrence, KS 66045, USA 2OSF Healthcare Cancer Institute, Peoria, IL 61603, USA 3University of Illinois College of Medicine, Peoria, IL 61605 USA 4B