弱监督框架结合LLM标签精炼抽取流离失所文档中的数据集提及
原标题:Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement
AI 摘要
研究团队提出一种弱监督框架,用于在强迫流离失所及脆弱、冲突与暴力(FCV)领域文档中抽取数据集提及。该方法先用通用研究文献上训练的轻量模型生成候选提及,再由前沿大语言模型在上下文中审核、修正边界,并补充合成与对比样本以微调轻量模型。在包含1706个文本片段的独立基准上,模型提及级精确率74.1%、召回率70.5%,含数据集引用片段精确率达89.5%,片段级准确率88.2%。
正文节选
Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement Abstract Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other