返回全部动态

弱监督框架结合LLM标签精炼抽取流离失所文档中的数据集提及

原标题:Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

arXiv cs.CL一手来源研究质量 78

AI 摘要

研究团队提出一种弱监督框架,用于在强迫流离失所及脆弱、冲突与暴力(FCV)领域文档中抽取数据集提及。该方法先用通用研究文献上训练的轻量模型生成候选提及,再由前沿大语言模型在上下文中审核、修正边界,并补充合成与对比样本以微调轻量模型。在包含1706个文本片段的独立基准上,模型提及级精确率74.1%、召回率70.5%,含数据集引用片段精确率达89.5%,片段级准确率88.2%。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement Abstract Development and humanitarian organizations produce and support surveys, administrative registries, and other data resources to inform research, policy, and operations, yet systematically identifying where these datasets are referenced remains difficult. Such references are dispersed across research papers, project documents, humanitarian reports, and other


发布时间:2026-09-14 12:00
抓取时间:2026-09-14 12:05
来源机构:arXiv
阅读原文arxiv.org