返回全部动态
PACE:基于智能体自动化的发布者自适应内容提取
原标题:PACE: Publisher-Adaptive Content Extraction via Agentic Automation
AI 摘要
PACE 是一个用于网页内容提取的智能体框架,通过 LLM 分析页面结构并学习发布者特定的提取配置,在推理时使用固定确定性提取器模板,无需额外 LLM 调用。实验表明,PACE 在文章正文、元数据和多模态提取上优于可扩展的非人工基线,接近人工构建的发布者特定解析器质量。该框架旨在自动化发布者特定提取,为 LLM 数据管道提供可扩展的页面表示。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
PACE: Publisher-Adaptive Content Extraction via Agentic Automation ††thanks: Citation: Authors. Title. Pages…. DOI:000000/11111. Abstract Web content extraction is essential for reliable LLM data pipelines, yet existing methods often struggle to jointly satisfy accuracy, scalability, and adaptability. General-purpose extractors can be applied broadly, but they are often brittle on publisher-specific layouts and richer extraction targets such as metadata, images, and tables. Direct LLM-based extr
发布时间:2026-08-31 12:00
抓取时间:2026-08-31 12:04
来源机构:arXiv