返回全部动态

ExtractBench:模式引导的企业文档提取基准

原标题:ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Hugging Face Daily Papers一手来源研究质量 81

AI 摘要

ExtractBench 是一个用于评估模式引导的企业文档提取任务的基准,包含 370 份企业文档、4869 页内容,覆盖 8 个业务领域和 67 种文档类型。该基准首次同时评估值准确性、记录完整性、溯源性和成本。研究发现商业 VLM 在短文档上表现良好但在长文档上会截断记录,而编码代理虽准确率高但成本更高,LlamaExtract Agentic Plus 在三个指标上均排名第一且成本较低。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Abstract Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata. We present ExtractBench, a benchmark for schema-guided extraction and, to our knowledge, the first to score value accuracy, record completeness at scale, grounding, and measured c


发布时间:—
抓取时间:2026-08-03 16:12
来源机构:Hugging Face
阅读原文huggingface.co