返回全部动态

InSight-doc:面向长文档理解的智能视觉感知框架

原标题:InSight-doc: Agentic Visual Perception for Long-Document Understanding

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

InSight-doc 是一种用于长文档理解的智能视觉感知框架,将视觉分辨率视为可自适应调整的推理时资源。它从低分辨率开始,无需外部检索器即可选择性放大高分辨率区域以获取更精细的证据。通过包含17.9K SFT示例和19.2K硬RL示例的训练,InSight-doc-8B在文档VQA基准上提升了4.3-16.4个准确率点,在长文档上将幻觉减少超过40%,推理延迟降低41%-68%,同时保持准确率优势。代码、数据集和模型已在GitHub上发布。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

InSight-doc: Agentic Visual Perception for Long-Document Understanding Abstract InSight-doc adaptively allocates visual resolution during reasoning to improve long-document understanding while reducing latency and hallucinations. Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception framework that treats visual resolution as an adaptive reasoning-time


发布时间:—
抓取时间:2026-08-12 18:13
来源机构:Hugging Face
阅读原文huggingface.co