Cohere 发布 Parse 5 多模态模型,提升复杂文档信息提取效率
原标题:Cohere’s Parse 5 Promises Efficient Multi-Modal Information Extraction From Complex Documents
AI 摘要
Cohere 于 2026 年 8 月 27 日发布了 Parse 5,一个 23 亿参数的多模态视觉语言模型,旨在从复杂企业文档(如 PDF)中提取结构化数据并转换为 Markdown,同时提供边界框坐标用于视觉定位。该模型基于开源的 North-Micro-Vision-Instruct 架构,在 ParseBench 基准上平均得分 79.2,优于 Mistral OCR 和 Gemini 3 Flash,但低于 LlamaParse Agentic Plus。API 可通过 Cohere 平台、Azure AI Foundry 和 AWS SageMaker 使用,支持 RAG 和智能体系统,社区反响积极。
正文节选
Cohere has officially released Parse 5 (parse-v5.0), a proprietary multimodal foundation model specifically engineered to address the persistent developer challenge of extracting structured data from complex enterprise documents. Launched on August 27th 2026, the 2.3-billion-parameter Vision Language Model (VLM) converts visually rich PDFs, including financial reports and scientific papers, into clean Markdown while providing precise bounding box coordinates for visual grounding. Architecturally