返回全部动态

Cohere 发布 Parse 5 多模态模型,提升复杂文档信息提取效率

原标题:Cohere’s Parse 5 Promises Efficient Multi-Modal Information Extraction From Complex Documents

InfoQ AI ML and Data Engineering模型发布质量 75

AI 摘要

Cohere 于 2026 年 8 月 27 日发布了 Parse 5,一个 23 亿参数的多模态视觉语言模型,旨在从复杂企业文档(如 PDF)中提取结构化数据并转换为 Markdown,同时提供边界框坐标用于视觉定位。该模型基于开源的 North-Micro-Vision-Instruct 架构,在 ParseBench 基准上平均得分 79.2,优于 Mistral OCR 和 Gemini 3 Flash,但低于 LlamaParse Agentic Plus。API 可通过 Cohere 平台、Azure AI Foundry 和 AWS SageMaker 使用,支持 RAG 和智能体系统,社区反响积极。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Cohere has officially released Parse 5 (parse-v5.0), a proprietary multimodal foundation model specifically engineered to address the persistent developer challenge of extracting structured data from complex enterprise documents. Launched on August 27th 2026, the 2.3-billion-parameter Vision Language Model (VLM) converts visually rich PDFs, including financial reports and scientific papers, into clean Markdown while providing precise bounding box coordinates for visual grounding. Architecturally


发布时间:2026-09-03 14:06
抓取时间:2026-09-03 14:40
来源机构:InfoQ
阅读原文infoq.com