返回全部动态

FineBooks:开源OCR模型能否解锁历史知识?

原标题:FineBooks: are open OCR models good enough to unlock historical knowledge?

Hugging Face Blog一手来源研究质量 82

AI 摘要

Hugging Face 与 EleutherAI 合作推出 FineBooks 项目,旨在评估开源 OCR 模型在处理历史文献上的表现,并重新处理公共领域书籍以提升 AI 训练数据质量。项目发布了包含 14 个模型的 BHL OCR 排行榜、基准数据集和评估工具,基于 2165 页专家转录的历史书籍页面进行测试。初步结果显示,现代 VLM 模型在历史文档上的准确率仍有提升空间,但成本较低,适合大规模应用。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

As technology has progressed, OCR models have improved dramatically, especially in the last few years with the advent of Visual Language Models (VLMs). Many of this new generation of VLM-based OCR models are published with open weights and under an open license and can therefore be used for free by anyone on any hardware or infrastructure. In most cases, however, these new OCR models have been primarily trained on and optimized for modern documents. Historical books are often considerably more c


发布时间:
抓取时间:2026-08-10 22:26
来源机构:Hugging Face
阅读原文huggingface.co