返回全部动态

FineBooks项目:开源OCR模型助力历史书籍文本识别,小模型表现惊艳

原标题:Old OCR text cripples language model training, and FineBooks wants to fix that at scale

THE DECODER研究质量 76

AI 摘要

Hugging Face 与 EleutherAI 合作发起 FineBooks 项目,在超过 2000 页历史书籍上测试了 14 个开源 OCR 模型,发现小模型表现优于大模型,最佳模型字符准确率超 97%,成本低于每千页 2 美元。研究认为这些模型足以用于 AI 训练数据,但尚不适合学术用途。团队计划用顶级模型重新处理约 20 万份 BHL 文档并开放数据集。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Old OCR text cripples language model training, and FineBooks wants to fix that at scale Key Points - The FineBooks project, a collaboration between Hugging Face and EleutherAI, benchmarked 14 open-source OCR models on over 2,000 historical book pages to evaluate how well they can convert scanned texts into clean training data for AI language models. - Smaller models frequently outperformed larger ones, with the top-performing model achieving over 97 percent character accuracy at a cost of less t


发布时间:2026-08-11 02:20
抓取时间:2026-08-11 02:28
来源机构:THE DECODER
阅读原文the-decoder.com