FineBooks项目:开源OCR模型助力历史书籍文本识别,小模型表现惊艳
原标题:Old OCR text cripples language model training, and FineBooks wants to fix that at scale
AI 摘要
Hugging Face 与 EleutherAI 合作发起 FineBooks 项目,在超过 2000 页历史书籍上测试了 14 个开源 OCR 模型,发现小模型表现优于大模型,最佳模型字符准确率超 97%,成本低于每千页 2 美元。研究认为这些模型足以用于 AI 训练数据,但尚不适合学术用途。团队计划用顶级模型重新处理约 20 万份 BHL 文档并开放数据集。
正文节选
Old OCR text cripples language model training, and FineBooks wants to fix that at scale Key Points - The FineBooks project, a collaboration between Hugging Face and EleutherAI, benchmarked 14 open-source OCR models on over 2,000 historical book pages to evaluate how well they can convert scanned texts into clean training data for AI language models. - Smaller models frequently outperformed larger ones, with the top-performing model achieving over 97 percent character accuracy at a cost of less t