Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale
Enables developers working with large-scale digitized text corpora to customize preprocessing while preserving linguistic and structural metadata.
AI Summary
Researchers release IB-HL-ET, an open-source multilingual pipeline for denoising, deduplicating, and annotating OCR text at scale (217B tokens, 983K volumes) to preserve metadata for machine parsing and human study.
Excerpt
Released in 2025, Institutional Books: Harvard Library (IB-HL) is a collection of 983,004 volumes (242B o200k_base tokens), originally digitized through Harvard Library's participation in the Google Books Library project. As researchers and developers have begun to use IB-HL, a tension has emerged between standard large-scale preprocessing practices and the goals of careful information stewardship. Many existing pipelines optimize for web text: as a result, they tend to aggressively filter, dedu
