AI and DIY save 1800 rare books by processing 526k scans


Featured image AI and DIY save 1800 rare books by processing 526k scans

From Photocopier to Algorithm: How Three Friends Digitized Lost Urdu Literature

In 2015, a group of Pakistani friends embarked on an ambitious mission driven by pure love for language: to digitize a vast collection of out-of-print Urdu books, many of which were fragile lithographs. This project, which eventually grew into the Ibteda Digital Library”>Ibteda Digital Library, began without a budget or institutional support, relying instead on their personal commitment and sheer determination.

The process was anything but simple. They started with a very analog approach, utilizing just a Nikon D5300 camera, some LED lights, and a sheet of glass taken from a photocopier to press the books flat. Every step required immense manual effort. Turning pages by hand and painstakingly post-processing the resulting images in Photoshop consumed countless hours, demonstrating that achieving archival-quality digital copies of rare texts is a tedious, painstaking endeavor.

The true difficulty lay in the unique nature of the material. The bulk of the work involved Urdu script, specifically the flowing Nastaliq calligraphy. This script, rich with diacritics and unique layouts, presented significant challenges. Unlike standard texts, the varying script formats and the presence of marks and blemishes required meticulous attention to ensure that the photographs captured the essence of the writing without introducing noise.

Each book was a unique puzzle. Because of the layout variations, the density of the script, and the presence of margin notes often found in old lithographs, maintaining consistent perspectives and accurate margins across multiple publications was nearly impossible. The consistency issues meant that no single photographic method could be applied universally; every book was a distinct corner case.

After capturing over 576,000 shutter counts on a D5300 and 326,000 on a D3300 camera, the team was left with over 526,000 dual-page photos to manage. Manually processing this massive volume was simply out of the question. This is where the project pivoted from physical labor to cutting-edge technology. One of the researchers turned to automation, harnessing the power of OpenCV”>OpenCV, a powerful computer vision library.

Initially, applying standard computer vision methods proved unsuccessful, as a rule set successful for one set of books failed spectacularly for the next. The breakthrough came when the team realized they could use their own expertly processed images to establish a source/target correspondence—the “finished pages became labels.” This led to an innovative approach to calculating the homography fit from the Photoshop files, which was then used to train a neural network.

Even with this advanced technique, challenges remained, such as correctly matching crop borders where dense Urdu print repeated strokes and patterns, and accurately identifying visible areas while accounting for deep gutter shadows. The researchers discovered an interesting twist: adding more books to the training set actually made the model’s pattern recognition worse. This demonstrated that the unique, manually determined crop margins for each book were crucial, underscoring the value of customized data in machine learning.

The result is a remarkable achievement: a digitally preserved library of rare Urdu texts. The entire process, from manual photography to advanced AI training, culminated in books now hosted securely in a ZFS pool with BLAKE3″>BLAKE3 manifests, ensuring their integrity. The story of the Ibteda Digital Library proves that passion, combined with a willingness to innovate, can turn analog dedication into a digital legacy.

You may also like: