Unveiling Ancient Secrets: AI's Revolutionary Impact on Medieval Manuscripts
In a groundbreaking development, AI technology has cracked open the mysteries of ancient manuscripts, offering a glimpse into the past that was previously unimaginable. The project, named CoMMA (Corpus of Multilingual Medieval Archives), is a testament to the power of modern innovation and its ability to revolutionize historical research.
Unlocking the Past
The challenge of deciphering medieval manuscripts is immense. Consider the sheer volume of work required: transcribing a single page of 12th-century Latin can take years, and a full manuscript might consume a researcher's entire career. Yet, a collaborative effort between Inria and philologists has automated this process, processing an astonishing 32,763 manuscripts in just four months.
The result? A freely accessible online archive, CoMMA, where transcriptions are displayed alongside digitized pages. This achievement is nothing short of remarkable, especially when you consider the complexities involved in recognizing handwritten text, let alone ancient scripts.
The Complexity of Ancient Scripts
One of the key challenges is the variability of ancient scripts. A single letter or word can take on drastically different forms depending on the century and document type. As Thibault Clérice, a researcher at Inria, puts it, "Lecture notes or administrative documents hastily scribbled will always be much harder to crack than a beautiful, regular manuscript copied out for a noble or a king."
And that's not even considering the issue of abbreviations, which are rampant in medieval texts. Up to 35-40% of words in 14th-century Latin manuscripts are abbreviated, and for Old French, the ratio is between 7 and 12%. In some technical works, only half the letters might be present on the page, with the rest left to be 'understood' by the reader!
Why Popular AI Language Models Fall Short
You might wonder why popular AI language models like GPT or Mistral couldn't tackle this task. The answer lies in the nature of these models, which generate text by predicting 'tokens' or character strings. They thrive on patterns and regularity, but medieval French, for example, had no established spelling rules. The model would be left guessing, with a high likelihood of hallucinating a word that sounds plausible but is completely made up.
A Graphic Interpretation
So, the team took a different approach. Instead of relying on meaning, they focused on shape. It's a graphic interpretation, where every sign is analyzed independently, with even accents counted as separate characters. For instance, the letter 'à' is treated as two signs to recognize.
This led to the CATMuS project, initiated by Ariane Pinche at CNRS. The goal was to build a consistent learning corpus before training algorithms. Over several years, the team transcribed 200,000 lines from 300 manuscripts in 11 languages, covering the 9th to the 16th centuries. The key rule? Don't correct anything. Abbreviations, spelling, and even scribal mistakes were left as they were, providing a raw and authentic representation of the manuscripts.
Building CoMMA
The algorithm was then trained using open-source tools like Kraken and eScriptorium, which are not based on language models, thus avoiding the issue of hallucinations. The team applied this model to a vast collection of digitized manuscripts from various sources, including the French National Library, ARCA, e-codices, the Bodleian Library, and the Bavarian State Library.
The CoMMA platform delivers transcriptions as-is, without human revision. The research team checked a sample of manuscripts and found an average error rate of 9.7%. Each document includes metadata with the exact percentage of lines correctly recognized, with scores often passing 80% or even 90%. Predictably, more cursive scripts, especially in later manuscripts, posed more challenges due to their scarcity in the training data.
A Treasure Trove for Scholars
The CoMMA collection is a treasure trove, comprising over three billion words, mostly in Latin and Old French. For Old French alone, the volume of available texts has increased fortyfold. This opens up a whole new world of research possibilities in historical linguistics, philology, and textual history, which were previously limited by a lack of material.
In my opinion, this project is a perfect example of how technology can enhance our understanding of the past. It's a fascinating development that showcases the potential for AI to revolutionize historical research and unlock centuries-old secrets.