Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Other literature type . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
ZENODO
Conference object . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Text Restoration of Historical Documents

Authors: Shibingfeng, Zhang;

Text Restoration of Historical Documents

Abstract

This PhD project investigates the application of pre-trained language models (PLMs) to the automated restoration of Latin diplomatic texts, with a focus on medieval notary documents. The project addresses a significant challenge in historical document studies: the reconstruction of damaged or missing text in low-resource Latin corpora. To this end, the project systematically evaluates a range of PLMs that vary in architecture, training language, and scale, to identify the most effective approach for this specialised restoration task. The project is structured around the following research questions: Does adding Ancient Greek and English during pre-training improve performance in Latin text restoration, or is monolingual pre-training exclusively on Latin more effective? How does the performance of smaller, domain-specific models fine-tuned on Latin compare to large-scale commercial large language models using few-shot prompting in the context of Latin text restoration? The experimental design distinguishes between two key settings based on whether the length of the missing text is known or unknown, which leads to the evaluation of both encoder-based models and encoder-decoder or decoder-only models. Controlled comparisons between model pairs which share identical architectures but differing in training data allow for a rigorous assessment of the effect of multilingual pre-training on downstream Latin text restoration tasks.

Keywords

digital diplomatics, language model, text restoration

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average