The restoration of lacunae in ancient Greek inscriptions remains a central challenge in digital epigraphy and historical document analysis. The difficulty is heightened in texts written in scriptio continua, where the absence of explicit word boundaries undermines token-based NLP assumptions. At the same time, epigraphic corpora are inherently low-resource: there are no large-scale aligned datasets of damaged and complete inscriptions suitable for training supervised restoration models. These constraints call for approaches that move beyond data-hungry architectures. We propose a low-resource framework that reconceptualizes lacunae filling as a constrained visual–linguistic inference task. First, a YOLO-based glyph detection model identifies and classifies characters directly from inscription images. From these detections, we estimate inter-letter spacing and measure the geometric extent of missing regions, deriving upper bounds on the number of characters that can plausibly fill each gap. This spatial constraint significantly reduces the search space of candidate restorations. Second, we exploit character-level n-gram statistics extracted from independent corpora of ancient Greek texts. Working at the character level avoids reliance on explicit tokenization and ensures robustness to scriptio continua. Candidate restorations are generated by combining probabilistic n-gram priors with the left and right character context surrounding each gap, producing ranked hypotheses that are both linguistically plausible and geometrically consistent. We evaluate the method on the Gortyn Law Code dataset using Top-k accuracy to measure restoration quality. Experimental results show that the proposed framework achieves strong performance while drastically reducing the number of candidate sequences that epigraphists must inspect, providing a practical and scalable tool for the restoration of fragmentary inscriptions.

Bridging the Gaps: Learning to Estimate Missing Text in Fragmentary Greek Inscriptions

Zunino, Maddalena
;
Mignosa, Valentina
;
2026

Abstract

The restoration of lacunae in ancient Greek inscriptions remains a central challenge in digital epigraphy and historical document analysis. The difficulty is heightened in texts written in scriptio continua, where the absence of explicit word boundaries undermines token-based NLP assumptions. At the same time, epigraphic corpora are inherently low-resource: there are no large-scale aligned datasets of damaged and complete inscriptions suitable for training supervised restoration models. These constraints call for approaches that move beyond data-hungry architectures. We propose a low-resource framework that reconceptualizes lacunae filling as a constrained visual–linguistic inference task. First, a YOLO-based glyph detection model identifies and classifies characters directly from inscription images. From these detections, we estimate inter-letter spacing and measure the geometric extent of missing regions, deriving upper bounds on the number of characters that can plausibly fill each gap. This spatial constraint significantly reduces the search space of candidate restorations. Second, we exploit character-level n-gram statistics extracted from independent corpora of ancient Greek texts. Working at the character level avoids reliance on explicit tokenization and ensures robustness to scriptio continua. Candidate restorations are generated by combining probabilistic n-gram priors with the left and right character context surrounding each gap, producing ranked hypotheses that are both linguistically plausible and geometrically consistent. We evaluate the method on the Gortyn Law Code dataset using Top-k accuracy to measure restoration quality. Experimental results show that the proposed framework achieves strong performance while drastically reducing the number of candidate sequences that epigraphists must inspect, providing a practical and scalable tool for the restoration of fragmentary inscriptions.
2026
Document Analysis and Recognition – ICDAR 2026
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/5124447
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact