This chapter explores the challenges of applying computational methods to Old Church Slavonic and Church Slavonic texts, focusing on the impact of non-standard character encoding and Unicode interoperability on semantic analysis. Developed within the ITSERR project, the study examines the creation of tools and workflows for processing texts from the Cyrillomethodiana portal, addressing issues related to Private Use Area (PUA) characters, font dependencies, and digital text standardisation. Through the development of a Unicode mapping system, a text converter, and the DIACU dataset, the research demonstrates how the normalisation of historical Slavic texts can improve their integration into Natural Language Processing (NLP) environments and semantic retrieval systems. The results highlight the importance of standardised digital resources for the preservation, accessibility, and computational analysis of Slavic religious corpora, providing a foundation for future linguistic, philological, and theological investigations.

Napolitano, M., Nawaz, U. (2025). From Characters to Meaning: Challenges and Limitations in the Semantic Analysis of Church Slavonic Texts. In A. Melloni, F. Cadeddu (a cura di), The Digital Turn in Religious Studies Research, Services, Infrastructures (pp. 41-65).

From Characters to Meaning: Challenges and Limitations in the Semantic Analysis of Church Slavonic Texts

Nawaz, Usman
2025-01-01

Abstract

This chapter explores the challenges of applying computational methods to Old Church Slavonic and Church Slavonic texts, focusing on the impact of non-standard character encoding and Unicode interoperability on semantic analysis. Developed within the ITSERR project, the study examines the creation of tools and workflows for processing texts from the Cyrillomethodiana portal, addressing issues related to Private Use Area (PUA) characters, font dependencies, and digital text standardisation. Through the development of a Unicode mapping system, a text converter, and the DIACU dataset, the research demonstrates how the normalisation of historical Slavic texts can improve their integration into Natural Language Processing (NLP) environments and semantic retrieval systems. The results highlight the importance of standardised digital resources for the preservation, accessibility, and computational analysis of Slavic religious corpora, providing a foundation for future linguistic, philological, and theological investigations.
2025
Napolitano, M., Nawaz, U. (2025). From Characters to Meaning: Challenges and Limitations in the Semantic Analysis of Church Slavonic Texts. In A. Melloni, F. Cadeddu (a cura di), The Digital Turn in Religious Studies Research, Services, Infrastructures (pp. 41-65).
File in questo prodotto:
File Dimensione Formato  
The Digital Turn in Religious Studies_Napolitano-Nawaz.pdf

Solo gestori archvio

Tipologia: Versione Editoriale
Dimensione 555.66 kB
Formato Adobe PDF
555.66 kB Adobe PDF   Visualizza/Apri   Richiedi una copia

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/710773
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact