Archivio istituzionale della ricerca dell'Università degli Studi di Palermo

Automatic Text Complexity Evaluation (ATE) is a research field that aims at creating new methodologies to make autonomous the process of the text complexity evaluation, that is the study of the text-linguistic features (e.g., lexical, syntactical, morphological) to measure the grade of comprehensibility of a text. ATE can affect positively several different contexts such as Finance, Health, and Education. Moreover, it can support the research on Automatic Text Simplification (ATS), a research area that deals with the study of new methods for transforming a text by changing its lexicon and structure to meet specific reader needs. In this paper, we illustrate an ATE approach named DeepEva, a Deep Learning based system capable of classifying both Italian and English sentences on the basis of their complexity. The system exploits the Treetagger annotation tool, two Long Short Term Memory (LSTM) neural unit layers, and a fully connected one. The last layer outputs the probability of a sentence belonging to the easy or complex class. The experimental results show the effectiveness of the approach for both languages, compared with several baselines such as Support Vector Machine, Gradient Boosting, and Random Forest.

Lo Bosco, G., Pilato, G., Schicchi, D. (2021). DeepEva: A deep neural network architecture for assessing sentence complexity in Italian and English languages. ARRAY, 12, 1-10 [10.1016/j.array.2021.100097].

DeepEva: A deep neural network architecture for assessing sentence complexity in Italian and English languages

Lo Bosco, Giosué;Pilato, Giovanni;Schicchi, Daniele

2021-01-01

Abstract

Automatic Text Complexity Evaluation (ATE) is a research field that aims at creating new methodologies to make autonomous the process of the text complexity evaluation, that is the study of the text-linguistic features (e.g., lexical, syntactical, morphological) to measure the grade of comprehensibility of a text. ATE can affect positively several different contexts such as Finance, Health, and Education. Moreover, it can support the research on Automatic Text Simplification (ATS), a research area that deals with the study of new methods for transforming a text by changing its lexicon and structure to meet specific reader needs. In this paper, we illustrate an ATE approach named DeepEva, a Deep Learning based system capable of classifying both Italian and English sentences on the basis of their complexity. The system exploits the Treetagger annotation tool, two Long Short Term Memory (LSTM) neural unit layers, and a fully connected one. The last layer outputs the probability of a sentence belonging to the easy or complex class. The experimental results show the effectiveness of the approach for both languages, compared with several baselines such as Support Vector Machine, Gradient Boosting, and Random Forest.

Scheda breve

Scheda completa

Scheda completa (DC)

	Data
	
				2021
			
	Titolo del periodico 
DATO PREVISTO SU LOGINMIUR
	
				ARRAY
			
	DOI del contributo 
DATO PREVISTO SU LOGINMIUR
	
				https://dx.doi.org/10.1016/j.array.2021.100097
			
	URL dell'editore (Open access ove possibile)
	
				https://www.sciencedirect.com/science/article/pii/S2590005621000424?via=ihub
			
	Citazione
	
				Lo Bosco, G., Pilato, G., Schicchi, D. (2021). DeepEva: A deep neural network architecture for assessing sentence complexity in Italian and English languages. ARRAY, 12, 1-10 [10.1016/j.array.2021.100097].
			
	Appare nelle tipologie:
	
				1.01 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
1-s2.0-S2590005621000424-definitive.pdf accesso aperto Tipologia: Versione Editoriale Dimensione 764.74 kB Formato Adobe PDF Visualizza/Apri	764.74 kB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/524419

Citazioni

ND

12

4

social impact