Reproducibility of radiomics quality score: an intra- and inter-rater reliability study

Akinci D'Antonoli, T.; Cavallo, A.U.; Vernuccio, F.; Stanzione, A.; Klontzas, M.E.; Cannella, R.; Ugga, L.; Baran, A.; Fanni, S.C.; Petrash, E.; Ambrosini, I.; Cappellini, L.A.; Van Ooijen, P.; Kotter, E.; Pinto Dos Santos, D.; Cuocolo, R.

doi:10.1007/s00330-023-10217-x

Objectives: To investigate the intra- and inter-rater reliability of the total radiomics quality score (RQS) and the reproducibility of individual RQS items’ score in a large multireader study. Methods: Nine raters with different backgrounds were randomly assigned to three groups based on their proficiency with RQS utilization: Groups 1 and 2 represented the inter-rater reliability groups with or without prior training in RQS, respectively; group 3 represented the intra-rater reliability group. Thirty-three original research papers on radiomics were evaluated by raters of groups 1 and 2. Of the 33 papers, 17 were evaluated twice with an interval of 1 month by raters of group 3. Intraclass coefficient (ICC) for continuous variables, and Fleiss’ and Cohen’s kappa (k) statistics for categorical variables were used. Results: The inter-rater reliability was poor to moderate for total RQS (ICC 0.30–055, p < 0.001) and very low to good for item’s reproducibility (k − 0.12 to 0.75) within groups 1 and 2 for both inexperienced and experienced raters. The intra-rater reliability for total RQS was moderate for the less experienced rater (ICC 0.522, p = 0.009), whereas experienced raters showed excellent intra-rater reliability (ICC 0.91–0.99, p < 0.001) between the first and second read. Intra-rater reliability on RQS items’ score reproducibility was higher and most of the items had moderate to good intra-rater reliability (k − 0.40 to 1). Conclusions: Reproducibility of the total RQS and the score of individual RQS items is low. There is a need for a robust and reproducible assessment method to assess the quality of radiomics research. Clinical relevance statement: There is a need for reproducible scoring systems to improve quality of radiomics research and consecutively close the translational gap between research and clinical implementation. Key Points: • Radiomics quality score has been widely used for the evaluation of radiomics studies. • Although the intra-rater reliability was moderate to excellent, intra- and inter-rater reliability of total score and point-by-point scores were low with radiomics quality score. • A robust, easy-to-use scoring system is needed for the evaluation of radiomics research.

Akinci D'Antonoli T., Cavallo A.U., Vernuccio F., Stanzione A., Klontzas M.E., Cannella R., et al. (2023). Reproducibility of radiomics quality score: an intra- and inter-rater reliability study. EUROPEAN RADIOLOGY [10.1007/s00330-023-10217-x].

Reproducibility of radiomics quality score: an intra- and inter-rater reliability study

Akinci D'Antonoli T.;Cavallo A. U.;Vernuccio F.;Stanzione A.;Klontzas M. E.;Cannella R.;Ugga L.;Baran A.;Fanni S. C.;Petrash E.;Ambrosini I.;Cappellini L. A.;van Ooijen P.;Kotter E.;Pinto dos Santos D.;Cuocolo R.

2023-09-01

Abstract

Objectives: To investigate the intra- and inter-rater reliability of the total radiomics quality score (RQS) and the reproducibility of individual RQS items’ score in a large multireader study. Methods: Nine raters with different backgrounds were randomly assigned to three groups based on their proficiency with RQS utilization: Groups 1 and 2 represented the inter-rater reliability groups with or without prior training in RQS, respectively; group 3 represented the intra-rater reliability group. Thirty-three original research papers on radiomics were evaluated by raters of groups 1 and 2. Of the 33 papers, 17 were evaluated twice with an interval of 1 month by raters of group 3. Intraclass coefficient (ICC) for continuous variables, and Fleiss’ and Cohen’s kappa (k) statistics for categorical variables were used. Results: The inter-rater reliability was poor to moderate for total RQS (ICC 0.30–055, p < 0.001) and very low to good for item’s reproducibility (k − 0.12 to 0.75) within groups 1 and 2 for both inexperienced and experienced raters. The intra-rater reliability for total RQS was moderate for the less experienced rater (ICC 0.522, p = 0.009), whereas experienced raters showed excellent intra-rater reliability (ICC 0.91–0.99, p < 0.001) between the first and second read. Intra-rater reliability on RQS items’ score reproducibility was higher and most of the items had moderate to good intra-rater reliability (k − 0.40 to 1). Conclusions: Reproducibility of the total RQS and the score of individual RQS items is low. There is a need for a robust and reproducible assessment method to assess the quality of radiomics research. Clinical relevance statement: There is a need for reproducible scoring systems to improve quality of radiomics research and consecutively close the translational gap between research and clinical implementation. Key Points: • Radiomics quality score has been widely used for the evaluation of radiomics studies. • Although the intra-rater reliability was moderate to excellent, intra- and inter-rater reliability of total score and point-by-point scores were low with radiomics quality score. • A robust, easy-to-use scoring system is needed for the evaluation of radiomics research.

Scheda breve

Scheda completa

Scheda completa (DC)

	Data
	
				set-2023
			
	Titolo del periodico 
DATO PREVISTO SU LOGINMIUR
	
				EUROPEAN RADIOLOGY
			
	DOI del contributo 
DATO PREVISTO SU LOGINMIUR
	
				https://dx.doi.org/10.1007/s00330-023-10217-x
			
	URL dell'editore (Open access ove possibile)
	
				https://link.springer.com/article/10.1007/s00330-023-10217-x
			
	URL alternativo rispetto a quello dell'editore 
DATO PREVISTO SU LOGINMIUR
	
				https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10957586/
			
	Citazione
	
				Akinci D'Antonoli T.,  Cavallo A.U.,  Vernuccio F.,  Stanzione A.,  Klontzas M.E.,  Cannella R., et al. (2023). Reproducibility of radiomics quality score: an intra- and inter-rater reliability study. EUROPEAN RADIOLOGY [10.1007/s00330-023-10217-x].
			
	Appare nelle tipologie:
	
				1.01 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
330_2023_Article_10217.pdf accesso aperto Tipologia: Versione Editoriale Dimensione 1.91 MB Formato Adobe PDF Visualizza/Apri	1.91 MB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10447/639660

Citazioni

29

47

12

Archivio istituzionale della ricerca dell'Università degli Studi di Palermo

Reproducibility of radiomics quality score: an intra- and inter-rater reliability study

Akinci D'Antonoli T.;Cavallo A. U.;Vernuccio F.;Stanzione A.;Klontzas M. E.;Cannella R.;Ugga L.;Baran A.;Fanni S. C.;Petrash E.;Ambrosini I.;Cappellini L. A.;van Ooijen P.;Kotter E.;Pinto dos Santos D.;Cuocolo R.

2023-09-01

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

Citazioni

social impact

Archivio istituzionale della ricerca dell'Università degli Studi di Palermo

Reproducibility of radiomics quality score: an intra- and inter-rater reliability study

Akinci D'Antonoli T.;Cavallo A. U.;Vernuccio F.;Stanzione A.;Klontzas M. E.;Cannella R.;Ugga L.;Baran A.;Fanni S. C.;Petrash E.;Ambrosini I.;Cappellini L. A.;van Ooijen P.;Kotter E.;Pinto dos Santos D.;Cuocolo R.

2023-09-01

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Informazioni

Citazioni

social impact

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)