Interpretable prioritization of splice variants in diagnostic next-generation sequencing

Danis, D.; Jacobsen, J.O.B.; Carmody, L.C.; Gargano, M.A.; Mcmurry, J.A.; Hegde, A.; Haendel, M.A.; Valentini, G.; Smedley, D.; Robinson, P.N.

doi:10.1016/j.ajhg.2021.06.014

A critical challenge in genetic diagnostics is the computational assessment of candidate splice variants, specifically the interpretation of nucleotide changes located outside of the highly conserved dinucleotide sequences at the 5′ and 3′ ends of introns. To address this gap, we developed the Super Quick Information-content Random-forest Learning of Splice variants (SQUIRLS) algorithm. SQUIRLS generates a small set of interpretable features for machine learning by calculating the information-content of wild-type and variant sequences of canonical and cryptic splice sites, assessing changes in candidate splicing regulatory sequences, and incorporating characteristics of the sequence such as exon length, disruptions of the AG exclusion zone, and conservation. We curated a comprehensive collection of disease-associated splice-altering variants at positions outside of the highly conserved AG/GT dinucleotides at the termini of introns. SQUIRLS trains two random-forest classifiers for the donor and for the acceptor and combines their outputs by logistic regression to yield a final score. We show that SQUIRLS transcends previous state-of-the-art accuracy in classifying splice variants as assessed by rank analysis in simulated exomes, and is significantly faster than competing methods. SQUIRLS provides tabular output files for incorporation into diagnostic pipelines for exome and genome analysis, as well as visualizations that contextualize predicted effects of variants on splicing to make it easier to interpret splice variants in diagnostic settings.

Interpretable prioritization of splice variants in diagnostic next-generation sequencing / D. Danis, J.O.B. Jacobsen, L.C. Carmody, M.A. Gargano, J.A. McMurry, A. Hegde, M.A. Haendel, G. Valentini, D. Smedley, P.N. Robinson. - In: AMERICAN JOURNAL OF HUMAN GENETICS. - ISSN 0002-9297. - 108:9(2021 Sep 02), pp. 1564-1577. [10.1016/j.ajhg.2021.06.014]

Interpretable prioritization of splice variants in diagnostic next-generation sequencing

Danis D.;Jacobsen J. O. B.;Carmody L. C.;Gargano M. A.;McMurry J. A.;Hegde A.;Haendel M. A.;G. Valentini^Methodology;Smedley D.;Robinson P. N.

2021

Abstract

A critical challenge in genetic diagnostics is the computational assessment of candidate splice variants, specifically the interpretation of nucleotide changes located outside of the highly conserved dinucleotide sequences at the 5′ and 3′ ends of introns. To address this gap, we developed the Super Quick Information-content Random-forest Learning of Splice variants (SQUIRLS) algorithm. SQUIRLS generates a small set of interpretable features for machine learning by calculating the information-content of wild-type and variant sequences of canonical and cryptic splice sites, assessing changes in candidate splicing regulatory sequences, and incorporating characteristics of the sequence such as exon length, disruptions of the AG exclusion zone, and conservation. We curated a comprehensive collection of disease-associated splice-altering variants at positions outside of the highly conserved AG/GT dinucleotides at the termini of introns. SQUIRLS trains two random-forest classifiers for the donor and for the acceptor and combines their outputs by logistic regression to yield a final score. We show that SQUIRLS transcends previous state-of-the-art accuracy in classifying splice variants as assessed by rank analysis in simulated exomes, and is significantly faster than competing methods. SQUIRLS provides tabular output files for incorporation into diagnostic pipelines for exome and genome analysis, as well as visualizations that contextualize predicted effects of variants on splicing to make it easier to interpret splice variants in diagnostic settings.

Scheda breve

Scheda completa

Scheda completa (DC)

	Parole chiave
	
				bioinformatics; cryptic splicing; exome sequencing; genome sequencing; machine learning; Mendelian genetics; random forest; sequence logo; splice mutation; splice variant; splicing
			
	Settori scientifico-disciplinari dell'articolo (sola visualizzazione)
	
				Settore INF/01 - Informatica
Settore BIO/18 - Genetica
Settore MED/03 - Genetica Medica
			
	Data di pubblicazione
	
				2-set-2021
			
	Rivista in ANCE
	
				AMERICAN JOURNAL OF HUMAN GENETICS
			
	DOI
	
				https://dx.doi.org/10.1016/j.ajhg.2021.06.014
			
	Tipologia
	
				Article (author)
			
	Appare nelle tipologie:
	
				01 - Articolo su periodico

File in questo prodotto:

File	Dimensione	Formato
SQUIRLS-AJHG.pdf accesso riservato Descrizione: Articolo principale Tipologia: Publisher's version/PDF Dimensione 2.01 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	2.01 MB	Adobe PDF	Visualizza/Apri Richiedi una copia
1-s2.0-S000292972100238X-main.pdf accesso riservato Tipologia: Publisher's version/PDF Dimensione 1.92 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	1.92 MB	Adobe PDF	Visualizza/Apri Richiedi una copia

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/867351

Citazioni

7

41

38

ND

IRIS Institutional Research Information System - AIR Archivio Istituzionale della Ricerca