Mendelian genetic diseases comprise approximatelydescribed disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model,T-ResNet(Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model,TIDE-Var(Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show thatT-ResNetachieves an average Area Under the Precision-Recall Curve (AUPRC) of, comparable to that ofHyperSMURF, a state-of-the-art method. In contrast,TIDE-Varyields significantly better results thanHyperSMURF(AUPRC) and slightly better thanXGBoost, one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.

Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases / F. Stacchietti, M.N.. - In: NEUROCOMPUTING. - ISSN 0925-2312. - (2026), pp. 134682.1-134682.20. [Epub ahead of print] [10.1016/j.neucom.2026.134682]

Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases

F. Stacchietti
Primo
;
M. Nicolini
Secondo
;
E. Casiraghi
Penultimo
;
G. Valentini
Ultimo
2026

Abstract

Mendelian genetic diseases comprise approximatelydescribed disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model,T-ResNet(Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model,TIDE-Var(Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show thatT-ResNetachieves an average Area Under the Precision-Recall Curve (AUPRC) of, comparable to that ofHyperSMURF, a state-of-the-art method. In contrast,TIDE-Varyields significantly better results thanHyperSMURF(AUPRC) and slightly better thanXGBoost, one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.
Deep modular neural networks; Ensembles of neural networks; Pathogenic variant prediction in non-coding genome; Mendelian genetic diseases;
Settore INFO-01/A - Informatica
Settore MEDS-24/A - Statistica medica
   Metodi di AI per diagnostica PRECOce del carcinoma squamoso del Cavo Orale basato su imaging label-free (PRECOCO)
   PRECOCO
   UNIVERSITA' DEGLI STUDI DI MILANO-BICOCCA
2026
1-ago-2026
Article (author)
File in questo prodotto:
File Dimensione Formato  
NeuroComputingTRemm.pdf

accesso aperto

Tipologia: Post-print, accepted manuscript ecc. (versione accettata dall'editore)
Licenza: Creative commons
Dimensione 7.23 MB
Formato Adobe PDF
7.23 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1264755
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact