Mendelian genetic diseases comprise approximatelydescribed disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model,T-ResNet(Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model,TIDE-Var(Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show thatT-ResNetachieves an average Area Under the Precision-Recall Curve (AUPRC) of, comparable to that ofHyperSMURF, a state-of-the-art method. In contrast,TIDE-Varyields significantly better results thanHyperSMURF(AUPRC) and slightly better thanXGBoost, one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.
Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases / F. Stacchietti, M.N.. - In: NEUROCOMPUTING. - ISSN 0925-2312. - (2026), pp. 134682.1-134682.20. [Epub ahead of print] [10.1016/j.neucom.2026.134682]
Tabular implicit deep neural networks ensembles for the prediction of pathogenic genetic variants in Mendelian diseases
F. StacchiettiPrimo
;M. NicoliniSecondo
;E. CasiraghiPenultimo
;G. ValentiniUltimo
2026
Abstract
Mendelian genetic diseases comprise approximatelydescribed disorders, yet the genetic basis is known only for about half of them, and a molecular diagnosis often remains difficult or unresolved. In this context, machine learning methods play a key role. However, identifying pathogenic variants in non-coding regions of the human genome is particularly challenging, as they are vastly outnumbered by neutral variants. This extreme imbalance causes standard machine learning approaches to exhibit a strong predictive bias towards the majority class, significantly limiting their sensitivity. Building on recent advances in imbalance-aware and ensemble learning methods, we propose two novel deep learning models for predicting pathogenic non-coding variants in Mendelian diseases. The first model,T-ResNet(Tabular Residual Neural Network), adopts a modular architecture with residual connections, along with a mini-batch balancing strategy to address class imbalance. This design simplifies hyperparameter optimization while mitigating vanishing-gradient effects. The second model,TIDE-Var(Tabular Implicit Deep neural network Ensembles for Variant prediction), leverages the TabM and BatchEnsemble models to build an implicit ensemble of deep neural networks trained jointly by minimizing a common objective function, and partially sharing learning parameters. Learner-specific adapter parameters promote diversity among the implicit base learners while keeping the ensemble computationally efficient. Genome-wide experiments show thatT-ResNetachieves an average Area Under the Precision-Recall Curve (AUPRC) of, comparable to that ofHyperSMURF, a state-of-the-art method. In contrast,TIDE-Varyields significantly better results thanHyperSMURF(AUPRC) and slightly better thanXGBoost, one of the top-methods for the classification of tabular data. Ablation studies confirm that regularized joint ensemble learning with partially shared and base-learner specific adapter learning parameters are key factors in achieving high predictive performance, essential to improve the diagnostic yield for patients with rare genetic diseases.| File | Dimensione | Formato | |
|---|---|---|---|
|
NeuroComputingTRemm.pdf
accesso aperto
Tipologia:
Post-print, accepted manuscript ecc. (versione accettata dall'editore)
Licenza:
Creative commons
Dimensione
7.23 MB
Formato
Adobe PDF
|
7.23 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




