Automatic depression detection with deep learning has shown promise, but often suffers from limited generalization due to domain shift arising from interspeaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional long short-term memory (BiLSTM) with intramodal and cross-modal attention (CMA) mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by domain-adversarial training of neural networks (DANNs), which promotes domain-invariant representations by adversarially limiting the model’s ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a fivefold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-s segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1 -score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1 -score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.
Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network With Domain-Adversarial Training / A. Tabaraei, F.S.. - In: IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS. - ISSN 2162-237X. - (2026). [Epub ahead of print] [10.1109/tnnls.2026.3714047]
Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network With Domain-Adversarial Training
A. Tabaraei
Primo
;F. SimonettaPenultimo
;S. NtalampirasUltimo
2026
Abstract
Automatic depression detection with deep learning has shown promise, but often suffers from limited generalization due to domain shift arising from interspeaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional long short-term memory (BiLSTM) with intramodal and cross-modal attention (CMA) mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by domain-adversarial training of neural networks (DANNs), which promotes domain-invariant representations by adversarially limiting the model’s ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a fivefold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-s segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1 -score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1 -score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.| File | Dimensione | Formato | |
|---|---|---|---|
|
2607.22794v1_compressed.pdf
accesso aperto
Tipologia:
Pre-print (manoscritto inviato all'editore)
Licenza:
Altro
Dimensione
4.24 MB
Formato
Adobe PDF
|
4.24 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




