Automatic depression detection with deep learning has shown promise, but often suffers from limited generalization due to domain shift arising from interspeaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional long short-term memory (BiLSTM) with intramodal and cross-modal attention (CMA) mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by domain-adversarial training of neural networks (DANNs), which promotes domain-invariant representations by adversarially limiting the model’s ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a fivefold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-s segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1 -score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1 -score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network With Domain-Adversarial Training / A. Tabaraei, F.S.. - In: IEEE TRANSACTIONS ON NEURAL NETWORKS AND LEARNING SYSTEMS. - ISSN 2162-237X. - (2026). [Epub ahead of print] [10.1109/tnnls.2026.3714047]

Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network With Domain-Adversarial Training

A. Tabaraei
Primo
;
F. Simonetta
Penultimo
;
S. Ntalampiras
Ultimo
2026

Abstract

Automatic depression detection with deep learning has shown promise, but often suffers from limited generalization due to domain shift arising from interspeaker variability. To address this critical issue, we present the first patient-independent multimodal depression detection framework that incorporates domain generalization (DG), jointly leveraging both acoustic and textual modalities. The proposed model integrates bidirectional long short-term memory (BiLSTM) with intramodal and cross-modal attention (CMA) mechanisms, accompanied by segment-level fusion for decision-making. Generalization is further enhanced by applying a gradient reversal layer inspired by domain-adversarial training of neural networks (DANNs), which promotes domain-invariant representations by adversarially limiting the model’s ability to identify individual speakers, effectively reducing patient-specific bias. Conducting experiments on the Androids-Corpus dataset with a fivefold cross-validation (CV) protocol, various pairings of audio and text feature extractors were evaluated over different segment durations, determining MelSpec and ItalianBERT as the optimal baseline at a 30-s segment duration. The addition of DG to this baseline yields a 2.5% increase in accuracy and 3.3% in F1 -score, achieving 93.2% accuracy, 93.2% precision, 96.2% recall, and 94.2% F1 -score, surpassing all existing benchmarks. Extensive ablation studies assess the impact of multimodal fusion, deep architectural choices, and DG, highlighting their combined contribution to robust and generalizable depression detection.
Adversarial training; audio–text analysis; depression detection; domain generalization (DG); multimodal learning
Settore INFO-01/A - Informatica
2026
24-lug-2026
Article (author)
File in questo prodotto:
File Dimensione Formato  
2607.22794v1_compressed.pdf

accesso aperto

Tipologia: Pre-print (manoscritto inviato all'editore)
Licenza: Altro
Dimensione 4.24 MB
Formato Adobe PDF
4.24 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1266755
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
  • OpenAlex 0
social impact