Accurate vehicle detection in drone imagery remains a critical challenge due to substantial object scale variation and severe feature degradation arising from small target sizes and complex urban environments. To address these challenges, a novel detection framework, termed the structured integrated cascade network (SICNet), is proposed to improve detection performance in aerial scenes, particularly under dense traffic conditions and small-object scenarios. The core methodological innovation of SICNet lies in a structured attention-guided feature interaction paradigm, which explicitly models the coordinated enhancement of global semantic representations and local spatial details across hierarchical features. Specifically, the proposed structured integrated balanced attention path aggregation network (SIBA PAN) reformulates conventional feature fusion by introducing an attention-guided mechanism for efficient semantic–spatial information exchange. Within this architecture, a dual-branch attention module, referred to as the structured integrated bal , anced attention module (SIBAM), is constructed, where the structured grouped attention module and the integrated spatial attention module operate in parallel to capture complementary channel-wise dependencies and spatial responses, respectively, thereby enhancing discriminative feature representation while suppressing background interference. In addition, a spatial cross stage module (SCM) is integrated into the backbone to strengthen multi-scale context modeling through the coordinated use of cross-stage connections, ELAN-style aggregation, and depth wise separable convolutions, achieving efficient feature fusion with minimal computational overhead. Extensive experiments conducted on the VisDrone dataset demonstrate that SICNet achieves a 3.6% improvement in mAP0.5 over the baseline, with notable gains in detection accuracy for small and densely distributed vehicles, thereby validating the effectiveness of the proposed structured attention mechanism and multi-scale feature interaction strategy. The code and dataset are available at https:// github.com/yikuizhai/SICNet.

SICNet: Structured Integrated Cascade Network for UAV-Based Vehicle Detection / W. Qiu, C.D.. - In: IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS. - ISSN 1524-9050. - (2026). [Epub ahead of print] [10.1109/tits.2026.3717450]

SICNet: Structured Integrated Cascade Network for UAV-Based Vehicle Detection

P. Coscia;A. Genovese
Penultimo
;
2026

Abstract

Accurate vehicle detection in drone imagery remains a critical challenge due to substantial object scale variation and severe feature degradation arising from small target sizes and complex urban environments. To address these challenges, a novel detection framework, termed the structured integrated cascade network (SICNet), is proposed to improve detection performance in aerial scenes, particularly under dense traffic conditions and small-object scenarios. The core methodological innovation of SICNet lies in a structured attention-guided feature interaction paradigm, which explicitly models the coordinated enhancement of global semantic representations and local spatial details across hierarchical features. Specifically, the proposed structured integrated balanced attention path aggregation network (SIBA PAN) reformulates conventional feature fusion by introducing an attention-guided mechanism for efficient semantic–spatial information exchange. Within this architecture, a dual-branch attention module, referred to as the structured integrated bal , anced attention module (SIBAM), is constructed, where the structured grouped attention module and the integrated spatial attention module operate in parallel to capture complementary channel-wise dependencies and spatial responses, respectively, thereby enhancing discriminative feature representation while suppressing background interference. In addition, a spatial cross stage module (SCM) is integrated into the backbone to strengthen multi-scale context modeling through the coordinated use of cross-stage connections, ELAN-style aggregation, and depth wise separable convolutions, achieving efficient feature fusion with minimal computational overhead. Extensive experiments conducted on the VisDrone dataset demonstrate that SICNet achieves a 3.6% improvement in mAP0.5 over the baseline, with notable gains in detection accuracy for small and densely distributed vehicles, thereby validating the effectiveness of the proposed structured attention mechanism and multi-scale feature interaction strategy. The code and dataset are available at https:// github.com/yikuizhai/SICNet.
Drone; vehicle detection; feature fusion; attention mechanism
Settore INFO-01/A - Informatica
2026
21-ago-2026
Article (author)
File in questo prodotto:
File Dimensione Formato  
tits26.pdf

accesso riservato

Tipologia: Publisher's version/PDF
Licenza: Nessuna licenza
Dimensione 5.25 MB
Formato Adobe PDF
5.25 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1268498
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact