Accurate vehicle detection in drone imagery remains a critical challenge due to substantial object scale variation and severe feature degradation arising from small target sizes and complex urban environments. To address these challenges, a novel detection framework, termed the structured integrated cascade network (SICNet), is proposed to improve detection performance in aerial scenes, particularly under dense traffic conditions and small-object scenarios. The core methodological innovation of SICNet lies in a structured attention-guided feature interaction paradigm, which explicitly models the coordinated enhancement of global semantic representations and local spatial details across hierarchical features. Specifically, the proposed structured integrated balanced attention path aggregation network (SIBA PAN) reformulates conventional feature fusion by introducing an attention-guided mechanism for efficient semantic–spatial information exchange. Within this architecture, a dual-branch attention module, referred to as the structured integrated bal , anced attention module (SIBAM), is constructed, where the structured grouped attention module and the integrated spatial attention module operate in parallel to capture complementary channel-wise dependencies and spatial responses, respectively, thereby enhancing discriminative feature representation while suppressing background interference. In addition, a spatial cross stage module (SCM) is integrated into the backbone to strengthen multi-scale context modeling through the coordinated use of cross-stage connections, ELAN-style aggregation, and depth wise separable convolutions, achieving efficient feature fusion with minimal computational overhead. Extensive experiments conducted on the VisDrone dataset demonstrate that SICNet achieves a 3.6% improvement in mAP0.5 over the baseline, with notable gains in detection accuracy for small and densely distributed vehicles, thereby validating the effectiveness of the proposed structured attention mechanism and multi-scale feature interaction strategy. The code and dataset are available at https:// github.com/yikuizhai/SICNet.
SICNet: Structured Integrated Cascade Network for UAV-Based Vehicle Detection / W. Qiu, C.D.. - In: IEEE TRANSACTIONS ON INTELLIGENT TRANSPORTATION SYSTEMS. - ISSN 1524-9050. - (2026). [Epub ahead of print] [10.1109/tits.2026.3717450]
SICNet: Structured Integrated Cascade Network for UAV-Based Vehicle Detection
P. Coscia;A. GenovesePenultimo
;
2026
Abstract
Accurate vehicle detection in drone imagery remains a critical challenge due to substantial object scale variation and severe feature degradation arising from small target sizes and complex urban environments. To address these challenges, a novel detection framework, termed the structured integrated cascade network (SICNet), is proposed to improve detection performance in aerial scenes, particularly under dense traffic conditions and small-object scenarios. The core methodological innovation of SICNet lies in a structured attention-guided feature interaction paradigm, which explicitly models the coordinated enhancement of global semantic representations and local spatial details across hierarchical features. Specifically, the proposed structured integrated balanced attention path aggregation network (SIBA PAN) reformulates conventional feature fusion by introducing an attention-guided mechanism for efficient semantic–spatial information exchange. Within this architecture, a dual-branch attention module, referred to as the structured integrated bal , anced attention module (SIBAM), is constructed, where the structured grouped attention module and the integrated spatial attention module operate in parallel to capture complementary channel-wise dependencies and spatial responses, respectively, thereby enhancing discriminative feature representation while suppressing background interference. In addition, a spatial cross stage module (SCM) is integrated into the backbone to strengthen multi-scale context modeling through the coordinated use of cross-stage connections, ELAN-style aggregation, and depth wise separable convolutions, achieving efficient feature fusion with minimal computational overhead. Extensive experiments conducted on the VisDrone dataset demonstrate that SICNet achieves a 3.6% improvement in mAP0.5 over the baseline, with notable gains in detection accuracy for small and densely distributed vehicles, thereby validating the effectiveness of the proposed structured attention mechanism and multi-scale feature interaction strategy. The code and dataset are available at https:// github.com/yikuizhai/SICNet.| File | Dimensione | Formato | |
|---|---|---|---|
|
tits26.pdf
accesso riservato
Tipologia:
Publisher's version/PDF
Licenza:
Nessuna licenza
Dimensione
5.25 MB
Formato
Adobe PDF
|
5.25 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




