To address challenges in UAV aerial imagery, such as complex backgrounds, drastic scale variations, and partial occlusions. We propose DCA-YOLO, a high-performance vehicle detection framework featuring a newly designed neck architecture named the dynamic cross-attention path aggregation network (DCA-PAN). Our approach introduces three synergistic modules in the neck to enhance feature expressiveness for small vehicle detection: dynamic weighted adaptive feature fusion, cross-layer dynamic interaction, and coordinate-aware dynamic attention. These modules synergistically improve multi-scale representation, seamlessly integrate shallow details with deep semantics, and dynamically emphasize key object features. To this end, we leverage the efficient and powerful FasterNet as the backbone to achieve an effective balance between computational overhead and detection accuracy. Experiments on the VisDrone and UAV-HS datasets demonstrate that DCA-YOLO achieves an optimal accuracy-efficiency tradeoff among compact detectors. On the challenging VisDrone dataset, DCA-YOLO significantly outperforms the baseline YOLOv8n, boosting mAP50 from 57.2% to 60.9%, mAP50:95 from 38.5% to 41.6%, and mAPs from 14.3% to 16.1%.
DCA-YOLO: A Dynamic Cross-Attention Network for Small Vehicle Detection in UAV Aerial Imagery / C. Dong, Y.W.. - In: IEEE TRANSACTIONS ON AEROSPACE AND ELECTRONIC SYSTEMS. - ISSN 0018-9251. - (2026), pp. 1-18. [Epub ahead of print] [10.1109/taes.2026.3715596]
DCA-YOLO: A Dynamic Cross-Attention Network for Small Vehicle Detection in UAV Aerial Imagery
P. Coscia;A. GenovesePenultimo
;
2026
Abstract
To address challenges in UAV aerial imagery, such as complex backgrounds, drastic scale variations, and partial occlusions. We propose DCA-YOLO, a high-performance vehicle detection framework featuring a newly designed neck architecture named the dynamic cross-attention path aggregation network (DCA-PAN). Our approach introduces three synergistic modules in the neck to enhance feature expressiveness for small vehicle detection: dynamic weighted adaptive feature fusion, cross-layer dynamic interaction, and coordinate-aware dynamic attention. These modules synergistically improve multi-scale representation, seamlessly integrate shallow details with deep semantics, and dynamically emphasize key object features. To this end, we leverage the efficient and powerful FasterNet as the backbone to achieve an effective balance between computational overhead and detection accuracy. Experiments on the VisDrone and UAV-HS datasets demonstrate that DCA-YOLO achieves an optimal accuracy-efficiency tradeoff among compact detectors. On the challenging VisDrone dataset, DCA-YOLO significantly outperforms the baseline YOLOv8n, boosting mAP50 from 57.2% to 60.9%, mAP50:95 from 38.5% to 41.6%, and mAPs from 14.3% to 16.1%.| File | Dimensione | Formato | |
|---|---|---|---|
|
taes26(1)_compressed.pdf
accesso aperto
Tipologia:
Pre-print (manoscritto inviato all'editore)
Licenza:
Publisher
Dimensione
475.47 kB
Formato
Adobe PDF
|
475.47 kB | Adobe PDF | Visualizza/Apri |
|
DCA-YOLO_A_Dynamic_Cross-Attention_Network_for_Small_Vehicle_Detection_in_UAV_Aerial_Imagery(1).pdf
accesso aperto
Tipologia:
Post-print, accepted manuscript ecc. (versione accettata dall'editore)
Licenza:
Creative commons
Dimensione
19.64 MB
Formato
Adobe PDF
|
19.64 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




