Cross-and-Diagonal Networks: An Indirect Self-Attention Mechanism for Image Classification

Jiahang Lyu; Rongxin Zou; Qin Wan; Wang Xi; Qinglin Yang; Sarath Kodagoda; Shifeng Wang

doi:10.3390/s24072055

Cross-and-Diagonal Networks: An Indirect Self-Attention Mechanism for Image Classification

Sensors (Basel). 2024 Mar 23;24(7):2055. doi: 10.3390/s24072055.

Authors

Jiahang Lyu¹, Rongxin Zou¹, Qin Wan¹, Wang Xi¹, Qinglin Yang¹, Sarath Kodagoda², Shifeng Wang^{1

3}

Affiliations

¹ School of Optoelectronic Engineering, Changchun University of Science and Technology, Changchun 130022, China.
² Faculty of Engineering & Information Technology, University of Technology Sydney, Sydney, NWS 2007, Australia.
³ Zhongshan Institute of Changchun University of Science and Technology, Zhongshan 528400, China.

Abstract

In recent years, computer vision has witnessed remarkable advancements in image classification, specifically in the domains of fully convolutional neural networks (FCNs) and self-attention mechanisms. Nevertheless, both approaches exhibit certain limitations. FCNs tend to prioritize local information, potentially overlooking crucial global contexts, whereas self-attention mechanisms are computationally intensive despite their adaptability. In order to surmount these challenges, this paper proposes cross-and-diagonal networks (CDNet), innovative network architecture that adeptly captures global information in images while preserving local details in a more computationally efficient manner. CDNet achieves this by establishing long-range relationships between pixels within an image, enabling the indirect acquisition of contextual information. This inventive indirect self-attention mechanism significantly enhances the network's capacity. In CDNet, a new attention mechanism named "cross and diagonal attention" is proposed. This mechanism adopts an indirect approach by integrating two distinct components, cross attention and diagonal attention. By computing attention in different directions, specifically vertical and diagonal, CDNet effectively establishes remote dependencies among pixels, resulting in improved performance in image classification tasks. Experimental results highlight several advantages of CDNet. Firstly, it introduces an indirect self-attention mechanism that can be effortlessly integrated as a module into any convolutional neural network (CNN). Additionally, the computational cost of the self-attention mechanism has been effectively reduced, resulting in improved overall computational efficiency. Lastly, CDNet attains state-of-the-art performance on three benchmark datasets for similar types of image classification networks. In essence, CDNet addresses the constraints of conventional approaches and provides an efficient and effective solution for capturing global context in image classification tasks.

Keywords: CNN; computer vision; image classification; self-attention mechanism.

Grants and funding

This work is funded by the International Cooperation Foundation of Jilin Province (20210402074GH).