Using Improved Transformer Model Considering Multi-domain Features for CT Image Recognitionof Connective Tissue-Associated Interstitial Lung Disease
Gao Jing1, Li Lei2, Zhu Lingyan3*, Qian Tianhe1, Wang Yongfu3
1(School of Artificial Intelligence,Capital University of Economics and Business, Beijing 100070, China) 2(School of Management Engineering,Capital University of Economics and Business,Beijing 100070, China) 3(Key Laboratory of Autoimmunology of Inner Mongolia Autonomous Region,Department of Rheumatology and Immunology,The First Affiliated Hospital of Baotou Medical College,Baotou 014010,Inner Mongolia Autonomous Region, China)
Abstract:Connectivetissue disease-related interstitial lung disease (CTD-ILD) is a type of chronic respiratory disease with an increasing number of patients. Early detection and diagnosis of this disease can effectively improve patients' treatment outcomes and survival rates. Pulmonary imaging is a primary method for the early diagnosis of CTD-ILD. To extract valuable information quickly from many medical images and determine the severity of the disease, an improved Transformer model considering multi-domain features was proposed in this work. This model was based on multi-domain feature collaborative learning, integrating the ResNet and Vision Transformer (ViT). It used the self-attention mechanism in the complex lesion area and depthwise separable convolution in the simple non-lesion area for feature extraction. The experimental dataset comprised 2,752 CTD-ILD lung CT images labeled by professional physicians. The model's validity was systematically evaluated through comparative experiments, ablation studies, and robustness assessments under noisy conditions. On the self-built dataset, the model achieved an accuracy of 96.71 %, precision of 96.74 %, recall of 96.71 %, and F1 score of 96.69 %, representing a 5.83 % accuracy improvement over ResNet50 and outperforming 9 classic models. The average single training epoch took 17 minutes, 5 minutes shorter than the unimproved model, with only 28.7 seconds required to train a single image. Through multi-domain feature fusion and optimization of the ViT architecture, the proposed method has demonstrated excellent performance in medical image recognition for CTD-ILD, achieving a dual improvement in the image recognition performance and training efficiency. This method is expected to assist clinicians in diagnosis and improve diagnostic efficiency.
[1] Barnes H,Holland AE,Westall GP,et al. Cyclophosphamide for connective tissue disease-associated interstitial lung disease [J]. Chest,2018,1(1):CD010908. [2] Ohno Y, Aoyagi K, Takenaka D, et al. Machine learning for lung CT texture analysis:Improvement of inter-observer agreement for radiological finding classification in patients with pulmonary diseases [J]. European Journal of Radiology, 2021, 134: 109410. [3] 张璇,张飞,李铭麟,等. 智能机器人在基层慢性病管理中的应用与挑战 [J]. 中国全科医学, 2025, 28(1):7-12,19. [4] 周涛,刘赟璨,侯森宝,等. REC-ResNet:面向COVID-19辅助诊断的特征增强模型 [J]. 光学精密工程, 2023, 31(14): 2093-2110. [5] 白浩田,谷宇,杨立东,等. 改进知识蒸馏Transformer的新冠肺炎医学影像分类 [J]. 激光杂志, 2024, 45(2): 152-160. [6] Ha PN, Doucet A, Tran GS. Vision transformer for pneumonia classification in X-ray images [C]//Proceedings of the 2023 8th International Conference on Intelligent Information Technology. Da Nang:ACM, 2023: 185-192. [7] Sun Ruina, Pang Yuexin, Li Wenfa. Efficient lung cancer image classification and segmentation algorithm based on an improved swin transformer [J]. Electronics, 2023, 12(4): 1024. [8] 李兴珺,李双蓉,王楠,等. 特发性炎性肌病患者临床特点及发生肺间质病变的危险因素研究 [J]. 中国全科医学, 2024, 27(13): 1623-1629. [9] Chen Jianxun, Shen Yuchen, Peng Shinlei, et al. Pattern classification of interstitial lung diseases from computed tomography images using a ResNet-based network with a split-transform-merge strategy and split attention [J]. Physical and Engineering Sciences in Medicine, 2024: 1-13. [10] Dianat B, La Torraca P, Manfredi A, et al. Classification of pulmonary sounds through deep learning for the diagnosis of interstitial lung diseases secondary to connective tissue diseases [J]. Computers in Biology and Medicine, 2023, 160:106928. [11] Su Ningling, Hou Fan, Zheng Wen, et al. Computed tomography-based deep learning model for assessing the severity of patients with connective tissue disease-associated interstitial lung disease [J]. Journal of Computer Assisted Tomography, 2023, 47(5): 738-745. [12] He Kaiming, Zhang Xiangyu, Ren Shaoqing, et al. Deep residual learning for image recognition [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Las Vegas, IEEE, 2016: 770-778. [13] 王贺,张震. 基于ResNeSt和改进Transformer的多标签图像分类算法 [J]. 测试技术学报, 2024, 38(1): 48-53. [14] Dosovitskiy A, Beyer L, Kolesnikov A, et al. An image is worth 16x16 words: transformers for image recognition at scale[J/OL]. https://arxiv.org/abs/2010.11929, 2020-10-22/2025-03-16. [15] Deng Jia, Dong Wei, Socher R, et al. Imagenet:A large-scale hierarchical image database [C]//2009 IEEE Conference on Computer Vision and Pattern Recognition. Florida, IEEE, 2009: 248-255. [16] Zhuang Fuzhen, Qi Zhiyuan, Duan Keyu, et al. A comprehensive survey on transfer learning [J]. Proceedings of the IEEE, 2020, 109(1): 43-76. [17] 唐子猗,易婷,王聃,等. 我国川东北地区多发性肌炎/皮肌炎合并间质性肺疾病患者的临床特征及其影响因素研究 [J]. 中国全科医学, 2019, 22(16): 1960-1965. [18] Krizhevsky A, Sutskever I, Hinton GE. Imagenet classification with deep convolutional neural networks [J]. Communications of the ACM, 2017, 60(6): 84-90. [19] Simonyan K, Zisserman A. Very deep convolutional networks for large-scale image recognition[J/OL]. https://arxiv.org/abs/1409.1556, 2015-04-10/2025-04-10. [20] Szegedy C, Liu W, Jia Y Q, et al. Going deeper with convolutions [C]//Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition. Boston, IEEE, 2015: 1-9. [21] Huang Gao, Liu Zhuang, Van Der Maaten L, et al. Densely connected convolutional networks [C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Hawaii, IEEE, 2017: 4700-4708. [22] Wang Yan, Liu Yi, Zhao Shijie, et al. CAMixerSR: Only details need more “attention” [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Seattle, IEEE, 2024: 25837-25846. [23] Zhu Jiachen, Chen Xinlei, He Kaiming, et al. Transformers without normalization[J/OL]. https://arxiv.org/abs/2503.10622, 2025-03-17/2025-04-28. [24] Liu Zhuang, Mao Hanzi, Wu Chaoyuan, et al. A convnet for the 2020s [C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. New Orleans, IEEE, 2022: 11976-11986. [25] Chen Chunfu, Fan Quanfu, Panda R. Crossvit: Cross-attention multi-scale vision transformer for image classification [C]//Proceedings of the IEEE/CVF International Conference on Computer Vision. Montreal, IEEE, 2021: 357-366. [26] 王珂,巨文璇,张萌,等. CTD-ILD患者多参数肺损伤风险评估模型构建与内部验证 [J]. CT理论与应用研究, 2025.113:1-10. [27] 冯灿,史卫亚,李岩超,等. 基于CNN和Transformer双编码器的皮肤病变分割算法 [J]. 科学技术与工程, 2025, 25(23): 9900-9910. [28] 任宇,杨鹏,范小琴,等. 基于轻量级多尺度CNN-Transformer网络的鼻咽癌诊断方法 [J].中国生物医学工程学报, 2025, 44(3): 279-290.