Article(id=1279511861326484060, tenantId=1146029695717560320, journalId=1278651732997652489, issueId=1279511628118986881, articleNumber=null, orderNo=null, doi=10.12086/oee.2026.250292, pmid=null, cstr=32245.14.oee.2026.250292, oa=null, hot=null, price=null, onlineType=0, articleFormat=0, articleType=null, articleTypeStr=null, receivedDate=1758902400000, receivedDateStr=2025-09-27, revisedDate=1769529600000, revisedDateStr=2026-01-28, acceptedDate=1768492800000, acceptedDateStr=2026-01-16, onlineDate=1782988999921, onlineDateStr=2026-07-02, pubDate=1776960000000, pubDateStr=2026-04-24, doiRegisterDate=null, doiRegisterDateStr=null, onlineIssueDate=1782988999921, onlineIssueDateStr=2026-07-02, onlineJustAcceptDate=null, onlineJustAcceptDateStr=null, onlineFirstDate=null, onlineFirstDateStr=null, sourceXml=null, magXml=null, createTime=1782988999921, creator=13701087609, updateTime=1782988999921, updator=13701087609, issue=Issue{id=1279511628118986881, tenantId=1146029695717560320, journalId=1278651732997652489, year='2026', volume='53', issue='4', pageStart='250244', pageEnd='250340', issueExtLink='null', onlineDate='null', pubDate='1776960000000', pubDateStr='2026-04-24', beforeIssueId=null, nextIssueId=null, price=null, status=1, issueComplete=1, articleOrder=1, issueType=1, specialIssue=null, createTime=1782988944320, creator='13701087609', updateTime=1782988944320, updator='13701087609', preIssue=null, nextIssue=null, articleTotal=null, ext=null, issueFiles=null, downloadFileDto=null}, startPage=250292, endPage=, ext={EN=ArticleExt(id=1279511863645934173, articleId=1279511861326484060, tenantId=1146029695717560320, journalId=1278651732997652489, language=EN, title=YOLO-based adaptive multi-scale infrared target detection network, columnId=1279511634116841602, journalTitle=Opto-Electronic Engineering, columnName=Article, runingTitle=null, highlight=null, articleAbstract=
Objective

Infrared target detection plays an indispensable role in numerous critical domains, such as security surveillance, autonomous driving, and military reconnaissance, owing to its unique perceptual capability under complex environments (e.g., low-light conditions and severe weather). However, infrared images inherently suffer from low contrast, blurred details, and significant noise interference, which often lead to ambiguous target edges, missing texture features, and other challenges during the detection process. Existing deep learning-based infrared target detection algorithms (ITDA) exhibit inadequate performance in feature extraction and processing for infrared images, resulting in relatively high rates of missed detection and false detection. Moreover, our systematic analysis of infrared target detection tasks reveals that algorithms tailored for small infrared targets rely heavily on high-sensitivity feature extraction to capture subtle characteristics. Nevertheless, as the scale of detected targets increases, these algorithms tend to encounter overfitting to local textures and elevated false detection rates, thereby degrading overall performance. In practical applications, detection environments are dynamically changing with targets of varying scales; thus, multi-scale detection capability is critical to ensuring algorithms maintain high reliability and adaptability in complex real-world scenarios. Unfortunately, most state-of-the-art algorithms are optimized for single-scale targets, making it challenging to simultaneously satisfy the requirements of high-precision localization for small targets and effective semantic understanding for large targets.

Methods

To address the above issues, this paper proposes an adaptive multi-scale infrared target detection network based on YOLO (AFITDYOLO). This network is designed to receive infrared target images of different scales and employs a multi-layer feature extraction module and a multi-layer feature fusion module to enhance its multi-scale infrared target detection capability. Firstly, a multi-scale feature fusion module (MFFM) is proposed. This module enhances the correlation between features of different layers in the feature pyramid network (FPN), coordinates deep semantic features with shallow spatial detail features more effectively, and thereby improves the representational ability of multi-scale feature fusion. Secondly, a multi-kernel feature extraction convolution (MFEConv) is constructed. By utilizing heterogeneous convolution groups, MFEConv expands the receptive field and strengthens the model's feature extraction capability. Additionally, a cross-attention fusion module (CAFM) is designed. Through the comparative interaction of feature maps output by different layers in the detection network, CAFM leverages the complementary information among these feature maps to suppress infrared noise in images and further enhance feature representation capability.

Results and Discussions

To validate the effectiveness of the proposed method in improving detection performance, extensive training and evaluation are conducted on the CTIR dataset, which comprises road pedestrians and vehicles with multi-scale infrared targets. To further verify the adaptability of the method, additional experiments are performed on the SIRST-UAVB dataset—a single-frame UAV bird dataset characterized by more complex backgrounds and smaller target scales. Experimental results on these two datasets demonstrate that AFITDYOLO achieves mean average precision at 50% intersection over union (mAP50) of 88.9% and 90.7%, respectively, representing significant improvements of 5.6% and 6.5% compared with YOLOv10n. In terms of lightweight optimization, the proposed method achieves higher inference speed (measured in frames per second, FPS) while utilizing fewer model parameters (params) and floating-point operations (FLOPs). When compared with current mainstream methods, AFITDYOLO exhibits the highest detection accuracy, the lowest parameter count and FLOPs, and the fastest inference speed, demonstrating distinct advantages. Additionally, to evaluate the generalization ability of the proposed method, cross-dataset experiments are carried out on the HIT-UAV dataset (a high-altitude UAV infrared thermal imaging dataset) and the IRSTD-1k dataset (a classic infrared small target dataset). Experimental results indicate that while the precision (P) value of AFITDYOLO is slightly inferior to that of DEIM-N, it outperforms all other mainstream methods in all remaining evaluation metrics. These findings confirm that the proposed method achieves improved detection accuracy on the infrared datasets used in the generalization experiments, validating its strong generalization capability and further demonstrating its feasibility for cross-scenario deployment. Overall, the proposed method simultaneously achieves enhanced detection accuracy and lightweight optimization of the detection model, fully meeting the requirements of real-time detection applications.

Conclusions

The AFITDYOLO network proposed in this paper, which is an adaptive multi-scale infrared target detection network based on YOLO, enhances the detection accuracy of infrared targets of different scales under various backgrounds with a relatively small number of parameters. The proposed MFFM enhances the model's representational ability in multi-scale feature fusion by improving the correlation between features of different layers in FPN. Additionally, the lightweight convolution module MFEConv is designed to achieve an efficient and larger receptive field with minimal parameters by leveraging the target distribution characteristics of infrared images. Furthermore, the CAFM is introduced to highlight important feature information, filter out irrelevant background information, and suppress noise through the comparative interaction of feature maps output by different layers, thereby further boosting the model's feature representation capability. Experimental results demonstrate that the proposed method outperforms current mainstream algorithms, exhibiting excellent detection accuracy, lightweight performance, and generalization ability, along with cross-scenario deployment capabilities.

, authors=Jiaxu Wang1, 2, Jun Yang2, *, Congyuan Xu2, authorsList=Jiaxu Wang, Jun Yang, Congyuan Xu, authorCompany=null, correspAuthors=Jun Yang, authorNote=null, correspAuthorsNote=
, copyrightStatement=Copyright © 2026 Opto-Electronic Engineering. All rights reserved., copyrightOwner=null, extLink=null, articleAbsUrl=null, sourceXml=null, magXml=null, pdfUrl=null, pdf=null, pdfFileSize=null, pdfExtLink=null, richHtmlUrl=null, mobilePdfUrl=null, reviewReport=null, pdfFirstPage=null, abstractGraph=null, abstractGraphContent=null, abstractVideo=null, citation=null, cebUrl=null, magXmlContent=null, mapNumber=null, fund=null), CN=ArticleExt(id=1279511926011040537, articleId=1279511861326484060, tenantId=1146029695717560320, journalId=1278651732997652489, language=CN, title=基于YOLO的自适应多尺度红外目标检测网络, columnId=1279511637858160772, journalTitle=光电工程, columnName=科研论文, runingTitle=null, highlight=null, articleAbstract=

现有基于深度学习的红外目标检测算法(infrared target detection algorithm, ITDA)对红外图像中特征的提取和处理能力不足,导致漏检、误检率较高。多数算法不能兼顾不同尺度的红外目标,红外小目标检测依赖高灵敏度特征提取以捕捉细微特征,当检测目标尺度增大时,因需要更强的全局理解能力,易出现局部纹理过拟合,导致性能降低,使算法在跨场景部署时准确性降低。为解决上述问题,本文提出了一种基于YOLO的自适应多尺度红外目标检测的网络(adaptive multi-scale infrared target detection network based on YOLO, AFITDYOLO),该网络通过多层特征提取模块及多层特征融合模块增强多尺度红外目标检测能力。提出多尺度特征融合模块MFFM,通过增强特征金字塔中不同层特征的相关性,提升了多尺度特征融合的表达能力;提出多核特征提取卷积MFEConv,通过异构卷积组增大感受野,更好地与不同尺度目标的空间分布保持一致;提出交叉注意力融合模块CAFM,通过网络中不同层输出特征图的对比交互,增强重要特征信息,提升特征表达能力。为了评估AFITDYOLO的性能,在无人机飞鸟数据集SIRST-UAVB和道路行人车辆数据集CTIR上进行实验,mAP50分别达到了88.9%和90.7%,对比YOLOv10n分别提高了5.6%和6.5%,与目前主流方法对比,本文方法检测精度最佳。本文提出的红外目标检测算法,在多尺度红外目标检测中,展现了优异的准确性和适应性,可跨场景部署。

, authors=汪佳旭1, 2, 杨俊2, *, 许聪源2, authorsList=汪佳旭, 杨俊, 许聪源, authorCompany=null, correspAuthors=杨俊, authorNote=

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

, correspAuthorsNote=
杨俊,
, copyrightStatement=版权所有©《光电工程》编辑部 2026, copyrightOwner=null, extLink=null, articleAbsUrl=null, sourceXml=zDitpQhCHxbvZvLZXdryfQ==, magXml=3flkh/pzlUr/NbwtEOvJ5w==, pdfUrl=null, pdf=gKSSR+mJkPwRXSzVMhZVkw==, pdfFileSize=9338094, pdfExtLink=null, richHtmlUrl=null, mobilePdfUrl=null, reviewReport=null, pdfFirstPage=null, abstractGraph=Ggvg/di8T6ytS58Y6h6VLA==, abstractGraphContent=null, abstractVideo=null, citation=null, cebUrl=null, magXmlContent=9z39xJ6QYnf4qf7V1D78+A==, mapNumber=null, fund=null)}, authors=[Author(id=1280951146235802136, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=0, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=2936598421@qq.com, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1280951146323882523, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146235802136, language=EN, stringName=Jiaxu Wang, firstName=Jiaxu, middleName=null, lastName=Wang, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=1, 2, address=1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China
2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146399379996, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146235802136, language=CN, stringName=汪佳旭, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=1, 2, address=1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018
2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"RDKWMmEB7nCI4VxjQJrwEQ==","content":"

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

"}, bioImg=RDKWMmEB7nCI4VxjQJrwEQ==, bioContent=

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146026086929, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=1, ext=[AuthorCompanyExt(id=1280951146034475538, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China), AuthorCompanyExt(id=1280951146042864147, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018)]), AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])]), Author(id=1280951146466488862, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=1, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=juneryoung@zjxu.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=1, authorType=1, ext={EN=AuthorExt(id=1280951146541986336, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146466488862, language=EN, stringName=Jun Yang, firstName=Jun, middleName=null, lastName=Yang, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, *, address=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146609095201, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146466488862, language=CN, stringName=杨俊, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, *, address=2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"uHEHZcx2EFUFfNFDYfM2/w==","content":"

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

"}, bioImg=uHEHZcx2EFUFfNFDYfM2/w==, bioContent=

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])]), Author(id=1280951146676204067, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=2, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=cyxu@zjxu.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1280951146890113573, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146676204067, language=EN, stringName=Congyuan Xu, firstName=Congyuan, middleName=null, lastName=Xu, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, address=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146953028134, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146676204067, language=CN, stringName=许聪源, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, address=2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"jDy0t5HTKKb3mneNkFo5fQ==","content":"

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

"}, bioImg=jDy0t5HTKKb3mneNkFo5fQ==, bioContent=

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])])], keywords=[Keyword(id=1280951147078857255, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, orderNo=1, keyword=infrared target detection), Keyword(id=1280951147158549032, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, orderNo=2, keyword=feature fusion), Keyword(id=1280951147221463593, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, orderNo=3, keyword=lightweight network architecture), Keyword(id=1280951147280183850, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, orderNo=4, keyword=attention mechanism), Keyword(id=1280951147334709803, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, orderNo=5, keyword=YOLO), Keyword(id=1280951147414401580, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, orderNo=1, keyword=红外目标检测), Keyword(id=1280951147485704749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, orderNo=2, keyword=特征融合), Keyword(id=1280951147557007918, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, orderNo=3, keyword=轻量级网络结构), Keyword(id=1280951147645088303, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, orderNo=4, keyword=注意力机制), Keyword(id=1280951147708002864, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, orderNo=5, keyword=YOLO)], refs=[Reference(id=1280951151143137861, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=1, rfOrder=0, authorNames=null, journalName=null, refType=null, unstructuredReference=Liu R M, Lu Y H, Gong C L, et al. Infrared point target detection with improved template matching[J]. Infrared Phys Technol, 2012, 55(4): 380−387., articleTitle=null, refAbstract=null), Reference(id=1280951151239606854, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=2, rfOrder=1, authorNames=null, journalName=null, refType=null, unstructuredReference=Rivest J F, Fortin R. Detection of dim targets in digital infrared imagery by morphological image processing[J]. Opt Eng, 1996, 35(7): 1886−1893., articleTitle=null, refAbstract=null), Reference(id=1280951151319298631, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=3, rfOrder=2, authorNames=null, journalName=null, refType=null, unstructuredReference=Dai Y M, Wu Y Q, Zhou F, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2021, 59(11): 9813−9824., articleTitle=null, refAbstract=null), Reference(id=1280951151386407496, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=4, rfOrder=3, authorNames=null, journalName=null, refType=null, unstructuredReference=Du P, Hamdulla A. Infrared small target detection using homogeneity-weighted local contrast measure[J]. IEEE Geosci Remote Sensing Lett, 2020, 17(3): 514−518., articleTitle=null, refAbstract=null), Reference(id=1280951151453516361, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=5, rfOrder=4, authorNames=null, journalName=null, refType=null, unstructuredReference=Liu Y J, Liu X Y, Hao X Y, et al. Single-frame infrared small target detection by high local variance, low-rank and sparse decomposition[J]. IEEE Trans Geosci Remote Sensing, 2023, 61: 5614317., articleTitle=null, refAbstract=null), Reference(id=1280951151541596746, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=6, rfOrder=5, authorNames=null, journalName=null, refType=null, unstructuredReference=Li R H, Shen Y. YOLOSR-IST: a deep learning method for small target detection in infrared remote sensing images based on super-resolution and YOLO[J]. Signal Proc, 2023, 208: 108962., articleTitle=null, refAbstract=null), Reference(id=1280951151612899915, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=7, rfOrder=6, authorNames=null, journalName=null, refType=null, unstructuredReference=Yang B, Zhang X Y, Zhang J, et al. EFLNet: enhancing feature learning network for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2024, 62: 5906511., articleTitle=null, refAbstract=null), Reference(id=1280951151688397388, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=8, rfOrder=7, authorNames=null, journalName=null, refType=null, unstructuredReference=GUPTA A, GUPTA U. Real time target detection for infrared images[C]//2020 Fourth International Conference on Inventive Systems and Control (ICISC), 2020: 570–574. https://doi.org/10.1109/ICISC47916.2020.9171208., articleTitle=null, refAbstract=null), Reference(id=1280951151751311949, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=9, rfOrder=8, authorNames=null, journalName=null, refType=null, unstructuredReference=Yang J N, Liu S L, Wu J J, et al. Pinwheel-shaped convolution and scale-based dynamic loss for infrared small target detection[C]//Proceedings of the 39th AAAI Conference on Artificial Intelligence, 2025: 9202–9210. https://doi.org/10.1609/aaai.v39i9.32996., articleTitle=null, refAbstract=null), Reference(id=1280951151835198030, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=10, rfOrder=9, authorNames=null, journalName=null, refType=null, unstructuredReference=赵斌, 王春平, 付强, 等. 基于深度注意力机制的多尺度红外行人检测[J]. 光学学报, 2020, 40(5): 0504001., articleTitle=null, refAbstract=null), Reference(id=1280951151902306895, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=10, rfOrder=10, authorNames=null, journalName=null, refType=null, unstructuredReference=Zhao B, Wang C P, Fu Q, et al. Multi-scale infrared pedestrian detection based on deep attention mechanism[J]. Acta Opt Sin, 2020, 40(5): 0504001., articleTitle=null, refAbstract=null), Reference(id=1280951151969415760, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=11, rfOrder=11, authorNames=null, journalName=null, refType=null, unstructuredReference=Dalal N, Triggs B. Histograms of oriented gradients for human detection[C]//2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2005: 886–893. https://doi.org/10.1109/CVPR.2005.177., articleTitle=null, refAbstract=null), Reference(id=1280951152044913233, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=12, rfOrder=12, authorNames=null, journalName=null, refType=null, unstructuredReference=何洪英, 姚建刚, 蒋正龙, 等. 基于支持向量机的高压绝缘子污秽等级红外热像检测[J]. 电力系统自动化, 2005, 29(24): 70−74,82., articleTitle=null, refAbstract=null), Reference(id=1280951152112022098, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=12, rfOrder=13, authorNames=null, journalName=null, refType=null, unstructuredReference=He H Y, Yao J G, Jiang Z L, et al. Infrared thermal image detecting of high voltage insulator contamination grades based on support vector machine[J]. Autom Electr Power Syst, 2005, 29(24): 70−74,82., articleTitle=null, refAbstract=null), Reference(id=1280951152183325267, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=13, rfOrder=14, authorNames=null, journalName=null, refType=null, unstructuredReference=Felzenszwalb P F, Girshick R B, McAllester D, et al. Object detection with discriminatively trained part-based models[J]. IEEE Trans Pattern Anal Mach Intell, 2010, 32(9): 1627−1645., articleTitle=null, refAbstract=null), Reference(id=1280951152275599956, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=14, rfOrder=15, authorNames=null, journalName=null, refType=null, unstructuredReference=Mathieu M, LeCun Y, Fergus R, et al. OverFeat: integrated recognition, localization and detection using convolutional networks[C]//International Conference on Learning Representations, 2014., articleTitle=null, refAbstract=null), Reference(id=1280951152380457557, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=15, rfOrder=16, authorNames=null, journalName=null, refType=null, unstructuredReference=Liu S, Qi L, Qin H F, et al. Path aggregation network for instance segmentation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018: 8759–8768. https://doi.org/10.1109/CVPR.2018.00913., articleTitle=null, refAbstract=null), Reference(id=1280951152443372118, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=16, rfOrder=17, authorNames=null, journalName=null, refType=null, unstructuredReference=Ghiasi G, Lin T Y, Le Q V. NAS-FPN: learning scalable feature pyramid architecture for object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019: 7029–7038. https://doi.org/10.1109/CVPR.2019.00720., articleTitle=null, refAbstract=null), Reference(id=1280951152506286679, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=17, rfOrder=18, authorNames=null, journalName=null, refType=null, unstructuredReference=Tan M X, Pang R M, Le Q V. EfficientDet: scalable and efficient object detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020: 10778–10787. https://doi.org/10.1109/CVPR42600.2020.01079., articleTitle=null, refAbstract=null), Reference(id=1280951152573395544, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=18, rfOrder=19, authorNames=null, journalName=null, refType=null, unstructuredReference=Min K, Lee G H, Lee S W. Attentional feature pyramid network for small object detection[J]. Neural Netw, 2022, 155: 439−450., articleTitle=null, refAbstract=null), Reference(id=1280951152644698713, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=19, rfOrder=20, authorNames=null, journalName=null, refType=null, unstructuredReference=Dai Y M, Li X, Zhou F, et al. One-stage cascade refinement networks for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2023, 61: 5000917., articleTitle=null, refAbstract=null), Reference(id=1280951152711807578, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=20, rfOrder=21, authorNames=null, journalName=null, refType=null, unstructuredReference=Dai Y M, Wu Y Q, Zhou F, et al. Asymmetric contextual modulation for infrared small target detection[C]. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 2021: 949–958. https://doi.org/10.1109/WACV48630.2021.00099., articleTitle=null, refAbstract=null), Reference(id=1280951152774722139, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=21, rfOrder=22, authorNames=null, journalName=null, refType=null, unstructuredReference=Ren D D, Li J B, Han M, et al. DNANet: dense nested attention network for single image dehazing[C]//ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021: 2035–2039. https://doi.org/10.1109/ICASSP39728.2021.9414179., articleTitle=null, refAbstract=null), Reference(id=1280951152833442396, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=22, rfOrder=23, authorNames=null, journalName=null, refType=null, unstructuredReference=Fan W Q, Xu X M, Cai B L, et al. ISNet: individual standardization network for speech emotion recognition[J]. IEEE/ACM Trans Audio Sp Lang Process, 2022, 30: 1803−1814., articleTitle=null, refAbstract=null), Reference(id=1280951152913134173, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=23, rfOrder=24, authorNames=null, journalName=null, refType=null, unstructuredReference=Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module[C]//15th European Conference on Computer Vision, 2018: 3–19. https://doi.org/10.1007/978-3-030-01234-2_1., articleTitle=null, refAbstract=null), Reference(id=1280951152980243038, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=24, rfOrder=25, authorNames=null, journalName=null, refType=null, unstructuredReference=Dai T, Cai J R, Zhang Y B, et al. Second-order attention network for single image super-resolution[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019: 11057–11066. https://doi.org/10.1109/CVPR.2019.01132., articleTitle=null, refAbstract=null), Reference(id=1280951153047351903, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=25, rfOrder=26, authorNames=null, journalName=null, refType=null, unstructuredReference=Dai T, Zha H, Jiang Y, et al. Image super-resolution via residual block attention networks[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019: 3879–3886. https://doi.org/10.1109/ICCVW.2019.00481., articleTitle=null, refAbstract=null), Reference(id=1280951153135432288, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=26, rfOrder=27, authorNames=null, journalName=null, refType=null, unstructuredReference=Niu B, Wen W L, Ren W Q, et al. Single image super-resolution via a holistic attention network[C]//16th European Conference on Computer Vision – ECCV 2020, 2020: 191–207. https://doi.org/10.1007/978-3-030-58610-2_12., articleTitle=null, refAbstract=null), Reference(id=1280951153198346849, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=27, rfOrder=28, authorNames=null, journalName=null, refType=null, unstructuredReference=Wang L, Shen J, Tang E, et al. Multi-scale attention network for image super-resolution[J]. J Vis Commun Image Repres, 2021, 80: 103300., articleTitle=null, refAbstract=null), Reference(id=1280951153261261410, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=28, rfOrder=29, authorNames=null, journalName=null, refType=null, unstructuredReference=Varghese R, M S. YOLOv8: a novel object detection algorithm with enhanced performance and robustness[C]//2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), 2024: 1–6. https://doi.org/10.1109/ADICS58448.2024.10533619., articleTitle=null, refAbstract=null), Reference(id=1280951154934788707, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=29, rfOrder=30, authorNames=null, journalName=null, refType=null, unstructuredReference=Terven J, Córdova-Esparza D M, Romero-González J A. A comprehensive review of YOLO architectures in computer vision: from YOLOv1 to YOLOv8 and YOLO-NAS[J]. Mach Learn Knowl Extr, 2023, 5(4): 1680−1716., articleTitle=null, refAbstract=null), Reference(id=1280951155010286180, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=30, rfOrder=31, authorNames=null, journalName=null, refType=null, unstructuredReference=Chen H, Chen K, Ding G G, et al. YOLOv10: real-time end-to-end object detection[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024: 107984–108011. https://doi.org/10.52202/079017-3429., articleTitle=null, refAbstract=null), Reference(id=1280951155081589349, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=31, rfOrder=32, authorNames=null, journalName=null, refType=null, unstructuredReference=Zhao F, Li S J, Zhang J J, et al. Convolution transformer fusion splicing network for hyperspectral image classification[J]. IEEE Geosci Remote Sensing Lett, 2023, 20: 5501005., articleTitle=null, refAbstract=null), Reference(id=1280951155152892518, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=32, rfOrder=33, authorNames=null, journalName=null, refType=null, unstructuredReference=Glorot X, Bengio Y. Understanding the difficulty of training deep feedforward neural networks[C]//Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, 2010: 249–256., articleTitle=null, refAbstract=null), Reference(id=1280951155215807079, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=33, rfOrder=34, authorNames=null, journalName=null, refType=null, unstructuredReference=Ioffe S, Szegedy C. Batch normalization: accelerating deep network training by reducing internal covariate shift[C]//Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37, 2015: 448–456., articleTitle=null, refAbstract=null), Reference(id=1280951155295498856, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=34, rfOrder=35, authorNames=null, journalName=null, refType=null, unstructuredReference=Bieder F, Sandkühler R, Cattin P C. Comparison of methods generalizing max- and average-pooling[Z]. arXiv: 2103.01746, 2021. https://doi.org/10.48550/arXiv.2103.01746., articleTitle=null, refAbstract=null), Reference(id=1280951155362607721, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=35, rfOrder=36, authorNames=null, journalName=null, refType=null, unstructuredReference=Kusupati A, Ramanujan V, Somani R, et al. Soft threshold weight reparameterization for learnable sparsity[C]//Proceedings of the 37th International Conference on Machine Learning, 2020: 5544–5555., articleTitle=null, refAbstract=null), Reference(id=1280951155429716586, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=36, rfOrder=37, authorNames=null, journalName=null, refType=null, unstructuredReference=Chollet F. Xception: deep learning with depthwise separable convolutions[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017: 1800–1807. https://doi.org/10.1109/CVPR.2017.195., articleTitle=null, refAbstract=null), Reference(id=1280951155509408363, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=37, rfOrder=38, authorNames=null, journalName=null, refType=null, unstructuredReference=Xu W, Wan Y. ELA: efficient local attention for deep convolutional neural networks[Z]. arXiv: 2403.01123, 2024. https://doi.org/10.48550/arXiv.2403.01123., articleTitle=null, refAbstract=null), Reference(id=1280951155589100140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=38, rfOrder=39, authorNames=null, journalName=null, refType=null, unstructuredReference=Fløystad G. Shift modules, strongly stable ideals, and their dualities[J]. Trans Am Math Soc Ser B, 2023, 10(21): 670−714., articleTitle=null, refAbstract=null), Reference(id=1280951155656209005, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=39, rfOrder=40, authorNames=null, journalName=null, refType=null, unstructuredReference=Zhang X, Song Y Z, Song T T, et al. AKConv: convolutional kernel with arbitrary sampled shapes and arbitrary number of parameters[Z]. arXiv: 2311.11587, 2023. https://doi.org/10.48550/arXiv.2311.11587., articleTitle=null, refAbstract=null), Reference(id=1280951155735900782, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=40, rfOrder=41, authorNames=null, journalName=null, refType=null, unstructuredReference=Luo Y, Wong Y, Kankanhalli M, et al. G-softmax: improving intraclass compactness and interclass separability of features[J]. IEEE Trans Neural Netw Learn Syst, 2020, 31(2): 685−699., articleTitle=null, refAbstract=null), Reference(id=1280951155815592559, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=41, rfOrder=42, authorNames=null, journalName=null, refType=null, unstructuredReference=Zhao X F, Zhang W W, Zhang H, et al. ITD-YOLOv8: an infrared target detection model based on YOLOv8 for unmanned aerial vehicles[J]. Drones, 2024, 8(4): 161., articleTitle=null, refAbstract=null), Reference(id=1280951155895284336, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=42, rfOrder=43, authorNames=null, journalName=null, refType=null, unstructuredReference=Li Y X, Zou W B, Wei Q M, et al. Multi-level feature fusion network for lightweight stereo image super-resolution[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024: 6489–6498. https://doi.org/10.1109/CVPRW63382.2024.00649., articleTitle=null, refAbstract=null), Reference(id=1280951155979170417, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=43, rfOrder=44, authorNames=null, journalName=null, refType=null, unstructuredReference=Wang X, He N, Hong C, et al. Improved YOLOX-X Based UAV aerial photography object detection algorithm[J]. Image Vision Comput, 2023, 135: 104697., articleTitle=null, refAbstract=null), Reference(id=1280951156050473586, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=44, rfOrder=45, authorNames=null, journalName=null, refType=null, unstructuredReference=Maity M, Banerjee S, Chaudhuri S S. Faster R-CNN and YOLO based vehicle detection: a survey[C]//2021 5th International Conference on Computing Methodologies and Communication (ICCMC), 2021: 1442–1447. https://doi.org/10.1109/ICCMC51019.2021.9418274., articleTitle=null, refAbstract=null), Reference(id=1280951156113388147, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=45, rfOrder=46, authorNames=null, journalName=null, refType=null, unstructuredReference=Sani A R, Zolfagharian A, Kouzani A Z. Automated defects detection in extrusion 3D printing using YOLO models[J]. J Intell Manuf, 2024. https://doi.org/10.1007/s10845-024-02543-8., articleTitle=null, refAbstract=null), Reference(id=1280951156188885620, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=46, rfOrder=47, authorNames=null, journalName=null, refType=null, unstructuredReference=Alif M A R, Hussain M. YOLOv12: a breakdown of the key architectural features[Z]. arXiv: 2502.14740, 2025. https://doi.org/10.48550/arXiv.2502.14740., articleTitle=null, refAbstract=null), Reference(id=1280951156268577397, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=47, rfOrder=48, authorNames=null, journalName=null, refType=null, unstructuredReference=Peng Y S, Li H B, Wu P X, et al. D-FINE: redefine regression task in DETRs as fine-grained distribution refinement[Z]. arXiv: 2410.13842, 2024. https://doi.org/10.48550/arXiv.2410.13842., articleTitle=null, refAbstract=null), Reference(id=1280951156331491958, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=48, rfOrder=49, authorNames=null, journalName=null, refType=null, unstructuredReference=Huang S H, Lu Z C, Cun X, et al. DEIM: DETR with improved matching for fast convergence[C]//2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025: 15162–15171. https://doi.org/10.1109/CVPR52734.2025.01412., articleTitle=null, refAbstract=null), Reference(id=1280951156406989431, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, doi=null, pmid=null, pmcid=null, year=null, volume=null, issue=null, pageStart=null, pageEnd=null, url=null, language=null, rfNumber=49, rfOrder=50, authorNames=null, journalName=null, refType=null, unstructuredReference=Zhang M J, Zhang R, Yang Y X, et al. ISNet: shape matters for infrared small target detection[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022: 867–876. https://doi.org/10.1109/CVPR52688.2022.00095., articleTitle=null, refAbstract=null)], funds=null, companyList=[AuthorCompany(id=1280951146026086929, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=1, ext=[AuthorCompanyExt(id=1280951146034475538, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China), AuthorCompanyExt(id=1280951146042864147, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018)]), AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])], figs=[ArticleFig(id=1280951147888357937, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.1, caption=The structures of AFITDYOLO, figureFileSmall=R0wMwXZRMMiyj8spqG4mFw==, figureFileBig=FdVY5+DlNv54MYvVHA1j/A==, tableContent=null), ArticleFig(id=1280951147955466802, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图1, caption=AFITDYOLO结构, figureFileSmall=R0wMwXZRMMiyj8spqG4mFw==, figureFileBig=FdVY5+DlNv54MYvVHA1j/A==, tableContent=null), ArticleFig(id=1280951148051935795, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.2, caption=The structures of MFFM, figureFileSmall=jfV1DPbZkDY+3vXrRWU08A==, figureFileBig=qaRY8oQ6XQQHtc6BRmaIkg==, tableContent=null), ArticleFig(id=1280951148135821876, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图2, caption=MFFM架构, figureFileSmall=jfV1DPbZkDY+3vXrRWU08A==, figureFileBig=qaRY8oQ6XQQHtc6BRmaIkg==, tableContent=null), ArticleFig(id=1280951148198736437, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.3, caption=The structures of MFEConv, figureFileSmall=VkeCsYWzpmIQNuvGwMoZSQ==, figureFileBig=Qh9hP5o+66oDDvsHAcTixg==, tableContent=null), ArticleFig(id=1280951148270039606, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图3, caption=MFEConv架构, figureFileSmall=VkeCsYWzpmIQNuvGwMoZSQ==, figureFileBig=Qh9hP5o+66oDDvsHAcTixg==, tableContent=null), ArticleFig(id=1280951148332954167, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.4, caption=The structures of CAFM, figureFileSmall=9ERpNT6LZu64XLVN9+SSzw==, figureFileBig=ddRmrSFNGSZ9TikaTx91GQ==, tableContent=null), ArticleFig(id=1280951148404257336, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图4, caption=CAFM架构, figureFileSmall=9ERpNT6LZu64XLVN9+SSzw==, figureFileBig=ddRmrSFNGSZ9TikaTx91GQ==, tableContent=null), ArticleFig(id=1280951148492337721, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.5, caption=Infrared data sample illustration, figureFileSmall=c91NqXMD3eYTvGyasFJPUA==, figureFileBig=zv2H493Tvs5H7ORDkJJEog==, tableContent=null), ArticleFig(id=1280951148563640890, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图5, caption=数据集, figureFileSmall=c91NqXMD3eYTvGyasFJPUA==, figureFileBig=zv2H493Tvs5H7ORDkJJEog==, tableContent=null), ArticleFig(id=1280951148643332667, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Fig.6, caption=Compare visualization results, figureFileSmall=RtU3BlAEXGX3gZdTQNQV6Q==, figureFileBig=R62/5xg/Xlz1NH0XFIK25w==, tableContent=null), ArticleFig(id=1280951148706247228, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=图6, caption=对比可视化, figureFileSmall=RtU3BlAEXGX3gZdTQNQV6Q==, figureFileBig=R62/5xg/Xlz1NH0XFIK25w==, tableContent=null), ArticleFig(id=1280951148769161789, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Tab.1, caption=

The ablation results for each module

, figureFileSmall=null, figureFileBig=null, tableContent=
Module+Params/M
FPS/FLOPs
CTIR/%SIRST-UAVB/%
MFFMMFEConvCAFMPRmAP50mAP95PRmAP50mAP95
2.71119/8.479.275.284.259.291.274.683.355.1
2.69122/8.483.877.988.363.391.875.984.756.5
2.32127/7.582.577.487.762.792.576.785.557.2
2.79119/8.882.678.887.361.393.376.986.957.8
2.32124/8.084.379.288.763.893.277.887.658.1
2.56122/8.284.780.490.764.293.777.888.958.8
), ArticleFig(id=1280951150438494782, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=表1, caption=

各个模块消融结果

, figureFileSmall=null, figureFileBig=null, tableContent=
Module+Params/M
FPS/FLOPs
CTIR/%SIRST-UAVB/%
MFFMMFEConvCAFMPRmAP50mAP95PRmAP50mAP95
2.71119/8.479.275.284.259.291.274.683.355.1
2.69122/8.483.877.988.363.391.875.984.756.5
2.32127/7.582.577.487.762.792.576.785.557.2
2.79119/8.882.678.887.361.393.376.986.957.8
2.32124/8.084.379.288.763.893.277.887.658.1
2.56122/8.284.780.490.764.293.777.888.958.8
), ArticleFig(id=1280951150572712511, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Tab.2, caption=

Comparison of MFEConv with other convolutions

, figureFileSmall=null, figureFileBig=null, tableContent=
ModelParams/MFPS/FLOPsCTIR/%SIRST-UAVB/%
PRmAP50mAP95PRmAP50mAP95
Conv2.71119/8.479.275.284.259.291.274.683.355.1
AKConv2.69126/7.880.274.783.658.891.174.381.454.9
DSConv2.32133/7.477.975.483.959.089.774.882.654.7
LSKConv2.72117/8.479.776.785.260.290.375.683.955.6
PConv2.32129/7.682.977.986.561.392.676.885.757.6
MFEConv(3)k=52.38121/7.781.277.586.962.192.376.784.656.3
MFEConv(3)k=32.32127/7.581.277.386.361.692.276.784.356.1
MFEConv(6)k=72.5112/8.683.277.787.762.992.877.386.157.6
MFEConv(6)k=52.41118/8.082.977.387.863.192.676.685.657.2
MFEConv(6)k=32.32127/7.583.177.487.762.792.576.785.557.2
), ArticleFig(id=1280951150648209984, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=表2, caption=

MFEConv与其他卷积的比较

, figureFileSmall=null, figureFileBig=null, tableContent=
ModelParams/MFPS/FLOPsCTIR/%SIRST-UAVB/%
PRmAP50mAP95PRmAP50mAP95
Conv2.71119/8.479.275.284.259.291.274.683.355.1
AKConv2.69126/7.880.274.783.658.891.174.381.454.9
DSConv2.32133/7.477.975.483.959.089.774.882.654.7
LSKConv2.72117/8.479.776.785.260.290.375.683.955.6
PConv2.32129/7.682.977.986.561.392.676.885.757.6
MFEConv(3)k=52.38121/7.781.277.586.962.192.376.784.656.3
MFEConv(3)k=32.32127/7.581.277.386.361.692.276.784.356.1
MFEConv(6)k=72.5112/8.683.277.787.762.992.877.386.157.6
MFEConv(6)k=52.41118/8.082.977.387.863.192.676.685.657.2
MFEConv(6)k=32.32127/7.583.177.487.762.792.576.785.557.2
), ArticleFig(id=1280951150748873281, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Tab.3, caption=

Comparison results with mainstream models

, figureFileSmall=null, figureFileBig=null, tableContent=
ModelParams/MFPS/FLOPsCTIR/%SIRST-UAVB/%
PRmAP50mAP95PRmAP50mAP95
FR-CNN41.3856/18.380.376.283.958.887.274.883.555.1
YOLOv10s7.2195/24.482.776.585.159.291.375.283.755.3
YOLOv10m15.4383/6482.97785.359.491.175.184.155.6
YOLOv10n2.71119/8.479.275.284.259.291.274.683.355.1
YOLOv11n2.63133/6.682.978.987.762.79176.984.957.4
YOLOv12n2.61141/6.077.375.581.358.286.371.580.454.1
D-FINE-N3.9776/7.080.277.886.662.191.778.186.457.5
DEIM-N5.3581/7.0382.578.286.562.391.378.386.357.2
AFITDYOLO2.56122/8.284.780.490.764.293.777.788.958.8
), ArticleFig(id=1280951150832759362, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=表3, caption=

与主流模型的比较

, figureFileSmall=null, figureFileBig=null, tableContent=
ModelParams/MFPS/FLOPsCTIR/%SIRST-UAVB/%
PRmAP50mAP95PRmAP50mAP95
FR-CNN41.3856/18.380.376.283.958.887.274.883.555.1
YOLOv10s7.2195/24.482.776.585.159.291.375.283.755.3
YOLOv10m15.4383/6482.97785.359.491.175.184.155.6
YOLOv10n2.71119/8.479.275.284.259.291.274.683.355.1
YOLOv11n2.63133/6.682.978.987.762.79176.984.957.4
YOLOv12n2.61141/6.077.375.581.358.286.371.580.454.1
D-FINE-N3.9776/7.080.277.886.662.191.778.186.457.5
DEIM-N5.3581/7.0382.578.286.562.391.378.386.357.2
AFITDYOLO2.56122/8.284.780.490.764.293.777.788.958.8
), ArticleFig(id=1280951150925034051, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=EN, label=Tab.4, caption=

Generalization experiment results/%

, figureFileSmall=null, figureFileBig=null, tableContent=
DatasetModelPRmAP50mAP95
HIT-UAVYOLOv10n91.169.277.849.2
YOLOv12n90.368.576.548.9
DEIM-N92.173.282.551.3
AFITDYOLO91.774.783.752.3
IRSTD-1kYOLOv10n89.083.787.456.3
YOLOv12n89.282.986.855.7
DEIM-N90.383.389.858.3
AFITDYOLO90.783.990.958.8
), ArticleFig(id=1280951150996337220, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, language=CN, label=表4, caption=

泛化性实验结果/%

, figureFileSmall=null, figureFileBig=null, tableContent=
DatasetModelPRmAP50mAP95
HIT-UAVYOLOv10n91.169.277.849.2
YOLOv12n90.368.576.548.9
DEIM-N92.173.282.551.3
AFITDYOLO91.774.783.752.3
IRSTD-1kYOLOv10n89.083.787.456.3
YOLOv12n89.282.986.855.7
DEIM-N90.383.389.858.3
AFITDYOLO90.783.990.958.8
)], attaches=null, journal=Journal(id=1278641367198941188, delFlag=0, nameCn=光电工程, nameEn=Opto-Electronic Engineering, nameHistory1=null, nameHistory2=null, issn=1003-501X, eissn=2097-4019, cn=51-1346/O4, coden=null, periodic=0, language=CN, oaType=null, ccby=null, superviseOffice=null, ownerOffice=null, pubOffice=null, editorOffice=null, officeType=null, aims=null, clcCode=null, officeProv=null, officeCity=null, officeAddr=null, officeZip=null, officeEmail=null, officePhone=null, editDirector=null, officeDirector=null, officeDirectorPhone=null, officeStaffNum=null, officeEmpNum=null, coverPicUrl=4Vimkd+qXLWNxtdpr9mFNw==, journalPrice=null, startedYear=null, abbrevIsoEn=Opto-Electronic Engineering, journalRemark=null, publicationField=null, createdTime=1782781457950, updatedTime=1784021784694, createdBy=18614031015, updatedBy=13041195026, firstLetterCn=G, firstLetterEn=G, subjectCode=Engineering, subjectName=null, subjectCodeEn=Engineering, subjectNameEn=null, picCn=4Vimkd+qXLWNxtdpr9mFNw==, picEn=vAq9s20WLs1ODDfbWq+Gjg==, jcr=null, cjcr=null, exts=[JournalExt(id=1283843675721536154, language=CN, name=光电工程, nameHistory1=null, nameHistory2=null, managedBy=, sponsoredBy=, publishedBy=, editorOffice=, officeProv=null, officeCity=null, officeAddr=, officeZip=, editDirector=, officeDirector=null, officePhone=null, coverPicUrl=null, journalRemark=, submitArticleUrl=null, websiteUrl=, createdTime=1784021784954, updatedTime=1784021784954, createdBy=13041195026, updatedBy=13041195026, submissionGuidelinesUrl=, submissionAuthorUrl=http://www.manuscripts.com.cn/gdgc, submissionEditorUrl=http://www.manuscripts.com.cn/gdgc, submissionReviewUrl=http://www.manuscripts.com.cn/gdgc, submissionCeEditorUrl=, submissionAeEditorUrl=, option={"copyright":""}), JournalExt(id=1283843675771867803, language=EN, name=Opto-Electronic Engineering, nameHistory1=null, nameHistory2=null, managedBy=, sponsoredBy=, publishedBy=, editorOffice=, officeProv=null, officeCity=null, officeAddr=, officeZip=, editDirector=, officeDirector=null, officePhone=null, coverPicUrl=null, journalRemark=, submitArticleUrl=null, websiteUrl=, createdTime=1784021784966, updatedTime=1784021784966, createdBy=13041195026, updatedBy=13041195026, submissionGuidelinesUrl=, submissionAuthorUrl=http://www.manuscripts.com.cn/gdgc, submissionEditorUrl=http://www.manuscripts.com.cn/gdgc, submissionReviewUrl=http://www.manuscripts.com.cn/gdgc, submissionCeEditorUrl=, submissionAeEditorUrl=, option={"copyright":""})], databaseList=null, tenantJournalId=1278651732997652489, websiteList=[Website(id=1278723867418018151, webName=null, webTitle=null, webDomain=null, webCopyrigh=null, webIpcNo=null, seoTitle=null, seoKeywords=null, seoDescription=null, tenantJournalId=null, journalId=1278651732997652489, journalNameCn=null, journalNameEn=null, grayFlag=null, tenantId=1146029695717560320, platformId=null, journalGroupId=null, journalGroupNameCn=null, journalGroupNameEn=null, type=1, domain=https://castjournals.cast.org.cn/joweb/oee/CN, language=CN, createTime=1782801127533, createBy=18614031015, updateTime=1782804494442, updateBy=18614031015, name=光电工程-中文, tplId=1146099689490845704, title=光电工程, delFlag=0, indexPage=/home, props=[WebsiteProps(id=1278738091150128034, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=articleTextType, value=kx, createTime=1782804518735, updateTime=1782804518735, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091120767903, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=banner, value=null, createTime=1782804518728, updateTime=1782804518728, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091171099557, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=grayFlag, value=0, createTime=1782804518740, updateTime=1782804518740, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091108184990, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=logo, value=https://castjournals.cast.org.cn/joweb/oee/CN/file/pic?fileId=A1C6uwqtMazluiWkEpR0Mg==, createTime=1782804518725, updateTime=1782804518725, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091179488167, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=minRunFlag, value=0, createTime=1782804518742, updateTime=1782804518742, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091141739425, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=picServerUrl, value=https://castjournals.cast.org.cn/joweb/oee/CN/file/pic, createTime=1782804518733, updateTime=1782804518733, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091175293862, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=silenceFlag, value=0, createTime=1782804518741, updateTime=1782804518741, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091129156512, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=staticResourcePath, value=https://castjournals.cast.org.cn/joweb/cast_kjdb_cn_619/, createTime=1782804518730, updateTime=1782804518730, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091154322339, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=themeColor, value=null, createTime=1782804518736, updateTime=1782804518736, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738091162710948, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867418018151, code=themeStyle, value=null, createTime=1782804518738, updateTime=1782804518738, creator=18614031015, updator=18614031015)]), Website(id=1278723867522875769, webName=null, webTitle=null, webDomain=null, webCopyrigh=null, webIpcNo=null, seoTitle=null, seoKeywords=null, seoDescription=null, tenantJournalId=null, journalId=1278651732997652489, journalNameCn=null, journalNameEn=null, grayFlag=null, tenantId=1146029695717560320, platformId=null, journalGroupId=null, journalGroupNameCn=null, journalGroupNameEn=null, type=1, domain=https://castjournals.cast.org.cn/joweb/oee/EN, language=EN, createTime=1782801127558, createBy=18614031015, updateTime=1782804490442, updateBy=18614031015, name=光电工程-英文, tplId=1146101810881728533, title=Opto-Electronic Engineering, delFlag=0, indexPage=/home, props=[WebsiteProps(id=1278738063660659607, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=articleTextType, value=kx, createTime=1782804512181, updateTime=1782804512181, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063635493780, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=banner, value=null, createTime=1782804512175, updateTime=1782804512175, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063924900762, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=grayFlag, value=0, createTime=1782804512244, updateTime=1782804512244, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063606133651, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=logo, value=https://castjournals.cast.org.cn/joweb/oee/EN/file/pic?fileId=A1C6uwqtMazluiWkEpR0Mg==, createTime=1782804512168, updateTime=1782804512168, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063937483676, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=minRunFlag, value=0, createTime=1782804512247, updateTime=1782804512247, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063652270998, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=picServerUrl, value=https://castjournals.cast.org.cn/joweb/oee/EN/file/pic, createTime=1782804512179, updateTime=1782804512179, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063933289371, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=silenceFlag, value=0, createTime=1782804512246, updateTime=1782804512246, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063639688085, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=staticResourcePath, value=https://castjournals.cast.org.cn/joweb/cast_kjdb_en_623/, createTime=1782804512176, updateTime=1782804512176, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063664853912, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=themeColor, value=null, createTime=1782804512182, updateTime=1782804512182, creator=18614031015, updator=18614031015), WebsiteProps(id=1278738063916512153, tenantId=1146029695717560320, journalId=null, journalGroupId=null, siteId=1278723867522875769, code=themeStyle, value=null, createTime=1782804512242, updateTime=1782804512242, creator=18614031015, updator=18614031015)])], journalTitle=光电工程, weixinUrl=null, journalUrl=https://www.oejournal.org/oee, iacademicId=null, status=1, seqNo=null, journalTitleEn=Opto-Electronic Engineering, journalPhotoCn=4Vimkd+qXLWNxtdpr9mFNw==, journalPhotoEn=vAq9s20WLs1ODDfbWq+Gjg==, journalFirstLetter=G, journalRecommend=null, journalNew=null, journalCollection=null, jcrJf=null, cjcrJf=null, jcrJfStr=null, cjcrJfStr=null, submissionFirstDecision=null, sciSubjectClassification=null, casSubjectClassification=null, citeScore=null, totalCitationFrequency=null, icpCode=null, psCode=null, advertisingLicenseCode=null, copyrightInformation=null, country=null, option=, provinceCode=null, provinceName=null, collectFlag=false, interPubPlatform=, interPubPlatformUrl=null), detailUrlCn=https://castjournals.cast.org.cn/joweb/oee/CN/10.12086/oee.2026.250292, detailUrlEn=https://castjournals.cast.org.cn/joweb/oee/EN/10.12086/oee.2026.250292, pdfUrlCn=https://castjournals.cast.org.cn/joweb/oee/CN/PDF/10.12086/oee.2026.250292, pdfUrlEn=https://castjournals.cast.org.cn/joweb/oee/EN/PDF/10.12086/oee.2026.250292, aliStartDate=0, aliEndDate=0, collectionFlag=false, citedCount=null, citedUrl=null, previewStatus=0, delFlag=0, hasFullText=1, orderTime=1776960000000, fullTextJson=null, articleText=null, reference=null)
收藏切换
基于YOLO的自适应多尺度红外目标检测网络
收藏切换
PDF下载
汪佳旭 1, 2 , 杨俊 2, * , 许聪源 2
光电工程 | 科研论文 2026,53(4): 250292
收起
收藏切换
光电工程 |科研论文 2026 , 53 (4) : 250292
基于YOLO的自适应多尺度红外目标检测网络
全屏
[Author(id=1280951146235802136, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=0, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=2936598421@qq.com, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1280951146323882523, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146235802136, language=EN, stringName=Jiaxu Wang, firstName=Jiaxu, middleName=null, lastName=Wang, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=1, 2, address=1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China
2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146399379996, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146235802136, language=CN, stringName=汪佳旭, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=1, 2, address=1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018
2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"RDKWMmEB7nCI4VxjQJrwEQ==","content":"

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

"}, bioImg=RDKWMmEB7nCI4VxjQJrwEQ==, bioContent=

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146026086929, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=1, ext=[AuthorCompanyExt(id=1280951146034475538, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China), AuthorCompanyExt(id=1280951146042864147, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146026086929, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018)]), AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])]), Author(id=1280951146466488862, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=1, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=juneryoung@zjxu.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=1, authorType=1, ext={EN=AuthorExt(id=1280951146541986336, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146466488862, language=EN, stringName=Jun Yang, firstName=Jun, middleName=null, lastName=Yang, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, *, address=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146609095201, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146466488862, language=CN, stringName=杨俊, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, *, address=2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"uHEHZcx2EFUFfNFDYfM2/w==","content":"

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

"}, bioImg=uHEHZcx2EFUFfNFDYfM2/w==, bioContent=

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])]), Author(id=1280951146676204067, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, orderNo=2, firstName=null, middleName=null, lastName=null, nameCn=null, orcid=null, stid=null, country=null, authorPic=null, dead=0, email=cyxu@zjxu.edu.cn, emailSecond=null, emailThird=null, correspondingAuthor=0, authorType=1, ext={EN=AuthorExt(id=1280951146890113573, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146676204067, language=EN, stringName=Congyuan Xu, firstName=Congyuan, middleName=null, lastName=Xu, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, address=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China, bio=null, bioImg=null, bioContent=null, aboutCorrespAuthor=null), CN=AuthorExt(id=1280951146953028134, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, authorId=1280951146676204067, language=CN, stringName=许聪源, firstName=null, middleName=null, lastName=null, prefix=null, suffix=null, authorComment=null, nameInitials=null, affiliation=null, department=null, xref=2, address=2嘉兴大学人工智能学院,浙江 嘉兴 314001, bio={"img":"jDy0t5HTKKb3mneNkFo5fQ==","content":"

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

"}, bioImg=jDy0t5HTKKb3mneNkFo5fQ==, bioContent=

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

, aboutCorrespAuthor=null)}, companyList=[AuthorCompany(id=1280951146139333140, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, xref=2, ext=[AuthorCompanyExt(id=1280951146147721749, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=EN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China), AuthorCompanyExt(id=1280951146156110358, tenantId=1146029695717560320, journalId=1278651732997652489, articleId=1279511861326484060, companyId=1280951146139333140, language=CN, country=null, province=null, city=null, postcode=null, companyName=null, departmentName=null, remark=2嘉兴大学人工智能学院,浙江 嘉兴 314001)])])]
汪佳旭1, 2 , 杨俊2, * , 许聪源2
作者信息
  • 1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018
  • 2嘉兴大学人工智能学院,浙江 嘉兴 314001
通讯作者:
作者简介:

汪佳旭(2001-),男,浙江理工大学信息科学与工程学院(网络安全学院)硕士研究生,研究方向为机器视觉、深度学习理论及应用。E-mail:

杨俊(1978-),男,工学博士,硕士研究生导师,嘉兴大学人工智能学院副教授,研究方向为智能多媒体信息分析与处理,机器学习,计算机视觉。E-mail:

许聪源(1990-),男,工学博士,嘉兴大学人工智能学院讲师,研究方向为机器视觉、网络空间安全。E-mail:

YOLO-based adaptive multi-scale infrared target detection network
Jiaxu Wang1, 2 , Jun Yang2, * , Congyuan Xu2
Affiliations
  • 1School of Information Science and Engineering(School of Cyber Science and Technology), Zhejiang Sci-Tech University, Hangzhou, Zhejiang 310018, China
  • 2College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China
出版时间: 2026-04-24 doi: 10.12086/oee.2026.250292
文章导航
收藏切换

现有基于深度学习的红外目标检测算法(infrared target detection algorithm, ITDA)对红外图像中特征的提取和处理能力不足,导致漏检、误检率较高。多数算法不能兼顾不同尺度的红外目标,红外小目标检测依赖高灵敏度特征提取以捕捉细微特征,当检测目标尺度增大时,因需要更强的全局理解能力,易出现局部纹理过拟合,导致性能降低,使算法在跨场景部署时准确性降低。为解决上述问题,本文提出了一种基于YOLO的自适应多尺度红外目标检测的网络(adaptive multi-scale infrared target detection network based on YOLO, AFITDYOLO),该网络通过多层特征提取模块及多层特征融合模块增强多尺度红外目标检测能力。提出多尺度特征融合模块MFFM,通过增强特征金字塔中不同层特征的相关性,提升了多尺度特征融合的表达能力;提出多核特征提取卷积MFEConv,通过异构卷积组增大感受野,更好地与不同尺度目标的空间分布保持一致;提出交叉注意力融合模块CAFM,通过网络中不同层输出特征图的对比交互,增强重要特征信息,提升特征表达能力。为了评估AFITDYOLO的性能,在无人机飞鸟数据集SIRST-UAVB和道路行人车辆数据集CTIR上进行实验,mAP50分别达到了88.9%和90.7%,对比YOLOv10n分别提高了5.6%和6.5%,与目前主流方法对比,本文方法检测精度最佳。本文提出的红外目标检测算法,在多尺度红外目标检测中,展现了优异的准确性和适应性,可跨场景部署。

红外目标检测  /  特征融合  /  轻量级网络结构  /  注意力机制  /  YOLO
Objective

Infrared target detection plays an indispensable role in numerous critical domains, such as security surveillance, autonomous driving, and military reconnaissance, owing to its unique perceptual capability under complex environments (e.g., low-light conditions and severe weather). However, infrared images inherently suffer from low contrast, blurred details, and significant noise interference, which often lead to ambiguous target edges, missing texture features, and other challenges during the detection process. Existing deep learning-based infrared target detection algorithms (ITDA) exhibit inadequate performance in feature extraction and processing for infrared images, resulting in relatively high rates of missed detection and false detection. Moreover, our systematic analysis of infrared target detection tasks reveals that algorithms tailored for small infrared targets rely heavily on high-sensitivity feature extraction to capture subtle characteristics. Nevertheless, as the scale of detected targets increases, these algorithms tend to encounter overfitting to local textures and elevated false detection rates, thereby degrading overall performance. In practical applications, detection environments are dynamically changing with targets of varying scales; thus, multi-scale detection capability is critical to ensuring algorithms maintain high reliability and adaptability in complex real-world scenarios. Unfortunately, most state-of-the-art algorithms are optimized for single-scale targets, making it challenging to simultaneously satisfy the requirements of high-precision localization for small targets and effective semantic understanding for large targets.

Methods

To address the above issues, this paper proposes an adaptive multi-scale infrared target detection network based on YOLO (AFITDYOLO). This network is designed to receive infrared target images of different scales and employs a multi-layer feature extraction module and a multi-layer feature fusion module to enhance its multi-scale infrared target detection capability. Firstly, a multi-scale feature fusion module (MFFM) is proposed. This module enhances the correlation between features of different layers in the feature pyramid network (FPN), coordinates deep semantic features with shallow spatial detail features more effectively, and thereby improves the representational ability of multi-scale feature fusion. Secondly, a multi-kernel feature extraction convolution (MFEConv) is constructed. By utilizing heterogeneous convolution groups, MFEConv expands the receptive field and strengthens the model's feature extraction capability. Additionally, a cross-attention fusion module (CAFM) is designed. Through the comparative interaction of feature maps output by different layers in the detection network, CAFM leverages the complementary information among these feature maps to suppress infrared noise in images and further enhance feature representation capability.

Results and Discussions

To validate the effectiveness of the proposed method in improving detection performance, extensive training and evaluation are conducted on the CTIR dataset, which comprises road pedestrians and vehicles with multi-scale infrared targets. To further verify the adaptability of the method, additional experiments are performed on the SIRST-UAVB dataset—a single-frame UAV bird dataset characterized by more complex backgrounds and smaller target scales. Experimental results on these two datasets demonstrate that AFITDYOLO achieves mean average precision at 50% intersection over union (mAP50) of 88.9% and 90.7%, respectively, representing significant improvements of 5.6% and 6.5% compared with YOLOv10n. In terms of lightweight optimization, the proposed method achieves higher inference speed (measured in frames per second, FPS) while utilizing fewer model parameters (params) and floating-point operations (FLOPs). When compared with current mainstream methods, AFITDYOLO exhibits the highest detection accuracy, the lowest parameter count and FLOPs, and the fastest inference speed, demonstrating distinct advantages. Additionally, to evaluate the generalization ability of the proposed method, cross-dataset experiments are carried out on the HIT-UAV dataset (a high-altitude UAV infrared thermal imaging dataset) and the IRSTD-1k dataset (a classic infrared small target dataset). Experimental results indicate that while the precision (P) value of AFITDYOLO is slightly inferior to that of DEIM-N, it outperforms all other mainstream methods in all remaining evaluation metrics. These findings confirm that the proposed method achieves improved detection accuracy on the infrared datasets used in the generalization experiments, validating its strong generalization capability and further demonstrating its feasibility for cross-scenario deployment. Overall, the proposed method simultaneously achieves enhanced detection accuracy and lightweight optimization of the detection model, fully meeting the requirements of real-time detection applications.

Conclusions

The AFITDYOLO network proposed in this paper, which is an adaptive multi-scale infrared target detection network based on YOLO, enhances the detection accuracy of infrared targets of different scales under various backgrounds with a relatively small number of parameters. The proposed MFFM enhances the model's representational ability in multi-scale feature fusion by improving the correlation between features of different layers in FPN. Additionally, the lightweight convolution module MFEConv is designed to achieve an efficient and larger receptive field with minimal parameters by leveraging the target distribution characteristics of infrared images. Furthermore, the CAFM is introduced to highlight important feature information, filter out irrelevant background information, and suppress noise through the comparative interaction of feature maps output by different layers, thereby further boosting the model's feature representation capability. Experimental results demonstrate that the proposed method outperforms current mainstream algorithms, exhibiting excellent detection accuracy, lightweight performance, and generalization ability, along with cross-scenario deployment capabilities.

infrared target detection  /  feature fusion  /  lightweight network architecture  /  attention mechanism  /  YOLO
汪佳旭, 杨俊, 许聪源. 基于YOLO的自适应多尺度红外目标检测网络. 光电工程, 2026 , 53 (4) : 250292 - . DOI: 10.12086/oee.2026.250292
Jiaxu Wang, Jun Yang, Congyuan Xu. YOLO-based adaptive multi-scale infrared target detection network[J]. Opto-Electronic Engineering, 2026 , 53 (4) : 250292 - . DOI: 10.12086/oee.2026.250292
红外目标检测凭借其在低光照、恶劣天气等复杂环境下独特的感知能力,在安防监控、自动驾驶、军事侦察等众多关键领域中发挥着举足轻重的作用。红外图像具有对比度低、细节模糊、噪声干扰明显等特性[1],导致其检测时易出现目标边缘不清晰、纹理特征缺失等情况。
早期红外目标检测采用传统模型驱动方法[2-5],需要基于先验知识并手动调整参数,难以有效提取特征,检测精度和实时性均面临瓶颈。随着深度卷积神经网络(convolutional neural network, CNN)在视觉任务中的成功应用,基于深度学习的红外目标检测算法[6-8]也由此兴起,此类算法具有高准确性和轻量化的优势,有助于提高红外目标检测精度,从而提高运营效率和成本效益。现有基于深度学习的红外目标检测算法通过新颖的网络架构提升了红外检测的性能,如Yang等人根据红外小目标的成像特点提出了风车形卷积PConv[9]并引入检测网络,增强了底层特征提取,提升了红外小目标检测精度;Zhao等人提出了一种基于深度注意力机制的红外行人检测方法[10],通过构建四尺度特征金字塔网络并结合注意力模块,显著提高了近距离行人目标的检测性能。我们在分析红外目标检测任务中得出,红外小目标检测算法依赖高灵敏度特征提取以捕捉细微特征,但在检测目标尺度增大时,易出现局部纹理过拟合和误检,导致性能降低。在实际应用中,检测环境复杂多变,目标尺度不一,多尺度检测能力是确保算法在复杂真实环境中保持高可靠性、适应性的关键。但目前多数算法均针对单一尺度目标进行改进,这些算法难以同时满足小目标的高精度定位与大目标的语义理解需求,本文针对该问题提出一种基于YOLO的自适应多尺度红外目标检测网络AFITDYOLO,目的是同时提升检测网络的适用性和性能,该网络接收不同尺度红外目标图像并使用多层特征提取模块及多层特征融合模块增强多尺度红外目标检测能力。本文的主要贡献如下:
1)提出一种基于YOLO的自适应多尺度红外目标检测网络AFITDYOLO,通过多层特征提取及融合增强多尺度红外目标检测能力。
2)提出多尺度特征融合模块(multi scale feature fusion module, MFFM),通过增强特征金字塔中不同层特征之间的相关性,提升多尺度特征融合的表达能力。
3)构建多核特征提取卷积 (multi-kernel feature extraction convolution, MFEConv),通过异构卷积组,增大感受野,增强模型特征提取能力,
4)设计交叉注意力融合模块(cross-attention fusion module, CAFM),通过检测网络中不同层输出特征图的对比交互,利用特征图之间的互补信息,抑制图像红外噪声,进一步增强特征表达能力。
传统目标检测方法依赖手工特征提取与模式分类这一核心流程。这些方法通常包含预处理、分割、特征提取和分类等步骤,核心思想是利用图像的特性,如目标与背景的差异、目标形状等进行检测。Dalal等人基于手工特征提出了HOG+SVM[11],该方法使用面向梯度HOG描述符的直方图,并将检测窗口与重叠的HOG描述符组合,最后使用基于SVM的窗口分类器对组合的特征向量进行分类,实现目标行人检测。He等人提出了一种基于SVM和红外热成像技术的高压绝缘子污染等级检测新方法[12],使用基于梯度信息的自适应平滑滤波去除图像噪声,分割并提取目标区域,从绝缘子表面提取温度、最大与最小温度对比度、表面温度标准差以及前10%亮度像素与总像素比例四类特征,采用多类SVM进行污染等级检测。Felzenszwalb等人在HOG基础上提出基于部件的模型(discriminatively trained part based models, DPM)[13],通过将目标对象分解为多个部件,然后对每个部件进行独立的训练,最后这些部件组合形成完整的目标检测模型。
传统的目标检测方法,提高了目标的检测精度,但泛化能力差,很难部署至红外目标检测,同时多步骤串联导致实时性受限。
早期由Sermanet等人提出的Overfeat算法[14]首次使用CNN代替手工特征来提取特征,显著提升了特征表达能力。近年来,许多基于深度学习的方法从不同角度分析和改进了红外目标检测。
特征金字塔(Feature pyramid networks, FPN)的构建是红外目标检测任务中的关键步骤,并且是现代检测器的组成部分,形成了解决目标多尺度问题的基础。对于较小的红外目标,特征图通常只包含来自几个甚至一个像素的有效信息;对于较大的红外目标,特征图则是由多个组成部分和多层次特征构成的复杂实体。因此,特征融合方法的研究对于准确表示红外目标的特征信息显得尤为重要。FPN构建了自上而下的路径,结合了各个级别的特征以实现多尺度特征融合。PANet[15]在FPN的基础上引入了自下而上的路径,有利于高分辨率信息与更强语义特征的融合。随后,NAS-FPN[16]和BiFPN[17]被提出来增强多尺度特征的融合。与许多特征融合方法不同,AFPN[18]探索聚合特征上的节点操作,利用注意力机制来指导特征融合。虽然这种方法增强了检测性能,但它显着增加了计算复杂度。在本文中,我们研究多尺度特征融合,提出多尺度特征融合模块MFFM,以低于FPN的计算复杂度实现了模型优越的检测性能。
在基于CNN的红外目标检测中,Dai等人提出了OSCAR单级联级优化网络[19],Yang等人引入了EFLNet[7]以解决目标与背景不平衡问题。ACM模块[20]将底层细节嵌入高级特征中,而DNANet[21]则通过融合机制挖掘上下文信息。这些基于CNN的网络专注于构建特征提取网络,但忽视了通过卷积模块增强ITDA性能的潜力。ISNet模型[22],它结合了可变形卷积,提高了检测性能,但增加了训练时间和网络参数。因此,本文通过构建轻量化卷积MFEConv,增强模型提取底层特征的能力。
红外图像包含丰富的多层次信息,不同信息对检测任务的贡献程度各异。为了更聚焦于图像中的重要特征与结构信息,Woo等人提出了CBAM[23],实现红外目标准确提取。此后,各种注意力机制[24-27]被证明是提升红外目标检测性能的有效手段。但这些方法未充分挖掘图像的互补信息,本文设计交叉注意力融合模块CAFM,通过检测网络中不同层输出特征图之间的对比交互,突出重要特征信息,抑制图像红外噪声,实现模型特征表达能力的增强。
本文提出一种基于YOLO的自适应多尺度红外目标检测网络 AFITDYOLO,由主干(backbone)、颈部(neck)和检测头(head)三个部分组成,整体结构如图1所示。
AFITDYOLO模型结合了C2MFE网络架构,实现高效的特征提取,该架构使用MFEConv模块替换C2f中Bottleneck[28-30]的标准卷积层,用于提升C2f模块的特征提取能力,同时实现更高效、更灵活的卷积运算。为了优化网络结构,增强模型多尺度特征融合环节,提出MFFM模块,用于合并不同层次的特征图,该模块通过改进特征金字塔中不同层次特征的相关性,提升了多尺度特征融合的表达能力。为了进一步提升网络检测性能,设计交叉注意力融合模块CAFM提高检测精确度,在检测头的P3 (高分辨率)、P4 (中尺度)、P5 (低分辨率)层输出前分别部署CAFM,针对性处理不同尺度的红外目标,有效提升了红外目标检测的精度与鲁棒性。
在红外目标检测任务中,通过多尺度特征融合机制[29]整合浅层空间细节与深层语义信息能够有效处理红外图像低分辨率、低对比度、目标尺度极端变化三大问题。但传统多尺度特征融合机制存在缺陷,仅通过堆叠和通道融合对低分辨率特征进行上采样并与相邻层融合,没有考虑它们的相关性,从而限制了各层相关特征的利用,针对该问题,本文设计了多尺度有效融合模块 MFFM,模块结构如图2所示。首先对输入特征进行预处理与全局特征融合,使用$ 1\times 1 $卷积层来获取模块的浅层特征并保持通道数的一致,将浅层特征进行元素级相加生成全局特征$ {\boldsymbol{X}}_{\rm{G}} $,使用三分支并行处理全局特征。
通道注意力机制CBAM通过引入注意力模块学习每个通道的权重以动态调整每个通道的重要性。第一个分支引入通道注意力机制CBAM对$ {\boldsymbol{X}}_{\rm{G}} $进行处理,得到通道注意力特征$ {\boldsymbol{Y}}_{1} $。该分支确保模块能有效聚焦在重要的通道特征上,同时保留原始特征细节,防止丢失关键目标信息。
第二个分支中通过Sigmoid激活函数[31]处理$ {\boldsymbol{X}}_{\rm{G}} $生成筛选系数S,并与全局特征相乘得到筛选后的特征$ {\boldsymbol{X}}_{\rm{S}} $。将$ {\boldsymbol{X}}_{\rm{S}} $沿通道维度均分为4组,得到$ {\boldsymbol{X}}_{s}^{i}, i \in[1,2,3,4] $。每组特征经由卷积模块$ {\mathrm{Conv}}2 $处理,产生一个用于捕捉通道间特征相关性的注意力掩码。该掩码被应用于特征细化过程,随后各组细化特征$ {\boldsymbol{X}}_{{\mathrm{proc}}}^{4} $拼接融合[32](Cat),以形成具有聚合性和高度相关性的相邻特征$ {\boldsymbol{X}}_{{\mathrm{proc}}} $$ ({\boldsymbol{Y}}_{2}) $,计算方法见式(1)。
$ \begin{split}& {\boldsymbol{X}}_{ \rm{proc }}^i={\boldsymbol{X}}_s^i \cdot {S oftmax}\left\{B N\left[{Conv}_{1 \times 1}\left({\boldsymbol{X}}_s^i\right)\right]\right\} \\& {\boldsymbol{Y}}_2={\boldsymbol{X}}_{\rm{proc }}={Cat}\left({\boldsymbol{X}}_{\rm{proc }}^1, \cdots, {\boldsymbol{X}}_{ \rm{proc }}^4\right)\end{split}\;. $
通过分组处理,模型能更灵活地捕捉红外图像的多尺度特征。
第三个分支将来自不同深度层级特征图中包含的丰富信息和较弱信息分离出来,并进行独立处理。具体而言,如图2中所示,输入特征$ {\boldsymbol{X}}_{{i-1}} $$ {\boldsymbol{X}}_{i} $$ 1\times 1 $卷积处理后,进行批量归一化(BN)[33]并由Sigmoid函数激活,分别生成每个特征对应的空间信息权重$ {\omega }_{1} $,$ {\omega }_{2} $,表明不同阶段特征的重要性。对$ {\boldsymbol{X}}_{\rm{G}} $进行自适应平均池化操作AP (average pooling)[34]并通过Sigmoid激活函数得到特征权重阈值$ {\omega }_{3} $,利用阈值函数Threshold[35],将来自不同阶段特征的权重信息$ {\omega }_{1} $,$ {\omega }_{2} $与特征权重阈值$ {\omega }_{3} $进行比较,获得捕捉空间信息强度的注意力图,计算方法如下:
$ \begin{split}& \left(\omega_1^{\text {low }}, \omega_1^{{\mathrm{u p}}}\right)={Threshold}\left(\omega_1, \omega_3\right), \\& \left(\omega_2^{\text {low }}, \omega_2^{{\mathrm{u p}}}\right)={Threshold}\left(\omega_2, \omega_3\right)\end{split}\;. $
通过注意力图将来自网络中不同层的强特征分离并聚合成强特征集$ {\boldsymbol{X}}_{\rm{up}} $,不同层的弱特征分离并聚合成弱特征集$ {\boldsymbol{X}}_{{\mathrm{low}}} $,计算方法见式(3),其中$\otimes $表示逐元素乘法:
$ \begin{split}{\boldsymbol{X}}_{\rm{up}}=&\left(\omega_1^{\rm{u p}}+\omega_2^{\rm{u p}}\right) \otimes {\boldsymbol{X}}_G \;,\\{\boldsymbol{X}}_{\rm{low}}=&\left(\omega_1^{\rm{low}}+\omega_2^{\rm{low}}\right) \otimes {\boldsymbol{X}}_G\end{split}\;. $
然后,将强、弱特征集分别进行独立处理,采用深度可分离卷积(DSConv)[36]和门控生成器GateGen处理弱特征集$ {\boldsymbol{X}}_{{\mathrm{low}}} $,生成具有更丰富语义信息的特征;应用$ 1\times 1 $卷积保持强特征集与处理后的弱特征集的通道一致。聚合得到此分支的最终输出$ {\boldsymbol{Y}}_{3} $。通过具有更丰富语义信息的特征与展示更多详细信息的强特征集进行聚合,使第三分支输出特征既包含详细信息,也包含跨通道交换信息。
将上述三个分支输出进行残差连接,并在输出前引入高效局部注意力(ELA)[37],通过动态地聚焦目标关键区域进而抑制无关噪声。MFFM增强FPN中不同层输出特征的相关性,使深层语义特征与浅层空间细节特征更加协调,提升了多尺度特征融合的表达能力。
基于卷积神经网络的红外目标检测方法多数采用标准卷积,从而忽略了红外目标分布的空间特征[8]。红外场景下往往存在多个待检测目标,且尺度不同、分布不均,标准卷积无法全面有效捕捉。本文针对上述问题设计了多核特征提取卷积MFEConv,模块架构如图3所示。
由通道移位模块(shift modules)[38]对输入特征进行预处理,对输入特征图进行不同方向的移位并将移位后的特征图进行叠加,通过该处理增强输入特征的多样性。在MFEConv中,对预处理后的特征图$ {\boldsymbol{X}}_{2} $按通道维度均分为6组,通过异构卷积组处理,卷积组由水平左右方向$ Con{v}_{1\times k} $、垂直上下方向$ Con{v}_{k\times 1} $及左右对角线方向$ Con{v}_{k\times k} $ 6个卷积核构成。为了提高训练稳定性和速度,本文在每次卷积后应用BN和Sigmoid线性单元(SiLU)[39],并行卷积计算见式(4)。
$ \begin{split}{\boldsymbol{X}}_{\rm{H}}^{1,2}&={\boldsymbol{X}}_{\rm{H}}^{2,2}=SiLU[BN({\boldsymbol{X}}_{\rm{H}}^{1}\ast Conv_{1\times k}^{{C}_{\rm{B}}})],\\{\boldsymbol{X}}_{\rm{V}}^{1,2}&={\boldsymbol{X}}_{\rm{V}}^{2,2}=SiLU[BN({\boldsymbol{X}}_{\rm{V}}^{1}\ast Conv_{k\times 1}^{{C}_{\rm{B}}})],\\{\boldsymbol{X}}_{\rm{S}}^{1,2}&={\boldsymbol{X}}_{\rm{S}}^{2,2}=SiLU[BN({\boldsymbol{X}}_{\rm{S}}^{1}\ast Conv_{k\times k}^{{C}_{\rm{B}}})]\end{split}\;, $
式中:$ \ast $表示卷积算子;通道数$ {C}_{\rm{B}}=C/6 $;k表示卷积核的长度,其取值越大模块的参数量越大,模块中k的值取3;下角标H、V、S分别表示水平方向、垂直方向、对角线方向;$ {\boldsymbol{X}}_{\rm{H}}^{1,2} $,$ {\boldsymbol{X}}_{\rm{H}}^{2,2} $分别表示由水平左右方向卷积处理后的特征,6分支并行卷积输出拼接融合(${ Cat } $):
$ {\boldsymbol{X}}_{3}=Cat({\boldsymbol{X}}_{\rm{H}}^{1,2},{\boldsymbol{X}}_{\rm{V}}^{1,2},{\boldsymbol{X}}_{\rm{S}}^{1,2},{\boldsymbol{X}}_{\rm{H}}^{2,2},{\boldsymbol{X}}_{\rm{V}}^{2,2},{\boldsymbol{X}}_{\rm{S}}^{2,2}) \;. $
通过异构卷积组增大感受野,捕获特征图中多方向分布特征,同时有效应对红外目标尺度变化问题。
设计联合注意力机制模块 (fusion attention mechanism, FAM) 引入MFEConv,进一步增强多尺度特征提取能力。具体而言,使用$ 1 \times 1 $卷积处理特征$ {\boldsymbol{X}}_{3} $压缩至单通道,生成FAM输入特征$ {\boldsymbol{X}}_{\rm{FAM}} $。应用$ 3 \times 3$卷积与Sigmoid函数处理$ {\boldsymbol{X}}_{{\mathrm{FAM}}} $生成热辐射权重W。引入Softmax函数[40]处理权重W与输入特征$ {\boldsymbol{X}}_{\rm{FAM}} $获取通道注意力$ \boldsymbol{A}_{\rm{C}} $,计算方法见式(6)。
$ \begin{split}{\boldsymbol{A}}_{\mathrm{C}}= & {S oftmax}\left\{\lambda \cdot {AvgPool}\left({\boldsymbol{X}}_3 \odot {\boldsymbol{W}}\right)\right. \\&+ \left.(1-\lambda) {MaxPool}\left[{\boldsymbol{X}}_3 \odot(1-{\boldsymbol{W}})\right]\right\}\end{split}\;, $
式中:AvgPool与MaxPool分别为热加权平均池化,反向加权最大池化;$ \lambda \in[0,1] $为可学习热融合系数,默认为0.5;$\odot $表示加权乘法。通道注意力聚焦主通道,利于大目标的语义理解。空间注意力$ {{\boldsymbol{A}}}_{\rm{S}} $计算方法见式(7),生成过程中,权重W值大的目标区域在计算$ {\boldsymbol{K}}^{\rm{T}}\boldsymbol{V} $时获得更高权重,权重W值小的目标区域则降低其空间关联性,从而强化高温区域细节,利于小目标的定位。其中$ \sqrt{C} $为温度缩放因子,用于稳定训练,上角标T表示矩阵转置:
$ \begin{split}& {\boldsymbol{A}}_{\mathrm{S}}={S igmoid}\left(\frac{{\boldsymbol{K}}^{\mathrm{T}} {\boldsymbol{V}}}{\sqrt{C}} \odot {\boldsymbol{W}}\right) \;,\\& {\boldsymbol{K}}={\boldsymbol{V}}={Conv}_{1 \times 1}\left({\boldsymbol{X}}_{\rm{FAM}}\right)\;. \end{split}$
MFEConv末端使用残差连接,其中$ \gamma $为可学习缩放参数,初始值为1,输出计算如下:
$ {\boldsymbol{Y}}={\boldsymbol{X}}_{3}+\gamma [({{\boldsymbol{A}}}_{\rm{S}}\otimes {{\boldsymbol{A}}}_{{\mathrm{C}}}\otimes {\boldsymbol{W}})\odot {\boldsymbol{X}}_{2}]\;. $
MFEConv中的异构卷积组,可增强特征图中不同方向分布特征的提取;利用联合注意力机制,可以在抑制噪声的同时,对目标区域进行特征强化。
C2f是YOLO网络中的高效特征提取模块,通过拼接不同Bottleneck模块的输出和原始特征图实现特征提取,Bottleneck的标准卷积层由一个$ 3\times 3 $卷积和$ 1\times 1 $卷积叠加组成。本文提出MFEConv替换Bottleneck中的标准卷积层,构建改进模块C2MFE,在提升效果的同时,实现轻量化改造,提高检测效率,克服C2f在红外场景下面临感受野不足、噪声敏感、细节丢失等问题。
传统目标检测模型对所有区域或通道赋予相同权重,导致背景噪声干扰或关键特征被稀释,利用注意力机制可以使模型聚焦于目标区域和重要通道。然而,在红外目标检测中,现有多数注意力机制如SE[41]、CBAM不能有效捕获红外目标必需的局部上下文,加之采用固定权重相加或拼接,无法根据目标尺寸动态调整融合策略,导致出现小目标漏检、大目标模糊。针对上述问题,本文设计了交叉注意力融合模块CAFM,由交叉交互注意力CA (cross-interaction attention)和动态简易门控模块DG (dynamic gating)构成,如图4所示。
CA建立左右分支特征(浅层特征与深层特征)的跨层级语义关联,增强多尺度特征交互能力。如图中CA部分所示,左右分支输入特征通过投影层分别生成查询矩阵$ {\boldsymbol{V}}_{\rm{R}} $$ {\boldsymbol{V}}_{\rm{L}} $,计算方法见式(9),查询矩阵$ {\boldsymbol{V}}_{\rm{R}} $进行转置后用于下一步运算。通过矩阵乘法计算后引入Softmax函数处理,生成跨分支的注意力权重矩阵$ {{\boldsymbol{Attention}}} $,计算方法见式(10),计算过程中引入$ \sqrt{C} $用于优化分布,防止点积过大导致梯度不稳。
$ {{\boldsymbol{V}}}_{{\mathrm{L}}}=DS Conv({{\boldsymbol{X}}}_{{\mathrm{L}}}),{{\boldsymbol{V}}}_{{\mathrm{R}}}=DSConv({{\boldsymbol{X}}}_{{\mathrm{R}}})\;, $
$ {{\boldsymbol{Attention}}}={ S oftmax }\left(\frac{{\boldsymbol{V}}_{\rm{R}}^{\mathrm{T}} {\boldsymbol{V}}_{\rm{L}}}{\sqrt{C}}\right) \in {\mathbb{R}}^{H \times W \times W}\;. $
利用注意力权重矩阵加权,实现有效的左右视图交互,增强左右输入特征并应用$ 1\times 1$卷积层保持左右增强特征通道数一致。投影层采用深度可分离卷积进行轻量化改造。CA通过左右分支特征互补捕获全局关联语义,实现图像的跨视图信息增强,特征增强计算如下:
$ \begin{split}{\boldsymbol{F}}_{{\mathrm{L}} \rightarrow {\mathrm{R}}} & ={Conv}_{1 \times 1}\left({ {{\boldsymbol{Attention}}} } \otimes {\boldsymbol{V}}_{\rm{R}}\right)\;, \\{\boldsymbol{F}}_{{\mathrm{R}} \rightarrow {\mathrm{L}}} & ={Conv}_{1 \times 1}\left({ {{\boldsymbol{Attention}}}}^T \otimes {\boldsymbol{V}}_{\rm{L}}\right)\end{split}\;, $
DG融合局部细节与全局语义线索,动态生成空间-通道自适应的权重矩阵,实现像素级特征优选,解决小目标细节丢失和噪声干扰问题,原始特征$ {{\boldsymbol{X}}}_{{\mathrm{L}}} $$ {{\boldsymbol{X}}}_{{\mathrm{R}}} $与交叉注意力增强特征$ {{\boldsymbol{F}}}_{{\mathrm{R}}\rightarrow {\mathrm{L}}} $$ {{\boldsymbol{F}}}_{{\mathrm{L}}\rightarrow {\mathrm{R}}} $沿通道维度拼接:
$ {\boldsymbol{X}}_{Concat}=Concat({\boldsymbol{X}}_{L},{\boldsymbol{X}}_{R},{F}_{R\rightarrow L},{F}_{L\rightarrow R})\;. $
使用内核长度为5的卷积来捕获局部上下文,并引入Sigmoid函数生成注意力权重$ \alpha$,让模型自身获得了根据输入内容动态调整感受野的能力。将该注意力权重归一化至$ [0,1]$$ \alpha$动态调节左右分支贡献比例生成抗噪门控输出:
${\boldsymbol{ Y}}=\alpha \odot {{\boldsymbol{X}}}_{{\mathrm{R}}}+(1-\alpha )\odot {{\boldsymbol{X}}}_{{\mathrm{L}}}\;. $
CAFM通过轻量化的跨视图交互机制,对左右视图互补信息进行有效融合,实现多尺度特征高效提取。AFITDYOLO颈部网络用于红外目标多尺度特征融合与提取,在颈部网络输出前,使用交叉注意力融合模块CAFM,优化检测头的输入特征质量。该模块左右输入分别对应浅层输出特征图和深层输出特征图,通过动态多尺度特征增强与噪声感知抑制,显著提升了网络的红外目标检测的精度与鲁棒性。
本节首先描述数据集、评估指标和训练细节,然后展示不同模块对模型的性能影响,并进行目标检测可视化,与现有主流模块模型对比,最后验证提出模型的泛化能力。
数据集与评价指标。为了验证本文方法对不同尺度的红外目标检测能力,使用两个不同复杂背景下的红外数据集,单帧无人机飞鸟数据集SIRST-UAVB[8]与红外监控摄像头实拍下的道路行人车辆数据集CTIR[42]。CTIR包含城市道路和高速公路场景下,在夜间及多种恶劣天气条件中采集的11938张红外热图像,灰度8~14,其中目标包含34078辆汽车(car)、31035个行人(pedestrian)、16524个骑自行车人(cyclist)、2404辆公共汽车(bus)和1886辆卡车(truck),每张红外图像中目标种类较多,且同时包含大小尺度目标。SIRST-UAVB,由3000张针对无人机和鸟类的红外图像组成,这些图像包括1742个鸟类(bird)和2955个无人机(UAV),是一年多来收集的不同季节、天气条件和复杂背景的图像。该数据集中目标方向、尺度不同且受遮挡等挑战,微小尺度目标占比较高。数据集中部分红外图像如图5所示。本文使用精度P和召回率R来评估模型检测的准确性,mAP50及mAP95 来评估检测目标位置准确性,参数量Params作为轻量化指标,浮点运算量FLOPs衡量网络计算复杂度,每秒处理帧率FPS衡量检测速度。如式(14)所示,mAP50是指在IoU阈值为0.5时计算得到的平均精度均值,mAP95是一项更为综合的评估指标,它在IoU从0.5到0.95的范围内,以0.05为步长,取多个不同阈值下的平均精度结果。TP表示真正例,FP为假正例,FN为假反例,n为缺陷的类别数量。
$ \begin{split}& P=\frac{T P}{T P+F P}, \\&R=\frac{T P}{T P+F N}, \\& A P=\int_0^1 P(R) {\mathrm{d}} R, \\&m A P=\frac{1}{n} \sum_{i=1}^n A P_i\end{split}\;. $
训练细节。为了验证本文提出的方法可以显著提高模型检测能力,在包含多尺度目标的CTIR数据集进行训练,同时为了进一步验证方法的自适应性,在背景更为复杂、目标尺度更小的SIRST-UAVB数据集上也进行训练。两个数据集均按6:2:2的比例划分为训练集、测试集和验证集。本文实验均使用RTX 3090 GPU上的PyTorch框架,在训练阶段,采用SGD作为优化器,初始学习率为0.01,耐心值设置为80,其余超参数均为默认。输入图像尺寸为640 pixel×640 pixel、batchsize大小为40、epochs为500。
为了评估每个组件对AFITDYOLO的单独影响,使用CTIR及SIRST-UAVB两个红外数据集,对模型进行消融实验。为保证消融实验的精度和准确性,保持各组实验的环境和超参数一致,实验结果如表1所示,其中最优性能数据以粗体和下划线突出显示。本文选择 YOLOv10n为基线。由表1中实验结果可知,第一组实验基于MFFM模块改进了颈部网络,增强了网络的多尺度特征融合的表达能力,各项指标均有提升,CTIR数据集中,mAP50、mAP95涨幅更明显,均达4.1%。使用MFEConv模块改造主干网络中的C2f模块,对模型轻量化改造,MFEConv模块实现了检测精度的提升和参数量的降低。使得主干网络输出特征更具灵活性和适应性,能够在复杂场景下保持较高性能,相比于基线PR、mAP50、mAP95均有着明显的增长,Params、FLOPs和 FPS均为最佳,提高了模型检测效率。模型颈部网络输出前使用CAFM,通过浅层特征与深层特征的对比交互处理,使模型在略微增加计算复杂度和模型参数量的情况下,显著提高了检测的准确性。通过任务对齐机制,以及分类和回归分支的动态选择机制,提高了特征利用率和检测精度,P与mAP50提升均大于3%,mAP95涨幅也达 2.7%。第5组在第2组基础上继续集成MFEConv模块,观察到使用双模块的实验结果中,各项指标中均优于单模块。第6组在第5组基础上集成CAFM模块,即为本文方法,以更少的参数量Params和浮点运算量FLOPs,实现了更高的推理速度FPS,效率优势明显。在CTIR数据集上,其余各项评估指标均达到最高。随后,在更具挑战的 SIRST-UAVB数据集上依旧表现稳健,进一步证明了本文方法高效、精准且适应性强。
AFITDYOLO模型主干、颈部和检测头网络的是基于YOLOv10模型上的改进,AFITDYOLO与YOLOv10结构相似。因此选择轻量化检测模型YOLOv10n作为可视化实验的对照组,实验结果如图6所示,图中每个数据集的第一行为YOLOv10n可视化结果,第二行为AFITDYOLO可可视化结果。可以看出,对于CTIR数据集,本文所提方法在背景复杂且目标较多的情况下可以更准确检测到不同尺度目标,准确检测到各种遮挡的车辆与行人目标,减少了漏检和误检的发生概率。在SIRST-UAVB数据集中,对于红外远距离的检测目标置信度更高,因此微小目标的定位也更精准。
AFITDYOLO很大程度上减少了漏检和误检的概率,该模型在多种复杂环境下的大小尺度目标的红外检测中,取得了良好的检测结果,可应用于多种环境中的检测任务,实现跨场景部署。
本文将MFEConv卷积模块与现主流卷积模块进行对比实验,结果如表2所示。本文选择以下卷积模块作为对照组:可变核卷积AKConv[39]、深度可分离卷积DSConv、大型选择性核卷积LSKConv[43]及风车型卷积PConv,为保证实验的有效性和可行性,其他模块均采用与本文MFEConv相同的模块集成方法,每个卷积模块替换C2f中Bottleneck的Conv层。对于SIRST-UAVB数据集,可以看出MFEConv在$ k=3 $时优于PConv之外的所有卷积模块,而在目标尺度更大的CTIR数据集中,其表现则优于包括PConv的所有卷积模块。MFEConv在$ k=7 $时性能达到最优,但Params和FLOPs为最高值,同时FPS下降明显。
通过堆叠奇数卷积核或增加卷积核长度能扩展感受野[41]。由表2中可以看出增加MFEConv的卷积核的长度并没有产生可观的性能增益,同时卷积核长度的增加,参数量和浮点计算量会相应增加,降低检测速度。对于MFEConv异构卷积组,采用6个方向的卷积核进行构建相比于3个方向的卷积核参数量略微增加,但检测精度大幅提高。MFEConv通过堆叠奇数卷积核来扩展感受野,并且构建了由6个短卷积核构成的卷积层,实现该模块感受野全覆盖。由表2可以看出,设定卷积核数量为6,卷积核长度$ k=3 $的情况下,在损失较少性能的同时实现最佳的模块轻量化效果,并且考虑到整体网络参数量,本文设定卷积核长为3,使整体网络的综合性能最佳。
为了进一步验证模型的有效性,按照相同的实验环境,在本文的数据集上,将本文所提算法与其他主流深度学习检测算法进行对比。实验结果由表3所示,可以看出,AFITDYOLO检测精度优于实验所提单模态目标检测模型(1~6行),其中FR-CNN[44]为经典的二阶检测算法,mAP50达到了84.0%,但其参数量大,检测速度低。在YOLO系列算法中,YOLOv11n[45]、YOLOv12n[46]通过轻量化改造,减少了模型的参数量,但损失了一定的检测性能。D-FINE-N[47]、DEIM[48]是基于Transformer架构的两个目标检测模型,在不增加额外推理和训练成本的情况下,实现良好的检测性能,但准确性和检测速度均低于本文方法。
综上,从实验结果可知,本文提出的模型参数量低于目前主流模型,轻量化效果优秀,同时实现了良好检测精度,能够满足实时检测需求。
本文选取高空无人机红外热成像数据集HIT-UAV[41]和经典红外小目标数据集IRSTD-1k[49]进行泛化性实验。HIT-UAV数据集包含不同场景下无人机拍摄的2898张红外图像。该数据集的图像包括人员、自行车、汽车、其他车辆等目标。IRSTD-1k数据集由1000张红外图像组成,图像尺寸为512 pixel×512 pixel,包含不同环境下的无人机、生物、船只和车辆四个目标。选择YOLOv10n、YOLOv12n与DEIM-N作为实验对照组。实验结果如表4所示,HIT-UAV数据集上,AFITDYOLO与其他算法相比,P值略差于DEIM-N,其余指标均达到最高值。在IRSTD-1k数据集上,各项指标均最佳。可以看出,本文方法在泛化实验所用的红外数据集上也实现了检测精度的提高,验证了本文所提方法的泛化性,进一步说明本文方法可以跨场景部署。
本文提出的基于YOLO的自适应多尺度红外目标检测网络AFITDYOLO,在较低参数量情况下提高了不同背景下,不同尺度红外目标的检测精度。本文提出多尺度特征融合模块MFFM,通过改进FPN中不同层特征之间的相关性,增强了模型多尺度特征融合的表达能力。本文还设计了轻量化卷积模块MFEConv,利用红外图像的目标分布特性,以最小的参数实现高效、更大的感受野。接着引入交叉注意力融合模块CAFM,通过不同层输出特征图的对比交互,突出重要特性信息,滤除无关背景信息和抑制噪声,进一步增强模型特征表达能力。实验结果表明,本文方法优于目前主流的方法,展示了良好的检测精度、轻量化程度和泛化能力,并且具有跨场景部署能力。

参考文献 引证文献
排序方式:
1
Liu R M, Lu Y H, Gong C L, et al. Infrared point target detection with improved template matching[J]. Infrared Phys Technol, 2012, 55(4): 380−387.
2
Rivest J F, Fortin R. Detection of dim targets in digital infrared imagery by morphological image processing[J]. Opt Eng, 1996, 35(7): 1886−1893.
3
Dai Y M, Wu Y Q, Zhou F, et al. Attentional local contrast networks for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2021, 59(11): 9813−9824.
4
Du P, Hamdulla A. Infrared small target detection using homogeneity-weighted local contrast measure[J]. IEEE Geosci Remote Sensing Lett, 2020, 17(3): 514−518.
5
Liu Y J, Liu X Y, Hao X Y, et al. Single-frame infrared small target detection by high local variance, low-rank and sparse decomposition[J]. IEEE Trans Geosci Remote Sensing, 2023, 61: 5614317.
6
Li R H, Shen Y. YOLOSR-IST: a deep learning method for small target detection in infrared remote sensing images based on super-resolution and YOLO[J]. Signal Proc, 2023, 208: 108962.
7
Yang B, Zhang X Y, Zhang J, et al. EFLNet: enhancing feature learning network for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2024, 62: 5906511.
8
GUPTA A, GUPTA U. Real time target detection for infrared images[C]//2020 Fourth International Conference on Inventive Systems and Control (ICISC), 2020: 570–574. https://doi.org/10.1109/ICISC47916.2020.9171208.
9
Yang J N, Liu S L, Wu J J, et al. Pinwheel-shaped convolution and scale-based dynamic loss for infrared small target detection[C]//Proceedings of the 39th AAAI Conference on Artificial Intelligence, 2025: 9202–9210. https://doi.org/10.1609/aaai.v39i9.32996.
10
赵斌, 王春平, 付强, 等. 基于深度注意力机制的多尺度红外行人检测[J]. 光学学报, 2020, 40(5): 0504001.
Zhao B, Wang C P, Fu Q, et al. Multi-scale infrared pedestrian detection based on deep attention mechanism[J]. Acta Opt Sin, 2020, 40(5): 0504001.
11
Dalal N, Triggs B. Histograms of oriented gradients for human detection[C]//2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2005: 886–893. https://doi.org/10.1109/CVPR.2005.177.
12
何洪英, 姚建刚, 蒋正龙, 等. 基于支持向量机的高压绝缘子污秽等级红外热像检测[J]. 电力系统自动化, 2005, 29(24): 70−74,82.
He H Y, Yao J G, Jiang Z L, et al. Infrared thermal image detecting of high voltage insulator contamination grades based on support vector machine[J]. Autom Electr Power Syst, 2005, 29(24): 70−74,82.
13
Felzenszwalb P F, Girshick R B, McAllester D, et al. Object detection with discriminatively trained part-based models[J]. IEEE Trans Pattern Anal Mach Intell, 2010, 32(9): 1627−1645.
14
Mathieu M, LeCun Y, Fergus R, et al. OverFeat: integrated recognition, localization and detection using convolutional networks[C]//International Conference on Learning Representations, 2014.
15
Liu S, Qi L, Qin H F, et al. Path aggregation network for instance segmentation[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018: 8759–8768. https://doi.org/10.1109/CVPR.2018.00913.
16
Ghiasi G, Lin T Y, Le Q V. NAS-FPN: learning scalable feature pyramid architecture for object detection[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019: 7029–7038. https://doi.org/10.1109/CVPR.2019.00720.
17
Tan M X, Pang R M, Le Q V. EfficientDet: scalable and efficient object detection[C]//2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020: 10778–10787. https://doi.org/10.1109/CVPR42600.2020.01079.
18
Min K, Lee G H, Lee S W. Attentional feature pyramid network for small object detection[J]. Neural Netw, 2022, 155: 439−450.
19
Dai Y M, Li X, Zhou F, et al. One-stage cascade refinement networks for infrared small target detection[J]. IEEE Trans Geosci Remote Sensing, 2023, 61: 5000917.
20
Dai Y M, Wu Y Q, Zhou F, et al. Asymmetric contextual modulation for infrared small target detection[C]. 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), 2021: 949–958. https://doi.org/10.1109/WACV48630.2021.00099.
21
Ren D D, Li J B, Han M, et al. DNANet: dense nested attention network for single image dehazing[C]//ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2021: 2035–2039. https://doi.org/10.1109/ICASSP39728.2021.9414179.
22
Fan W Q, Xu X M, Cai B L, et al. ISNet: individual standardization network for speech emotion recognition[J]. IEEE/ACM Trans Audio Sp Lang Process, 2022, 30: 1803−1814.
23
Woo S, Park J, Lee J Y, et al. CBAM: convolutional block attention module[C]//15th European Conference on Computer Vision, 2018: 3–19. https://doi.org/10.1007/978-3-030-01234-2_1.
24
Dai T, Cai J R, Zhang Y B, et al. Second-order attention network for single image super-resolution[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019: 11057–11066. https://doi.org/10.1109/CVPR.2019.01132.
25
Dai T, Zha H, Jiang Y, et al. Image super-resolution via residual block attention networks[C]//2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), 2019: 3879–3886. https://doi.org/10.1109/ICCVW.2019.00481.
26
Niu B, Wen W L, Ren W Q, et al. Single image super-resolution via a holistic attention network[C]//16th European Conference on Computer Vision – ECCV 2020, 2020: 191–207. https://doi.org/10.1007/978-3-030-58610-2_12.
27
Wang L, Shen J, Tang E, et al. Multi-scale attention network for image super-resolution[J]. J Vis Commun Image Repres, 2021, 80: 103300.
28
Varghese R, M S. YOLOv8: a novel object detection algorithm with enhanced performance and robustness[C]//2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), 2024: 1–6. https://doi.org/10.1109/ADICS58448.2024.10533619.
29
Terven J, Córdova-Esparza D M, Romero-González J A. A comprehensive review of YOLO architectures in computer vision: from YOLOv1 to YOLOv8 and YOLO-NAS[J]. Mach Learn Knowl Extr, 2023, 5(4): 1680−1716.
30
Chen H, Chen K, Ding G G, et al. YOLOv10: real-time end-to-end object detection[C]//Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024: 107984–108011. https://doi.org/10.52202/079017-3429.
31
Zhao F, Li S J, Zhang J J, et al. Convolution transformer fusion splicing network for hyperspectral image classification[J]. IEEE Geosci Remote Sensing Lett, 2023, 20: 5501005.
32
Glorot X, Bengio Y. Understanding the difficulty of training deep feedforward neural networks[C]//Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, 2010: 249–256.
33
Ioffe S, Szegedy C. Batch normalization: accelerating deep network training by reducing internal covariate shift[C]//Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37, 2015: 448–456.
34
Bieder F, Sandkühler R, Cattin P C. Comparison of methods generalizing max- and average-pooling[Z]. arXiv: 2103.01746, 2021. https://doi.org/10.48550/arXiv.2103.01746.
35
Kusupati A, Ramanujan V, Somani R, et al. Soft threshold weight reparameterization for learnable sparsity[C]//Proceedings of the 37th International Conference on Machine Learning, 2020: 5544–5555.
36
Chollet F. Xception: deep learning with depthwise separable convolutions[C]//2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017: 1800–1807. https://doi.org/10.1109/CVPR.2017.195.
37
Xu W, Wan Y. ELA: efficient local attention for deep convolutional neural networks[Z]. arXiv: 2403.01123, 2024. https://doi.org/10.48550/arXiv.2403.01123.
38
Fløystad G. Shift modules, strongly stable ideals, and their dualities[J]. Trans Am Math Soc Ser B, 2023, 10(21): 670−714.
39
Zhang X, Song Y Z, Song T T, et al. AKConv: convolutional kernel with arbitrary sampled shapes and arbitrary number of parameters[Z]. arXiv: 2311.11587, 2023. https://doi.org/10.48550/arXiv.2311.11587.
40
Luo Y, Wong Y, Kankanhalli M, et al. G-softmax: improving intraclass compactness and interclass separability of features[J]. IEEE Trans Neural Netw Learn Syst, 2020, 31(2): 685−699.
41
Zhao X F, Zhang W W, Zhang H, et al. ITD-YOLOv8: an infrared target detection model based on YOLOv8 for unmanned aerial vehicles[J]. Drones, 2024, 8(4): 161.
42
Li Y X, Zou W B, Wei Q M, et al. Multi-level feature fusion network for lightweight stereo image super-resolution[C]. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2024: 6489–6498. https://doi.org/10.1109/CVPRW63382.2024.00649.
43
Wang X, He N, Hong C, et al. Improved YOLOX-X Based UAV aerial photography object detection algorithm[J]. Image Vision Comput, 2023, 135: 104697.
44
Maity M, Banerjee S, Chaudhuri S S. Faster R-CNN and YOLO based vehicle detection: a survey[C]//2021 5th International Conference on Computing Methodologies and Communication (ICCMC), 2021: 1442–1447. https://doi.org/10.1109/ICCMC51019.2021.9418274.
45
Sani A R, Zolfagharian A, Kouzani A Z. Automated defects detection in extrusion 3D printing using YOLO models[J]. J Intell Manuf, 2024. https://doi.org/10.1007/s10845-024-02543-8.
46
Alif M A R, Hussain M. YOLOv12: a breakdown of the key architectural features[Z]. arXiv: 2502.14740, 2025. https://doi.org/10.48550/arXiv.2502.14740.
47
Peng Y S, Li H B, Wu P X, et al. D-FINE: redefine regression task in DETRs as fine-grained distribution refinement[Z]. arXiv: 2410.13842, 2024. https://doi.org/10.48550/arXiv.2410.13842.
48
Huang S H, Lu Z C, Cun X, et al. DEIM: DETR with improved matching for fast convergence[C]//2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025: 15162–15171. https://doi.org/10.1109/CVPR52734.2025.01412.
49
Zhang M J, Zhang R, Yang Y X, et al. ISNet: shape matters for infrared small target detection[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022: 867–876. https://doi.org/10.1109/CVPR52688.2022.00095.
2026年第53卷第4期
PDF下载
120
52
引用本文
BibTeX
文章信息
doi: 10.12086/oee.2026.250292
  • 接收时间:2025-09-27
  • 首发时间:2026-07-02
  • 出版时间:2026-04-24
补充材料
相关文章
文章信息
作者
出版历史
  • 收稿日期:2025-09-27
  • 修回日期:2026-01-28
  • 录用日期:2026-01-16
基金
作者信息
    1浙江理工大学信息科学与工程学院(网络安全学院),浙江 杭州 310018
    2嘉兴大学人工智能学院,浙江 嘉兴 314001

通讯作者:

参考文献
分享链接
https://castjournals.cast.org.cn/joweb/oee/CN/10.12086/oee.2026.250292
分享至
全文二维码

扫描看全文

引用本文
BibTeX
本文的引用情况
2种不同金属材料的力学参数

Family
属数
Number of
genus
种数
Number of
species
占总种数比例
Percentage of
total species (%)

Genus
种数
Number of
species
占总种数比例
Percentage of total
species (%)
鹅膏菌科Amanitaceae 2 11 5.26 鹅膏菌属 Amanita 10 4.78
小菇科 Mycenaceae 2 12 5.74 丝盖伞属 Inocybe 5 2.39
多孔菌科 Polyporaceae 8 14 6.70 蜡蘑属 Laccaria 5 2.39
红菇科 Russulaceae 3 23 11.00 小皮伞属 Marasmius 6 2.87
小菇属 Mycena 11 5.26
光柄菇属 Pluteus 5 2.39
红菇属 Russula 17 8.13
栓菌属 Trametes 5 2.39
关闭全屏