收藏切换
Remote sensing image super-resolution reconstruction via multi-scale augmented state space model
收藏切换
PDF
Jiacheng Chen1, Fei Wu1, *, Hangyao Tu2, Jiawei Jiang3, Wanliang Wang3
Opto-Electronic Engineering | 2026, 53(4) : 250304
Less
收藏切换
Opto-Electronic Engineering | 2026, 53(4): 250304
Article
Remote sensing image super-resolution reconstruction via multi-scale augmented state space model
Full
Jiacheng Chen1, Fei Wu1, *, Hangyao Tu2, Jiawei Jiang3, Wanliang Wang3
Affiliations
  • 1College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China
  • 2College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, 310015, China
  • 3College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou, Zhejiang 310023, China
Published: 2026-04-24 doi: 10.12086/oee.2026.250304
Outline
收藏切换

Convolutional neural networks (CNNs) and vision transformers (ViTs) represent the two dominant paradigms in remote sensing single image super-resolution (RSSISR), each with distinct strengths. While CNNs have long been the workhorse due to their inductive biases, vision transformers have recently demonstrated superior performance in many cases, primarily attributed to their exceptional capability in modeling long-range dependencies through the self-attention mechanism. However, this advantage comes at a significant cost: the self-attention mechanism suffers from quadratic computational complexity with respect to image size. This inherent limitation becomes a critical bottleneck in RSSISR, where generating high-resolution outputs from low-resolution inputs demands extensive computation, severely restricting the practical deployment of ViTs for large-area remote sensing imagery.

To effectively overcome this fundamental challenge, we propose a novel architecture named the multi-scale augmented state space model (MS3M). Our approach is grounded in the recent advancements of state space models (SSMs), which are renowned for their linear computational complexity and strong potential in capturing long-range interactions. Unlike existing SSM-based feature extraction methods that often rely on fixed, unidirectional scanning paths, our MS3M introduces a grouped parallel scanning strategy. This design efficiently captures comprehensive global and non-local features without being constrained by a single scanning direction, all while rigorously maintaining linear computational complexity, thereby ensuring high efficiency.

Furthermore, acknowledging the inherent and critically important multi-scale spatial structures present in remote sensing images—from fine-grained textures of buildings to extensive patterns of farmlands—we embed a multi-receptive-field aggregation mechanism directly into the state space model. This allows our network to seamlessly integrate contextual information across different scales, a capability essential for accurately reconstructing complex geographical objects. To further enhance the representation power of local features, we design a novel high-order moment channel affinity modulation module. This module moves beyond simple first-order statistics to optimize feature expressions, enabling more nuanced and powerful feature transformations within the network. The entire MS3M framework is constructed upon a U-shaped architecture to facilitate effective multi-level feature fusion across the encoder and decoder.

We conduct extensive experiments on several public remote sensing datasets. The results demonstrate that our proposed MS3M achieves state-of-the-art performance, outperforming existing leading methods in terms of objective metrics including PSNR, SSIM and LPIPS, as well as in subjective visual quality. The superior results validate the effectiveness of our architectural choices and underscore MS3M's advancement as a robust and efficient solution for the challenging task of remote sensing image super-resolution.

remote sensing image super-resolution  /  state space model  /  grouped parallel scanning  /  multi-receptive field aggregation  /  high-order moment statistics
Jiacheng Chen, Fei Wu, Hangyao Tu, Jiawei Jiang, Wanliang Wang. Remote sensing image super-resolution reconstruction via multi-scale augmented state space model[J]. Opto-Electronic Engineering, 2026 , 53 (4) : 250304 - . DOI: 10.12086/oee.2026.250304
Year 2026 volume 53 Issue 4
PDF
151
67
Cite this Article
BibTeX
Article Info
doi: 10.12086/oee.2026.250304
  • Receive Date:2025-10-10
  • Online Date:2026-07-02
  • Published:2026-04-24
Article Data
Affiliations
History
  • Received:2025-10-10
  • Revised:2026-01-13
  • Accepted:2026-01-14
Affiliations
    1College of Artificial Intelligence, Jiaxing University, Jiaxing, Zhejiang 314001, China
    2College of Computer Science and Technology, Zhejiang University, Hangzhou, Zhejiang, 310015, China
    3College of Computer Science and Technology, Zhejiang University of Technology, Hangzhou, Zhejiang 310023, China

Corresponding:

References
Share
https://castjournals.cast.org.cn/joweb/oee/EN/10.12086/oee.2026.250304
Share to
QR

Scan QR to access full text

Cite this article
BibTeX
Citations
表12种不同金属材料的力学参数

Family
属数
Number of
genus
种数
Number of
species
占总种数比例
Percentage of
total species (%)

Genus
种数
Number of
species
占总种数比例
Percentage of total
species (%)
鹅膏菌科Amanitaceae 2 11 5.26 鹅膏菌属 Amanita 10 4.78
小菇科 Mycenaceae 2 12 5.74 丝盖伞属 Inocybe 5 2.39
多孔菌科 Polyporaceae 8 14 6.70 蜡蘑属 Laccaria 5 2.39
红菇科 Russulaceae 3 23 11.00 小皮伞属 Marasmius 6 2.87
小菇属 Mycena 11 5.26
光柄菇属 Pluteus 5 2.39
红菇属 Russula 17 8.13
栓菌属 Trametes 5 2.39
关闭全屏
  • BibTeX
  • EndNote
  • RefWorks
  • TxT