Whole Semantic Sparse Coding Network for Remote Sensing Image-Text Retrieval
Sparse semantic coding across local and global cross-modal representations
IEEE Transactions on Geoscience and Remote Sensing, 2025

Method overview supplied by the project author.
I. Overview
Remote-sensing scenes contain multiple objects, land-cover patterns, and spatial relationships that may be described at different levels of detail. Whole semantic sparse coding seeks a compact representation that retains both global scene meaning and discriminative local semantics.
II. Key Contributions
- Integrates visual and textual features into a shared sparse semantic representation.
- Models information at multiple semantic scales instead of relying on one pooled feature.
- Uses structured cross-modal learning to improve the separation of semantically similar retrieval candidates.
III. Methodology
Image and text encoders extract modality-specific features, which are projected into a shared semantic space. The sparse coding stages shown in the supplied architecture aggregate complementary local and global information while encouraging a compact set of active semantic components for matching.
IV. Research Focus
The project investigates whether structured sparsity can reduce redundant scene information and emphasize the concepts most relevant to an image–text query.
Reference
Citation
BibTeX citation
@article{zheng2025whole,
title={Whole Semantic Sparse Coding Network for Remote Sensing Image-Text Retrieval},
author={Zheng, Chengyu and Wen, Qi and Li, Xiu and Yang, Chenxue and Nie, Jie and Guo, Yiyun and Qian, Yuntao and Wei, Zhiqiang},
journal={IEEE Transactions on Geoscience and Remote Sensing},
year={2025},
publisher={IEEE}
}

