Frequency- and Spatial-Domain Saliency Network for Remote Sensing Cross-Modal Retrieval
Complementary saliency modeling in spatial and frequency representations
IEEE Transactions on Geoscience and Remote Sensing, 2025

Method overview supplied by the project author.
I. Overview
Spatial features describe where visual structures occur, while frequency features expose complementary patterns in texture, edges, and repeated structures. This project combines both views to identify image content that is salient for a textual query.
II. Key Contributions
- Models remote-sensing imagery in both spatial and frequency domains.
- Learns saliency cues that suppress irrelevant background information before cross-modal matching.
- Fuses complementary domain features with language representations for retrieval.
III. Methodology
The supplied network overview separates visual processing into spatial and frequency branches. Each branch extracts domain-specific saliency information, and the resulting features are fused and aligned with encoded text in a shared retrieval space.
IV. Research Focus
The work explores how frequency-aware representations can complement conventional spatial features when scenes contain clutter, scale variation, and visually similar land-cover patterns.
Reference
Citation
BibTeX citation
@article{zheng2025frequency,
title={Frequency-and spatial-domain saliency network for remote sensing cross-modal retrieval},
author={Zheng, Chengyu and Nie, Jie and Yin, Bo and Li, Xiu and Qian, Yuntao and Wei, Zhiqiang},
journal={IEEE Transactions on Geoscience and Remote Sensing},
volume={63},
pages={1--13},
year={2025},
publisher={IEEE}
}

