Explicit Spatial Localization and Task-Adaptive Balancing for Remote Sensing Image-Text Retrieval
Spatial grounding and adaptive task coordination for cross-modal retrieval
ISPRS Journal of Photogrammetry and Remote Sensing, 2026

Method overview supplied by the project author.
I. Overview
Remote-sensing captions often refer to objects and relationships that occupy only a small part of a large, visually complex scene. This work brings explicit spatial localization into image–text retrieval and coordinates it with global semantic alignment.
II. Key Contributions
- Uses optimal-transport-guided query selection to identify visual queries relevant to a text description.
- Introduces explicit spatial cues that connect semantic entities to localized image regions.
- Balances retrieval and localization objectives adaptively instead of assigning them fixed weights.
III. Methodology
The supplied overview shows a visual–language backbone followed by Sinkhorn-based query selection. A bottom-up spatial pathway recovers object cues and anchors, while a top-down semantic pathway filters representations according to the text. A task-adaptive balancing module coordinates the spatial and retrieval losses during training.
IV. Research Focus
The method is designed for retrieval cases in which global scene similarity is insufficient and the queried content must be grounded in a particular region. The explicit localization task supplies additional supervision for learning spatially aware cross-modal representations.
Reference
Citation
BibTeX citation
@article{ZHENG2026109,
title = {Explicit spatial localization and task-adaptive balancing for remote sensing image-text retrieval},
journal = {ISPRS Journal of Photogrammetry and Remote Sensing},
volume = {241},
pages = {109-122},
year = {2026},
issn = {0924-2716},
doi = {https://doi.org/10.1016/j.isprsjprs.2026.08.022},
author = {Chengyu Zheng and Hanzhang Lu and Qianyu Shang and Yong Gao and Jie Nie and Shan Du},
}
