Multi-Perspective Subimage CLIP with Keyword Guidance
for Remote Sensing Image-Text Retrieval

1 Qinghai University
* First author. † Corresponding authors.

Abstract

Vision-Language Pre-training (VLP) models like CLIP have significantly advanced Remote Sensing Image-Text Retrieval (RSITR). However, existing methods predominantly rely on coarse-grained global alignment, which often overlooks the dense, multi-scale semantics inherent in overhead imagery. Moreover, adapting these heavy models via full fine-tuning incurs prohibitive computational costs and risks catastrophic forgetting.

To address these challenges, we propose MPS-CLIP, a parameter-efficient framework designed to shift the retrieval paradigm from global matching to keyword-guided fine-grained alignment. Specifically, we leverage a Large Language Model (LLM) to extract core semantic keywords, guiding the Segment Anything Model (SamGeo) to generate semantically relevant sub-perspectives. To efficiently adapt the frozen backbone, we introduce a Gated Global Attention (G2A) adapter, which captures global context and long-range dependencies with minimal overhead. Furthermore, a Multi-Perspective Representation (MPR) module aggregates these local cues into robust multi-perspective embeddings.

The framework is optimized via a hybrid objective combining multi-perspective contrastive and weighted triplet losses, which dynamically selects maximum-response perspectives to suppress noise and enforce precise semantic matching. Extensive experiments on the RSICD and RSITMD benchmarks demonstrate that MPS-CLIP achieves state-of-the-art performance with 35.18% and 48.40% mean Recall (mR), respectively, significantly outperforming full fine-tuning baselines and recent competitive methods.

Methodology

Pipeline
Figure 1. Overall pipeline of MPS-CLIP. Keywords from LLMs guide SamGeo to generate sub-images, processed by a shared CLIP backbone with G2A adapters.
Adapter Structure
Figure 2. Structure of the Gated Global Attention (G2A) adapter.

Experimental Results

Performance comparison with state-of-the-art methods on RSICD and RSITMD benchmarks.

Methods Text Retrieval Image Retrieval mR
R@1R@5R@10 R@1R@5R@10
Train from Scratch
KAMCL11.9927.1738.337.7825.2340.0225.08
VGSGN8.3321.8732.576.5323.1336.8521.55
DOVE8.6622.3534.956.0423.9540.3522.72
PIR-ITR10.8926.1737.797.1725.0741.0624.69
CLIP-based
SkyCLIP6.5916.1026.537.1422.3434.2918.83
CLIP-Adapter7.1119.4831.017.6724.8739.7321.65
Linear-probe CLIP8.4624.4137.727.8125.8942.4724.46
UniAdapter12.6530.8142.749.6130.0647.1628.84
PE-RSITR14.1331.5144.7811.6333.9250.7331.12
Full-FT CLIP13.5430.8343.4611.5533.1449.8330.39
HarMA16.3634.4847.7412.9237.1753.0733.62
MPS-CLIP (Ours) 18.30 37.42 50.32 13.28 37.04 54.73 35.18
Methods Text Retrieval Image Retrieval mR
R@1R@5R@10 R@1R@5R@10
Train from Scratch
KAMCL16.8134.9647.1214.6041.8659.8635.87
VGSGN14.1634.9650.6613.2342.5763.4136.50
DOVE16.8136.8050.9312.2049.9366.5037.73
PIR-ITR18.3642.0455.5313.3644.4761.7339.25
CLIP-based
SkyCLIP10.1825.4435.6210.8833.2749.8227.54
CLIP-Adapter12.8328.8439.0513.3040.2060.0632.38
Linear-probe CLIP17.0233.1248.3513.3341.8063.8936.25
UniAdapter19.8636.3251.2817.5444.8956.4639.23
PE-RSITR23.6744.0760.3620.1050.6367.9744.47
Full-FT CLIP24.1647.1261.2820.4050.5368.5445.33
HarMA25.8148.3760.6119.9253.2771.2146.53
MPS-CLIP (Ours) 27.88 51.11 61.06 22.61 56.59 71.15 48.40

Qualitative Results

Visual comparisons on remote sensing image-text retrieval.

Ablation Studies

Detailed analysis of component effectiveness. Click tabs to view details.

Effectiveness of the [CLS] token compared to Mean Pooling.

Method Text R@1Img R@1 mR
Mean Pool14.2710.5231.37
[CLS] Token18.3013.2835.18

Impact of Attention (Attn) and Gating mechanisms within the G2A adapter.

AttnGate mR Params (M)
××34.340.49
✓×34.590.51
×✓34.990.49
✓✓35.180.51

Effect of aggregating keyword-guided sub-perspectives via the MPR module.

MPR Module Text R@1Img R@1 mR
w/o MPR17.3813.3634.50
w/ MPR18.3013.2835.18

Contribution of Multi-Perspective Contrastive (MPC) and Weighted Triplet (MPT) losses.

LBaseLMPCLMPT mR
✓34.50
✓✓34.65
✓✓34.74
✓✓✓35.18

Citation

@article{MP_Subimage_CLIP_2025,
  title   = {Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval},
  author  = {Li, Yifan and Wang, Shiying and Huang, Jianqiang},
  journal = {International Conference on Multimedia and Expo (ICME)},
  year    = {2025}
}