Research questionHow can remote-sensing vision-language models align heterogeneous sensor imagery with text beyond RGB?Remote-sensing imagery varies in channel count and spectral content across sensor types. Models designed around RGB inputs may not transfer their visual representations while preserving a shared image-text semantic space. Latest papersRecent research connected to this question, newest first.Exploring the Potential of Contrastive Language-Image Pre-training for Multi-Source Remote Sensing DataApplies to remote-sensing vision-language modeling across RGB, SAR, multispectral, and hyperspectral inputs. The source reports evidence from image-text retrieval, zero-shot classification, and semantic localization experiments using the OmniRS5M corpus.research paper · Sep 3, 2026Lightweight Adaptation of General-Purpose VLMs for Multispectral and SAR Image UnderstandingThe reported protocol represents each observation as five optical views and one SAR view in a named multi-image prompt, then applies LoRA to the language network and selected visual-transformer blocks. Evidence covers a balanced six-class BigEarthNet-v2 land-cover benchmark, four VLM architectures, Sen1Floods11 flood verification, BigEarthNet.txt captioning, and image-removal and mismatch controls; Qwen3-VL reaches 0.8275 micro F1 on the land-cover benchmark.research paper · Sep 2, 2026