Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks8 from public data
- Masked Generative Adversarial Networks Are Data-Efficient Generation Learners33
This paper develops two masking strategies that work along orthogonal dimensions of training images, including a shifted spatial masking that masks the images in spatial dimensions with random shifts, and a balanced spectral masker that masks certain image spectral bands with self-adaptive probabilities.
- SARLANG-1M: A Benchmark for Vision–Language Modeling in SAR Image Understanding31
SARLANG-1M, a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality, is introduced, demonstrating that fine-tuning with SARLANG-1M significantly enhances their performance in SAR image interpretation, reaching performance comparable to human experts.
- Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation29
Concept Curation (CoCu) is proposed, a pipeline that leverages CLIP to compensate for the missing semantics of visual concepts and greatly boosts language-supervised segmentation baseline by a large margin, suggesting the value of bridging semantic gap in pre-training data.
- Segment Anything with Multiple Modalities14
MM-SAM is developed, an extension and expansion of SAM that supports cross-modal and multi-modal processing for robust and enhanced segmentation with different sensor suites, and addresses three main challenges: adaptation toward diverse non-RGB sensors for single-modal processing, synergistic processing of multi-modal data via sensor fusion, and mask-free training for different downstream tasks.
- 11
- GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing6
GeoMMBench is introduced, a comprehensive multimodal question-answering benchmark covering diverse RS disciplines, sensors, and tasks, enabling broader and more rigorous evaluation than prior benchmarks and proposed GeoMMAgent, a multi-agent framework that strategically integrates retrieval, perception, and reasoning through domain-specific RS models and tools.
- MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing6
This work presents MM-OVSeg, a multimodal Optical-SAR fusion framework for resilient open-vocabulary segmentation under adverse weather conditions, with two key designs: a cross-modal unification process for multi-sensor representation alignment, and a dual-encoder fusion module that integrates hierarchical features from multiple vision foundation models for text-aligned multimodal segmentation.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.