Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works8 from public data

TitleCited by
  • Masked Generative Adversarial Networks Are Data-Efficient Generation Learners

    Jiaxing Huang, Kaiwen Cui, Dayan Guan, Aoran Xiao, Fangneng Zhan, Shijian Lu, Shengcai Liao, Eric Xing

    neural information processing systems · 2022

    This paper develops two masking strategies that work along orthogonal dimensions of training images, including a shifted spatial masking that masks the images in spatial dimensions with random shifts, and a balanced spectral masker that masks certain image spectral bands with self-adaptive probabilities.

    33
  • SARLANG-1M: A Benchmark for Vision–Language Modeling in SAR Image Understanding

    Yimin Wei, Aoran Xiao, Yexian Ren, Yuting Zhu, Hongruixuan Chen, Junshi Xia, Naoto Yokoya

    IEEE Transactions on Geoscience and Remote Sensing · 2026

    SARLANG-1M, a large-scale benchmark tailored for multimodal SAR image understanding, with a primary focus on integrating SAR with textual modality, is introduced, demonstrating that fine-tuning with SARLANG-1M significantly enhances their performance in SAR image interpretation, reaching performance comparable to human experts.

    31
  • Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation

    Yun Xing, Jian Kang, Aoran Xiao, Jiahao Nie, Ling Shao, Shijian Lu

    neural information processing systems · 2023

    Concept Curation (CoCu) is proposed, a pipeline that leverages CLIP to compensate for the missing semantics of visual concepts and greatly boosts language-supervised segmentation baseline by a large margin, suggesting the value of bridging semantic gap in pre-training data.

    29
  • Segment Anything with Multiple Modalities

    Aoran Xiao, Weihao Xuan, Heli Qi, Y Xing, Naoto Yokoya, Shijian Lu

    arXiv · 2024

    MM-SAM is developed, an extension and expansion of SAM that supports cross-modal and multi-modal processing for robust and enhanced segmentation with different sensor suites, and addresses three main challenges: adaptation toward diverse non-RGB sensors for single-modal processing, synergistic processing of multi-modal data via sensor fusion, and mask-free training for different downstream tasks.

    14
  • PolarMix: A General Data Augmentation Technique for LiDAR Point Clouds

    Aoran Xiao, Jiaxing Huang, Dayan Guan, Kaiwen Cui, Shijian Lu, Ling Shao

    neural information processing systems · 2022

    11
  • GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

    Aoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren, Hongruixuan Chen, Naoto Yokoya

    arXiv · 2026

    GeoMMBench is introduced, a comprehensive multimodal question-answering benchmark covering diverse RS disciplines, sensors, and tasks, enabling broader and more rigorous evaluation than prior benchmarks and proposed GeoMMAgent, a multi-agent framework that strategically integrates retrieval, perception, and reasoning through domain-specific RS models and tools.

    6
  • MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

    Yimin Wei, Aoran Xiao, Hongruixuan Chen, Junshi Xia, Naoto Yokoya

    arXiv · 2026

    This work presents MM-OVSeg, a multimodal Optical-SAR fusion framework for resilient open-vocabulary segmentation under adverse weather conditions, with two key designs: a cross-modal unification process for multi-sensor representation alignment, and a dual-encoder fusion module that integrates hierarchical features from multiple vision foundation models for text-aligned multimodal segmentation.

    6
  • MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

    Kai Jiang, Jiaxing Huang, Jingyi Zhang, Weiying Xie, Yunsong Li, Yufei Wang, Aoran Xiao, Dacheng Tao

    Lecture notes in computer science · 2026

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.