Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works156 from public data

TitleCited by
  • Semantic Flow for Fast and Accurate Scene Parsing

    Xiang-Tai Li, Ansheng You, Zhen Zhu, Houlong Zhao, Maoke Yang, Kuiyuan Yang, Shao-Hua Tan, Yun-Hai Tong

    Lecture notes in computer science · 2020

    This paper proposes a Flow Alignment Module (FAM) to learn Semantic Flow between feature maps of adjacent levels, and broadcast high-level features to high resolution features effectively and efficiently and exhibits superior performance over other real-time methods even on light-weight backbone networks.

    456
  • Involution: Inverting the Inherence of Convolution for Visual Recognition

    Duo Li, Jie Hu, Changhu Wang, Xiang-Tai Li, Qi She, Lei Zhu, T. Zhang, Qifeng Chen

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation.

    418
  • Transformer-Based Visual Segmentation: A Survey

    Xiang-Tai Li, Henghui Ding, Haobo Yuan, Wenwei Zhang, Jiangmiao Pang, Guangliang Cheng, Kai Chen, Zi-Wei Liu, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    This survey provides a thorough overview of transformer-based visual segmentation, summarizing recent advancements and presents a meta-architecture that unifies all recent transformer-based approaches.

    342
  • Improving Semantic Segmentation via Decoupled Body and Edge Supervision

    Xiang-Tai Li, Xia Li, Li Zhang, Guangliang Cheng, Jian-Ping Shi, Zhou-Chen Lin, Shao-Hua Tan, Yun-Hai Tong

    Lecture notes in computer science · 2020

    This paper proposes a new paradigm for semantic segmentation that establishes new state of the art while retaining high efficiency in inference and shows that the proposed framework with various baselines or backbone networks leads to better object inner consistency and object boundaries.

    313
  • Rethinking Mobile Block for Efficient Attention-based Models

    Jiang-Ning Zhang, Xiang-Tai Li, Jian Li, Liang Liu, Zhu-Cun Xue, Boshen Zhang, Zhe Jiang, Tian-Xin Huang, +2 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    This work rethinks lightweight infrastructure from efficient IRB and effective components of Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMB) for lightweight model design.

    296
  • Towards Open Vocabulary Learning: A Survey

    Jianzong Wu, Xiang-Tai Li, Haobo Yuan, Henghui Ding, Yi-Bo Yang, Xia Li, Jiang-Ning Zhang, Yu Tong, +4 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    This paper thoroughly reviews open vocabulary learning, summarizing and analyzing recent developments in the field, and juxtaposing open vocabulary learning with analogous concepts such as zero-shot learning, open-set recognition, and out-of-distribution detection.

    285
  • 255
  • Gated Fully Fusion for Semantic Segmentation

    Xiang-Tai Li, Houlong Zhao, Lei Han, Yun-Hai Tong, Shao-Hua Tan, Kuiyuan Yang

    Proceedings of the AAAI Conference on Artificial Intelligence · 2020

    This paper proposes a new architecture, named Gated Fully Fusion(GFF), to selectively fuse features from multiple levels using gates in a fully connected way, and achieves the state of the art results on four challenging scene parsing datasets including Cityscapes, Pascal Context, COCO-stuff and ADE20K.

    249
  • TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers

    Qianyu Zhou, Xiang-Tai Li, Lu He, Yi-Bo Yang, Guangliang Cheng, Yun-Hai Tong, Li-Zhuang Ma, Da-Cheng Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    TransVOD is presented, the first end-to-end video object detection system based on simple yet effective spatial-temporal Transformer architectures that streamline the pipeline of current VOD, effectively removing the need for many hand-crafted components for feature aggregation.

    242
  • Dual Graph Convolutional Network for Semantic Segmentation.

    Li Zhang, Xiang-Tai Li, Anurag Arnab, Kuiyuan Yang, Yun-Hai Tong, P. Torr

    arXiv · 2019

    The Dual Graph Convolutional Network (DGCNet) models the global context of the input feature by modelling two orthogonal graphs in a single framework, which achieves state-of-the-art results on both Cityscapes and Pascal Context datasets.

    196
  • Neural Collapse Inspired Feature-Classifier Alignment for Few-Shot Class Incremental Learning

    Yibo Yang, Haobo Yuan, Xiang-Tai Li, Zhou-Chen Lin, P. Torr, Da-Cheng Tao

    arXiv · 2023

    A neural collapse inspired framework for FSCIL that holds the neural collapse optimality and does not break the feature-classifier alignment in an incremental fashion and experiments demonstrate that the proposed framework outperforms the state-of-the-art performances.

    191
  • SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model

    Zhenglin Huang, Jinwei Hu, Xiang-Tai Li, Yiwei He, Xing-Yu Zhao, Bei Peng, Baoyuan Wu, Xiaowei Huang, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This paper proposes a new image deepfake detection, localization, and explanation framework, named SIDA (Social media Image Detection, localization, and explanation Assistant), which not only discerns the authenticity of images, but also delineates tampered regions through mask prediction and provides textual explanations of the model’s judgment criteria.

    176
  • Sa2VA: Marrying SAM2 With MLLM for Dense Grounded Understanding of Images and Videos

    Haobo Yuan, Xiang-Tai Li, Tao Zhang, Yue-Yi Sun, Zi-Long Huang, Shilin Xu, Shun-Ping Ji, Yun-Hai Tong, +3 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2026

    Experiments show that Sa2VA achieves strong performance across multiple tasks, particularly in referring video object segmentation, highlighting its potential for complex real-world applications.

    168
  • CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

    Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Xiang-Tai Li, Wen-Tao Liu, C. Loy

    arXiv · 2023

    An in-depth analysis of the region-language alignment in CLIP models is embarked on, which proposes an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs.

    159
  • OMG-Seg: Is One Model Good Enough for all Segmentation?

    Xiang-Tai Li, Haobo Yuan, Wei Li, Henghui Ding, Size Wu, Wenwei Zhang, Yining Li, Kai Chen, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    It is shown that OMG-Seg, a transformer-based encoder-decoder architecture with task-specific queries and outputs, can support over ten distinct segmentation tasks and yet significantly reduce computational and parameter overhead across various tasks and datasets.

    140
  • Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation

    Xiang-Tai Li, Wenwei Zhang, Jiangmiao Pang, Kai Chen, Guangliang Cheng, Yun-Hai Tong, Chen Change Loy

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    Video K-Net is presented, a simple, strong, and unified framework for fully end-to-end video panoptic seg-mentation that achieves state-of-the-art videoPanoptic segmentation results on Citscapes-VPS and KITTI-STEP without bells and whistles and can serve as a new flexible baseline in video segmentation.

    125
  • Enhanced Boundary Learning for Glass-like Object Segmentation

    Hao He, Xiang-Tai Li, Guangliang Cheng, Jian-Ping Shi, Yun-Hai Tong, Gao-Feng Meng, V. Prinet, Lubin Weng

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    This paper proposes a novel refined differential module that outputs finer boundary cues and introduces an edge-aware point-based graph convolution network module to model the global shape along the boundary.

    123
  • End-to-End Video Object Detection with Spatial-Temporal Transformers

    Lu He, Qianyu Zhou, Xiang-Tai Li, Li Niu, Guangliang Cheng, Xiao Li, Wen-Xuan Liu, Yu Tong, +2 more

    ACM International Conference on Multimedia (ACM MM) · 2021

    TransVOD is presented, an end-to-end video object detection model based on a spatial-temporal Transformer architecture that streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks.

    116
  • Open-Vocabulary SAM: Segment and Recognize Twenty-Thousand Classes Interactively

    Haobo Yuan, Xiang-Tai Li, Chong Zhou, Yining Li, Kai Chen, C. Loy

    Lecture notes in computer science · 2024

    The Open-Vocabulary SAM is introduced, a SAM-inspired model designed for simultaneous interactive segmentation and recognition, leveraging two unique knowledge transfer modules: SAM2CLIP and CLIP2SAM.

    114
  • PointFlow: Flowing Semantics Through Points for Aerial Image Segmentation

    Xiang-Tai Li, Hao He, Xia Li, Duo Li, Guangliang Cheng, Jian-Ping Shi, Lubin Weng, Yun-Hai Tong, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    A point-wise affinity propagation module based on the Feature Pyramid Network (FPN) framework, named PointFlow, is proposed, which reduces the noise introduced by the background while keeping efficiency.

    111
  • RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation

    Peng Lu, Tao Jiang, Yining Li, Xiang-Tai Li, Kai Chen, Wen-Ming Yang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    RTMO is introduced, a one-stage pose estimation framework that seamlessly inte-grates coordinate classification by representing keypoints using dual I-D heatmaps within the YOLO architecture, achieving accuracy comparable to top-down methods while maintaining high speed.

    103
  • Multi-Task Learning With Multi-Query Transformer for Dense Prediction

    Yang-Yang Xu, Xiang-Tai Li, Haobo Yuan, Yi-Bo Yang, Jing Zhang, Yun-Hai Tong, Le-Fei Zhang, Da-Cheng Tao

    IEEE Transactions on Circuits and Systems for Video Technology · 2023

    This work proposes a simple pipeline named Multi-Query Transformer (MQTransformer) that is equipped with multiple queries from different tasks to facilitate the reasoning among multiple tasks and simplify the cross-task interaction pipeline.

    94
  • An Open and Comprehensive Pipeline for Unified Object Grounding and Detection

    Xiangyu Zhao, Yicheng Chen, Shilin Xu, Xiang-Tai Li, Xinjiang Wang, Yining Li, Haian Huang

    arXiv · 2024

    MM-Grounding-DINO is presented, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox, and outperforms the Grounding-DINO-Tiny baseline.

    93
  • Point Cloud Mamba: Point Cloud Learning via State Space Model

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    86
  • Toward Robust Referring Image Segmentation

    Jianzong Wu, Xiang-Tai Li, Xia Li, Heng-Hui Ding, Yun-Hai Tong, Dacheng Tao

    IEEE Transactions on Image Processing · 2024

    83
  • Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection

    Jiang-Ning Zhang, Xuhai Chen, Yabiao Wang, Chengjie Wang, Yong Liu, Xiang-Tai Li, Ming-Hsuan Yang, Da-Cheng Tao

    arXiv · 2023

    A novel ViT-based ViTAD structure is instantiated, designed incrementally from both global and local perspectives, and achieves state-of-the-art results and efficiency on MVTec AD, VisA, and Uni-Medical datasets.

    73
  • Panoptic Video Scene Graph Generation

    Jingkang Yang, Wen-Hsiao Peng, Xiang-Tai Li, Zu-Jin Guo, Liangyu Chen, Bo Li, Zheng Ma, Kaiyang Zhou, +3 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    A high-quality PVSG dataset is contributed, which consists of 400 videos (289 third-person + 111 egocentric videos) with totally 150K frames labeled with panoptic segmentation masks as well as fine, temporal scene graphs.

    73
  • 72
  • EdgeSAM: Prompt-In-the-Loop Distillation for SAM

    Chong Zhou, Xiang-Tai Li, C. Loy, Bo Dai

    International Journal of Computer Vision · 2025

    The approach involves distilling the original ViT-based SAM image encoder into a purely CNN-based architecture, better suited for edge devices, and incorporates a lightweight module within the encoder, which achieves a 37-fold speed increase compared to the original SAM.

    71
  • Towards Semantic Equivalence of Tokenization in Multimodal LLM

    Sheng-Qiong Wu, Hao Fei, Xiang-Tai Li, Jia-Yi Ji, Hanwang Zhang, Tat-Seng Chua, Shui-Cheng Yan

    arXiv · 2024

    A novel dynamic Semantic-Equivalent Vision Tokenizer (SeTok) is proposed, which groups visual features into semantic units via a dynamic clustering algorithm, flexibly determining the number of tokens based on image complexity.

    70
  • Sfnet: Faster and Accurate Semantic Segmentation Via Semantic Flow

    Xiang-Tai Li, Jiang-Ning Zhang, Yi-Bo Yang, Guangliang Cheng, Kuiyuan Yang, Yu Tong, Da-Cheng Tao

    International Journal of Computer Vision · 2023

    A Flow Alignment Module (FAM) is proposed to learn Semantic Flow between feature maps of adjacent levels and broadcast high-level features to high-resolution features effectively and efficiently and integrating this FAM to a standard feature pyramid structure exhibits superior performance over other real-time methods.

    70
  • Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

    Hao-Cheng Wang, Xiang-Tai Li, Zi-Long Huang, An-Ran Wang, Jia-Cong Wang, Tao Zhang, Jia-Ni Zheng, Sule Bai, +4 more

    arXiv · 2025

    TreeBench (Traceable Evidence Evaluation Benchmark), a diagnostic benchmark built on three principles: focused visual perception of subtle targets in complex scenes, traceable evidence via bounding box evaluation, and second-order reasoning to test object interactions and spatial hierarchies beyond simple object localization are proposed.

    66
  • Towards Robust Referring Image Segmentation

    Jianzong Wu, Xiang-Tai Li, Xia Li, Henghui Ding, Yu Tong, Da-Cheng Tao

    arXiv · 2022

    This work proposes a new formulation of RIS, named Robust Referring Image Segmentation (R-RIS), which considers the negative sentence inputs besides the regular positive text inputs, and proposes a new transformer-based model, called RefSegformer, with a token-based vision and language fusion module.

    66
  • MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation

    Jiahao Xie, Wei Li, Xiang-Tai Li, Zi-Wei Liu, Y. Ong, Chen Change Loy

    International Journal of Computer Vision · 2024

    Experimental results on the challenging LVIS long-tailed and open-vocabulary benchmarks demonstrate that MosaicFusion can significantly improve the performance of existing instance segmentation models, especially for rare and novel categories.

    61
  • UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

    Zhu-Cun Xue, Jiang-Ning Zhang, Teng Hu, Hao-Yang He, Yinan Chen, Yuxuan Cai, Yabiao Wang, Chengjie Wang, +3 more

    arXiv · 2025

    This paper proposes a high-quality open-sourced UHD-4K text-to-video dataset named UltraVideo, and expands Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of the data curation.

    59
  • Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

    Jinbin Bai, Tian Ye, Wei Chow, En-Xin Song, Qing-Guo Chen, Xiang-Tai Li, Zhen Dong, Lei Zhu, +1 more

    arXiv · 2024

    59
  • EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm

    Jiang-Ning Zhang, Xiang-Tai Li, Yabiao Wang, Chengjie Wang, Yi-Bo Yang, Yong Liu, Da-Cheng Tao

    International Journal of Computer Vision · 2024

    A novel pyramid EATFormer backbone that only contains the proposed EA-based transformer (EAT) block, which consists of three residual parts to model multi-scale, interactive, and individual information separately and design a task-related head docked with transformer backbone to complete final information fusion more flexibly.

    59
  • Exploring plain ViT features for multi-class unsupervised visual anomaly detection

    Jiang-Ning Zhang, Xuhai Chen, Yabiao Wang, Chengjie Wang, Yong Liu, Xiang-Tai Li, Ming-Hsuan Yang, Da-Cheng Tao

    Computer Vision and Image Understanding · 2025

    51
  • Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook

    Mu-Yi Bao, Shuchang Lyu, Zhaoyang Xu, Huiyu Zhou, Jin-Chang Ren, Shiming Xiang, Xiang-Tai Li, Guang-Liang Cheng

    arXiv · 2025

    This survey presents a comprehensive review of Mamba-based methodologies in remote sensing, systematically analyzing about 120 Mamba-based remote sensing studies to construct a holistic taxonomy of innovations and applications, and establishes Mamba as a transformative framework for remote sensing analysis.

    51
  • PolyphonicFormer: Unified Query Learning for Depth-Aware Video Panoptic Segmentation

    Haobo Yuan, Xiang-Tai Li, Yi-Bo Yang, Guangliang Cheng, Jing Zhang, Yun-Hai Tong, Le-Fei Zhang, D. Tao

    Lecture notes in computer science · 2022

    PolyphonicFormer, a vision transformer to unify these sub-tasks under the DVPS task and lead to more robust results, achieves state-of-the-art results on two DVPS datasets, and ranks 1st on the ICCV-2021 BMTT Challenge video + depth track.

    51
  • Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence

    Jia-Hao Meng, Xiang-Tai Li, Hao-Cheng Wang, Yue Tan, Tao Zhang, Lingdong Kong, Yun-Hai Tong, An-Ran Wang, +3 more

    arXiv · 2025

    This work introduces Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable.

    50
  • Towards Language-Driven Video Inpainting via Multimodal Large Language Models

    Jianzong Wu, Xiang-Tai Li, Chenyang Si, Shang-Chen Zhou, Jingkang Yang, Jiang-Ning Zhang, Yining Li, Kai Chen, +3 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work introduces a new task - language-driven video inpainting, which uses natural language instructions to guide the inpainting process, integrating Multimodal Large Language Models to understand and execute complex language-based inpaintingrequests effectively.

    48
  • Referring Image Editing: Object-Level Image Editing via Referring Expressions

    Chang Liu, Xiang-Tai Li, Henghui Ding

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    47
  • Explore In-Context Learning for 3D Point Cloud Understanding

    Zhongbin Fang, Xiang-Tai Li, Xia Li, J. Buhmann, Chen Change Loy, Mengyuan Liu

    arXiv · 2023

    This work introduces a novel framework, named Point-In-Context, designed especially for in-context learning in 3D point clouds, where both inputs and outputs are modeled as coordinates for each task.

    47
  • 45
  • Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

    Pei-Wen Sun, Si-Tong Cheng, Xiang-Tai Li, Zhen Ye, Hua-Dai Liu, Hong-Gang Zhang, Wei Xue, Yi-Ke Guo

    arXiv · 2024

    The SpatialSonic model, a spatial-aware encoders and azimuth state matrices model utilizing spatial-aware encoders and azimuth state matrices, is introduced utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.

    44
  • 41
  • Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation

    Jianzong Wu, Xiang-Tai Li, Henghui Ding, Xia Li, Guangliang Cheng, Yu Tong, Chen Change Loy

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    This work devise a joint Caption Grounding and Generation (CGG) framework, which incorporates a novel grounding loss that only focuses on matching object nouns to improve learning efficiency and introduces a caption generation head that enables additional supervision and contextual modeling as a complementation to the grounding loss.

    40
  • NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results

    Xin Li, Ye-Ying Jin, Xin Jin, Zongwei Wu, Bing-Chen Li, Yu-Fei Wang, Wenhan Yang, Yu Li, +22 more

    IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2025

    This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images to establish a new and powerful benchmark for the task of removing raindrops under varying lighting and focus conditions.

    37
  • Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation

    Xiang-Tai Li, Haobo Yuan, Wenwei Zhang, Guangliang Cheng, Jiangmiao Pang, C. Loy

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    Tube-Link is a near-online approach that takes a short subclip as input and outputs the corresponding spatial-temporal tube masks, and introduces temporal contrastive learning to instance-wise discriminative features for tube-level association.

    37
  • Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model

    Kuan-Chih Huang, Xiang-Tai Li, Lu Qi, Shui-Cheng Yan, Ming-Hsuan Yang

    International Conference on 3D Vision (3DV) · 2025

    Reason3D is introduced, a novel LLM designed for comprehensive 3D understanding that processes point cloud data and text prompts to produce textual responses and segmentation masks, enabling advanced tasks such as 3D reasoning segmentation, hierarchical searching, express referring, and question answering with detailed mask outputs.

    36
  • Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation

    Jinghao Wang, Zhengyu Wen, Xiang-Tai Li, Zu-Jin Guo, Jingkang Yang, Ziwei Liu

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    A novel framework is presented: Pair then Relation (Pair-Net), which uses a Pair Proposal Network (PPN) to learn and filter sparse pair-wise relationships between subjects and objects and achieves over 10% absolute gains compared to the baseline, PSGFormer.

    36
  • Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning

    Xinshun Wang, Zhongbin Fang, Xia Li, Xiang-Tai Li, Chen Chen, Mengyuan Liu

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work proposes Skeleton-in-Context (SiC), an effective framework for in-context skeleton sequence modeling that is able to handle multiple skeleton-based tasks simultaneously after a single training process and accomplish each task from context according to the given prompt.

    35
  • The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer

    Weixian Lei, Jiacong Wang, Hao-Cheng Wang, Xiang-Tai Li, J. Liew, Jia-Shi Feng, Zi-Long Huang

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    Sail is introduced, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture that eliminates the need for a separate vision encoder, presenting a more minimalist architecture design.

    34
  • 34
  • 4D Panoptic Scene Graph Generation

    Jingkang Yang, Jun Cen, Wen-Hsiao Peng, Shuai Liu, Fangzhou Hong, Xiang-Tai Li, Kaiyang Zhou, Qi-Feng Chen, +1 more

    arXiv · 2024

    To solve PSG-4D, the introduction of 4D Panoptic Scene Graph (PSG-4D), a new representation that bridges the raw visual data perceived in a dynamic 4D world and high-level visual understanding, and a real-world application example to demonstrate how the model can achieve dynamic scene understanding by integrating a large language model into the system.

    33
  • DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries

    Yi-Kang Zhou, Tao Zhang, Shun-Ping Ji, Shui-Cheng Yan, Xiang-Tai Li

    arXiv · 2024

    This work introduces Dynamic Anchor Queries (DAQ) to shorten the transition gap between the anchor and target queries by dynamically generating anchor queries based on the features of potential candidates and introduces a query-level object Emergence and Disappearance Simulation (EDS) strategy, which unleashes DAQ's potential without any additional cost.

    33
  • 33
  • 33
  • Electrocatalysis of oxygen reduction on carbon nanotubes with different surface functional groups in acid and alkaline solutions

    Hui-Juan Zhang, Haoliang Li, Xiang-Tai Li, Bin Zhao, Jun-He Yang

    International Journal of Hydrogen Energy · 2014

    33
  • BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

    Hai-Quan Wen, Yiwei He, Zhenglin Huang, Tianxiao Li, Zihan Yu, Xingru Huang, Lu Qi, Baoyuan Wu, +2 more

    arXiv · 2025

    This work introduces a comprehensive dataset, benchmark, and baseline model for video forgery detection, and develops BusterX, an MLLM baseline with RL training that outperforms several leading MLLMs in both detection accuracy and rationale quality.

    32
  • HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

    Jinbin Bai, Wei Chow, Ling Yang, Xiang-Tai Li, Juncheng Li, Hanwang Zhang, Shui-Cheng Yan

    arXiv · 2024

    High-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions, set a new versatile benchmark for instructional image editing datasets.

    32
  • DGMamba: Domain Generalization via Generalized State Space Model

    Shao-Cong Long, Qian-Yu Zhou, Xiang-Tai Li, Xue-Quan Lu, Chen-Hao Ying, Yuan Luo, Li-Zhuang Ma, Shui-Cheng Yan

    ACM International Conference on Multimedia (ACM MM) · 2024

    This paper proposes a novel framework for DG, named DGMamba, that excels in strong generalizability toward unseen domains and meanwhile has the advantages of global receptive fields, and efficient linear complexity.

    32
  • An Empirical Study of GPT-4o Image Generation Capabilities

    Si-Xiang Chen, Jinbin Bai, Zhuo-Ran Zhao, Tian Ye, Qingyu Shi, Dong-Hao Zhou, Wenhao Chai, Xin Lin, +11 more

    arXiv · 2025

    An empirical study of GPT-4o's image generation capabilities is conducted, benchmarking it against leading open-source and commercial models and identifying promising directions for future unified generative models, emphasizing the role of architectural design and data scaling.

    31
  • Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark

    Jiang-Ning Zhang, Chengjie Wang, Xiang-Tai Li, Guan-Zhong Tian, Zhu-Cun Xue, Yong Liu, Guan-Song Pang, Da-Cheng Tao

    International Journal of Computer Vision · 2026

    A simple but more powerful InvAD framework to achieve high-quality feature reconstruction and improves the effectiveness of reconstruction-based methods on popular MVTec AD, VisA, and the newly proposed COCO-AD datasets under a multi-class unsupervised setting.

    29
  • WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

    Ao Liang, Ling-Dong Kong, T. Yan, Hong-Si Liu, Wesley Yang, Zi-Qi Huang, Wei Yin, Jia-Long Zuo, +14 more

    arXiv · 2025

    This work introduces WorldLens, a full-spectrum benchmark evaluating how well a model builds, understands, and behaves within its generated world, and develops WorldLens-Agent, an evaluation model distilled from these annotations to enable scalable, explainable scoring.

    29
  • On Path to Multimodal Generalist: General-Level and General-Bench

    Hao Fei, Yuan Zhou, Juncheng Li, Xiang-Tai Li, Qingshan Xu, Bobo Li, Sheng-Qiong Wu, Yaoting Wang, +22 more

    arXiv · 2025

    This project introduces General-Level, an evaluation framework that defines 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI.

    29
  • Panoptic-PartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation

    Xiang-Tai Li, Shilin Xu, Yi-Bo Yang, Haobo Yuan, Guangliang Cheng, Yu Tong, Zhou-Chen Lin, Ming-Hsuan Yang, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    The first end-to-end unified framework, Panoptic-PartFormer is designed, designing a meta-architecture that decouples part features and things/stuff features, respectively and proposes a new metric Part-Whole Quality (PWQ), better to measure this task from pixel-region and part-whole perspectives.

    29
  • DST-Det: Open-Vocabulary Object Detection via Dynamic Self-Training

    Shilin Xu, Xiang-Tai Li, Size Wu, Wenwei Zhang, Yun-Hai Tong, C. Loy

    IEEE Transactions on Circuits and Systems for Video Technology · 2024

    29
  • PixelThink: Towards Efficient Chain-of-Pixel Reasoning

    Song Wang, Gongfan Fang, Lingdong Kong, Xiang-Tai Li, Jian-Yun Xu, Sheng Yang, Qiang Li, Jianke Zhu, +1 more

    arXiv · 2025

    28
  • Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation

    Yue Han, Jiang-Ning Zhang, Zhu-Cun Xue, Chao Xu, Xin-Tian Shen, Yabiao Wang, Cheng-Jie Wang, Yong Liu, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    A unified framework, Reference Twice (RefT), is introduced to exploit the relationship between support and query features for FSIS and related tasks, demonstrating that support object queries encode key factors after base training, allowing query features to be enhanced twice at both feature and query levels.

    28
  • PointRWKV: Efficient RWKV-Like Model for Hierarchical Point Cloud Learning

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    27
  • DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

    Zi-Tong Wang, Hang Zhao, Qianyu Zhou, Xue-Quan Lu, Xiang-Tai Li, Yi-Ren Song

    arXiv · 2025

    A novel task: Layer-Wise Decomposition of Alpha-Composited Images, aiming to recover constituent layers from single overlapped images under the condition of semi-transparent/transparent alpha layer non-linear occlusion, and presents DiffDecompose, a diffusion Transformer-based framework that learns the posterior over possible layer decompositions conditioned on the input image, semantic prompts, and blending type.

    26
  • LLAVADI: What Matters For Multimodal Large Language Models Distillation

    Shilin Xu, Xiang-Tai Li, Haobo Yuan, Lu Qi, Yun-Hai Tong, Ming-Hsuan Yang

    arXiv · 2024

    Results show that joint alignment for both tokens and logit alignment plays critical roles in teacher-student frameworks and even a 2.7B small-scale model can perform on par with larger models with 7B or 13B parameters.

    26
  • Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

    Tao Zhang, Xiang-Tai Li, Zi-Long Huang, Yanwei Li, Weixian Lei, Xueqing Deng, Shihao Chen, Shun-Ping Ji, +1 more

    arXiv · 2025

    The Pixel-SAIL, a single transformer for pixel-wise MLLM tasks, is presented and a novel visual prompt injection strategy to enable the single transformer to understand visual prompt inputs and benefit from the early fusion of visual prompt embeddings and vision tokens is proposed.

    25
  • BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model

    Yiran Song, Qian-Yu Zhou, Xiang-Tai Li, Deng-Ping Fan, Xue-Quan Lu, Li-Zhuang Ma

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    A Scalable Bias-Mode Attention Mask (BA-SAM) is proposed to enhance SAM's adaptability to varying image resolutions while eliminating the need for structure modifications, and a new scaling factor is introduced to ensure consistent magnitude in the attention layer's dot product values when the token sequence length changes.

    25
  • Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model

    Haobo Yuan, Xiang-Tai Li, Lu Qi, Tao Zhang, Ming-Hsuan Yang, Shui-Cheng Yan, C. Loy

    arXiv · 2024

    This work designs a mixed backbone that contains convolution and RWKV operation, which achieves the best for both accuracy and efficiency and designs an efficient decoder to utilize the multiscale tokens to obtain high-quality masks.

    25
  • DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

    Jianzong Wu, Chao Tang, Jingbo Wang, Yanhong Zeng, Xiang-Tai Li, Yun-Hai Tong

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work proposes a new task: customized manga generation and introduces DiffSensei, an innovative framework specifically designed for generating manga with dynamic multi-character control, marking a significant advancement in manga generation by enabling text- adaptable character customization.

    23
  • Global Aggregation Then Local Distribution for Scene Parsing

    Xiang-Tai Li, Li Zhang, Guangliang Cheng, Kuiyuan Yang, Yun-Hai Tong, Xia-Tian Zhu, T. Xiang

    IEEE Transactions on Image Processing · 2021

    A novel local distribution module is designed which models the affinity map between global and local relationship for each pixel adaptively and can be modularized as an end-to-end trainable block and easily plugged into existing semantic segmentation networks, giving rise to the GALD networks.

    23
  • Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models

    Shilin Xu, Yanwei Li, Rui Yang, Tao Zhang, Yue-Yi Sun, Wei Chow, Linfeng Li, Hang Song, +4 more

    arXiv · 2025

    A unified yet straightforward framework that contains a mixed reward function design (Mixed-Reward) and a mixed post-training dataset (Mixed-45K) and a new open-ended reward named Bidirectional Max-Average Similarity (BMAS) by leveraging tokenizer embedding matching between the generated response and the ground truth.

    20
  • DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency

    Meng-Shi Qi, Pengfei Zhu, Xiang-Tai Li, Xiaoyang Bi, Lu Qi, Hua-Dong Ma, Ming-Hsuan Yang

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025

    This work proposes a new Dual Consistency SAM (DC-SAM), a prompt-tuning framework that adapts SAM and SAM2 for both image and video in-context segmentation and introduces a novel cycle-consistent cross-attention mechanism to enforce alignment between fused features and visual prompts.

    19
  • Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

    Yi-Kang Zhou, Tao Zhang, Shilin Xu, Shihao Chen, Qian-Yu Zhou, Yun-Hai Tong, Shun-Ping Ji, Jiang-Ning Zhang, +2 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    19
  • AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

    Zhu-Cun Xue, Jiang-Ning Zhang, Xurong Xie, Yuxuan Cai, Yong Liu, Xiang-Tai Li, Da-Cheng Tao

    arXiv · 2025

    Experiments show that AdaVideoRAG significantly improves both efficiency and accuracy on long-video QA tasks and can be seamlessly plugged into existing MLLMs through lightweight APIs, establishing a new paradigm for adaptive retrieval-augmented video analysis.

    19
  • 19
  • Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

    Hao-Cheng Wang, Yuhao Wang, Tao Zhang, Yi-Kang Zhou, Yanwei Li, Jia-Cong Wang, Jia-Ni Zheng, Ye Tian, +8 more

    arXiv · 2025

    This work introduces Grasp Any Region (GAR) for region-level visual understanding, and constructs GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions.

    18
  • Neural Collapse Terminus: A Unified Solution for Class Incremental Learning and Its Variants

    Yi-Bo Yang, Haobo Yuan, Xiang-Tai Li, Jian-Long Wu, Le-Fei Zhang, Zhou-Chen Lin, P. Torr, Dacheng Tao, +1 more

    arXiv · 2023

    A unified solution to the misalignment dilemma in the three tasks of CIL, LTCIL and FSCIL is offered andoretical analysis indicates that the method holds the neural collapse optimality in an incremental fashion regardless of data imbalance or data scarcity.

    18
  • D$^2$GS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction

    Meixi Song, Xin Lin, Dizhe Zhang, Haodong Li, Xiang-Tai Li, Bo Du, Lu Qi

    arXiv · 2025

    A unified framework D$^2$GS is proposed, comprising a Depth-and-Density Guided Dropout strategy that suppresses overfitting by adaptively masking redundant Gaussians based on density and depth, and a Distance-Aware Fidelity Enhancement module that improves reconstruction quality in under-fitted far-field areas through targeted supervision.

    17
  • 17
  • VG4D: Vision-Language Model Goes 4D Video Recognition

    Zhichao Deng, Xiang-Tai Li, Xia Li, Yun-Hai Tong, Shen Zhao, Mengyuan Liu

    Proceedings - IEEE International Conference on Robotics and Automation/Proceedings · 2024

    This work proposes the Vision-Language Models Goes 4D (VG4D) framework, a framework to transfer VLM knowledge from visual-text pretrained models to a 4D point cloud network, and achieves improved recognition performance.

    17
  • Improving Video Instance Segmentation via Temporal Pyramid Routing

    Xiang-Tai Li, Hao He, Yi-Bo Yang, Henghui Ding, Kuiyuan Yang, Guangliang Cheng, Yun-Hai Tong, Da-Cheng Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    A Temporal Pyramid Routing strategy to conditionally align and conduct pixel-level aggregation from a feature pyramid pair of two adjacent frames is proposed, which is a light-weight and plug-and-play module and can be easily applied to existing instance segmentation methods.

    17
  • Rethinking Evaluation Metrics of Open-Vocabulary Segmentaion

    Hao Zhou, Tian-Cheng Shen, Xu Yang, Hai Huang, Xiang-Tai Li, Lu Qi, Ming-Hsuan Yang

    arXiv · 2023

    Novel evaluation metrics, namely Open mIoU, Open AP, and Open PQ, tailored for three open-vocabulary segmentation tasks are designed, which demonstrate that the relative subjectivity of similarity distance can still well evaluate the open ability of the existing open- vocabulary segmentation methods.

    16
  • Improving Video Segmentation via Dynamic Anchor Queries

    Lecture notes in computer science · 2024

    15
  • Unified Dense Prediction of Video Diffusion

    Lehan Yang, Lu Qi, Xiang-Tai Li, Sheng Li, Varun Jampani, Ming-Hsuan Yang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    A unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts and incorporates learnable task embeddings brings multiple dense prediction tasks into a single model, enhancing flexibility and further boosting performance.

    14
  • MEDIC: Zero-shot Music Editing with Disentangled Inversion Control

    Hua-Dai Liu, Jia-Lei Wang, Xiang-Tai Li, Wen Wang, Qian Chen, Rongjie Huang, Yang Liu, Jia-Yang Xu, +1 more

    arXiv · 2024

    MEDIC, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion, outperforms state-of-the-art inversion techniques in editing fidelity and content preservation.

    14
  • Convolution-Enhanced Evolving Attention Networks

    Yu-Jing Wang, Ya-Ming Yang, Zhuowan Li, Jian-Gang Bai, Ming-Liang Zhang, Xiang-Tai Li, J. Yu, Ce Zhang, +2 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2023

    This work proposes a novel and generic evolving attention mechanism, which directly models the evolution of inter-token relationships through a chain of residual convolutional modules, and is the first work that explicitly models the layer-wise evolution of attention maps.

    14
  • A Generalist FaceX via Learning Unified Facial Representation

    Yue Han, Jiang-Ning Zhang, Jun-Wei Zhu, Xiang-Tai Li, Yanhao Ge, Wei Li, Cheng-Jie Wang, Yong Liu, +2 more

    arXiv · 2023

    This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously, and introduces Facial Omni-Representation Decomposing (FORD) for seamless manipulation of various facial components, microscopically decomposing the core aspects of most facial editing tasks.

    14
  • RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection

    Jingtong Yue, Zhiwei Lin, Xin Lin, Xiao-Yu Zhou, Xiang-Tai Li, Lu Qi, Yongtao Wang, Ming-Hsuan Yang

    arXiv · 2025

    This work designs a 3D Gaussian Expansion module to mitigate inaccuracies in radar points, including position, Radar Cross-Section (RCS), and velocity, and introduces a weather-adaptive fusion module, which adaptively fuses radar and camera features based on camera signal confidence.

    13
  • From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond

    Hao Fei, Xiang-Tai Li, Haotian Liu, Fuxiao Liu, Zhuosheng Zhang, Hanwang Zhang, Shui-Cheng Yan

    ACM International Conference on Multimedia (ACM MM) · 2024

    This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs, focusing on three key areas: MLLM architecture design, instructional learning, and multimodal reasoning of MLLMs.

    13
  • DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection

    Shilin Xu, Xiang-Tai Li, Size Wu, Wenwei Zhang, Yining Li, Guangliang Cheng, Yun-Hai Tong, Kai Chen, +1 more

    arXiv · 2023

    13
  • SAMTok: Representing Any Mask with Two Words

    Yi-Kang Zhou, Tao Zhang, Dengxian Gong, Yuanzheng Wu, Ye Tian, Hao-Cheng Wang, Haobo Yuan, Jia-Cong Wang, +8 more

    arXiv · 2026

    A discrete mask tokenizer that converts any region mask into two special tokens and reconstructs the mask using these tokens with high fidelity, allowing base MLLMs to learn pixel-wise capabilities through standard next-token prediction and simple reinforcement learning, without architectural modifications and specialized loss design.

    12
  • 12
  • RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything

    Shilin Xu, Haobo Yuan, Qingyu Shi, Lu Qi, Jingbo Wang, Yi-Bo Yang, Yi-Ning Li, Kai Chen, +4 more

    arXiv · 2024

    This work explores a new real-time segmentation setting, named all-purpose segmentation in real-time, to transfer VFMs in real-time deployment and presents Real-Time All Purpose SAM (RAP-SAM), which contains an efficient encoder and an efficient decoupled decoder to perform prompt-driven decoding.

    12
  • Open-Vocabulary SAM3D: Towards Training-Free Open-Vocabulary 3D Scene Understanding

    Hanchen Tai, Qingdong He, Jiang-Ning Zhang, Yi-Jie Qian, Zhenyu Zhang, Xiao-Bin Hu, Xiang-Tai Li, Yabiao Wang, +1 more

    IEEE Transactions on Circuits and Systems for Video Technology · 2026

    This paper introduces OV-SAM3D, a training-free method that contains a universal framework for understanding open-vocabulary 3D scenes without requiring prior knowledge of the scene, and surpasses existing open-vocabulary methods in unknown open-world environments.

    11
  • DreamRelation: Bridging Customization and Relation Generation

    Qingyu Shi, Lu Qi, Jianzong Wu, Jinbin Bai, Jingbo Wang, Yun-Hai Tong, Xiang-Tai Li

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work introduces DreamRelation, a framework that disentangles identity and relation learning using a carefully curated dataset and proposes two key modules to tackle the two main challenges of generating accurate and natural relationships, especially when significant pose adjustments are required, and avoiding object confusion in cases of overlap.

    11
  • Exploring Self-Supervised Learning for Multi-Modal Remote Sensing Pre-Training via Asymmetric Attention Fusion

    Guozheng Xu, Xue Jiang, Xiang-Tai Li, Ze Zhang, Xing-Zhao Liu

    Remote Sensing · 2023

    The proposed Asymmetric Attention Fusion framework achieves an improvement of over 7% in all metrics compared to randomly initialized methods for both tasks, and when compared to early fusion and late fusion methods, AAF consistently outperforms in achieving superior improvements.

    11
  • Fast and Accurate Scene Parsing via Bi-Direction Alignment Networks

    Yanran Wu, Xiang-Tai Li, Chen Shi, Yun-Hai Tong, Hua Yang, Tao Song, Ruhui Ma, Hai-Bing Guan

    Proceedings - International Conference on Image Processing · 2021

    An effective method for fast and accurate scene parsing called Bidirectional Alignment Network (BiAlignNet) is proposed by aligning two-path information into each other through a learned flow field to avoid the noise and semantic gaps.

    11
  • Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark

    Haobo Yuan, Yue-Yi Sun, Yanwei Li, Tao Zhang, Xueqing Deng, Henghui Ding, Lu Qi, An-Ran Wang, +2 more

    arXiv · 2025

    The Visual Reasoning Tracer task is introduced, which requires models to not only localize the target object but also explicitly predict the intermediate objects that form the reasoning path, and reveals that while existing models often produce the correct final output, they struggle to ground their intermediate reasoning.

    10
  • A Masked Reference Token Supervision-Based Iterative Visual-Language Framework for Robust Visual Grounding

    Chun-Lei Wang, Wen-Quan Feng, Shuchang Lyu, Guang-Liang Cheng, Xiang-Tai Li, Binghao Liu, Qi Zhao

    IEEE Transactions on Circuits and Systems for Video Technology · 2024

    10
  • Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene

    Sheng-Qiong Wu, Hao Fei, Jing-Kang Yang, Xiang-Tai Li, Juncheng Li, Hanwang Zhang, Tat-Seng Chua

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    A 4D Large Language Model (4D-LLM) integrated with a 3D mask decoder for end-to-end generation of 4D-PSG and a 2D-to-4D visual scene transfer learning framework, where a spatial-temporal scene transcending strategy effectively transfers dimension-invariant features from abundant 2D SG annotations to 4D scenes, effectively compensating for data scarcity in 4D-PSG.

    9
  • MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation

    Ye Tian, Ling Yang, Jiong Yang, An-Ran Wang, Yu Tian, Jia-Ni Zheng, Hao-Cheng Wang, Zhiyang Teng, +5 more

    arXiv · 2025

    This work proposes a parallel multimodal diffusion framework, MMaDA-Parallel, that enables continuous, bidirectional interaction between text and images throughout the entire denoising trajectory, establishing a more robust paradigm for thinking-aware image synthesis.

    9
  • Rethinking Evaluation Metrics of Open-Vocabulary Segmentation

    Hao Zhou, Lu Qi, Tian-Cheng Shen, Hai Huang, Xu Yang, Xiang-Tai Li, Ming-Hsuan Yang

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025

    9
  • You Can't Ignore Either: Unifying Structure and Feature Denoising for Robust Graph Learning

    Tianmeng Yang, Jia-Hao Meng, Min Zhou, Ya-Ming Yang, Yu-Jing Wang, Xiang-Tai Li, Yun-Hai Tong

    ACM International Conference on Information and Knowledge Management (CIKM) · 2024

    This paper develops a unified graph denoising (UGD) framework to unravel the deadlock between structure and feature denoising, and proposes to refine noisy features with reconstruction based on a graph auto-encoder.

    9
  • Iterative Robust Visual Grounding with Masked Reference based Centerpoint Supervision

    Menghao Li, Chun-Lei Wang, W. Feng, Shuchang Lyu, Guangliang Cheng, Xiang-Tai Li, Binghao Liu, Qi Zhao

    IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) · 2023

    An Iterative Robust Visual Grounding (IRVG) framework with Masked Reference based Centerpoint Supervision (MRCS) is proposed with iterative multi-level vision-language fusion (IMVF) for better alignment and a multi-stage false-alarm sensitive decoder (MFSD) is presented to prevent the generation of false-alarm objects when presented with inaccurate expressions.

    9
  • Dense360: Dense Understanding from Omnidirectional Panoramas

    Yi-Kang Zhou, Tao Zhang, Dizhe Zhang, Shun-Ping Ji, Xiang-Tai Li, Lu Qi

    arXiv · 2025

    This work introduces an omnidirectional panoramas dataset featuring a comprehensive suite of reliability-scored annotations and introduces Dense360-Bench, the first benchmark for evaluating MLLMs on omnidirectional captioning and grounding, establishing a comprehensive framework for advancing dense visual-language understanding in panoramic settings.

    8
  • Explore In-Context Segmentation via Latent Diffusion Models

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    8
  • Continuous sPatial‐temporal deformable image registration and 4D frame interpolation

    Xia Li, Runzhao Yang, Mu-Heng Li, Xiang-Tai Li, A. Lomax, J. Buhmann, Ye Zhang

    Medical Physics · 2025

    This work presents a new approach to DIR implementation that addresses the issue of uncertainty in discrete volumetric motion representation, which affects the reliability of subsequent contour propagation and dose accumulation procedures.

    8
  • 8
  • Point-In-Context: Understanding Point Cloud via In-Context Learning

    Mengyuan Liu, Zhongbin Fang, Xia Li, J. Buhmann, De-Heng Ye, Xiang-Tai Li, C. Loy

    International Journal of Computer Vision · 2026

    This paper introduces Point-In-Context (PIC), a pioneering framework for 3D point cloud understanding that leverages in-context learning with a standard transformer architecture that uniquely enables the execution of multiple tasks after a single, unified training phase, eliminating the need for fine-tuning.

    7
  • The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA

    Quan-Zhu Niu, Dengxian Gong, Shihao Chen, Tao Zhang, Yi-Kang Zhou, Haobo Yuan, Lu Qi, Xiang-Tai Li, +1 more

    arXiv · 2025

    This paper proposes Segmentation Augmented and Selective Averaged Sa2VA (SaSaSa2VA) to address two key bottlenecks that limit segmentation performance: sparse frame sampling and reliance on a single [SEG] token for an entire video.

    7
  • From Masks to Worlds: A Hitchhiker's Guide to World Models

    Jinbin Bai, Yu Lei, Hecong Wu, Yuchen Zhu, Shufan Li, Yi Xin, Xiang-Tai Li, Mo-Lei Tao, +2 more

    arXiv · 2025

    This book follows one clear road: from early masked models that unified representation learning across modalities, to unified architectures that share a single paradigm, then to interactive generative models that close the action-perception loop, and finally to memory-augmented systems that sustain consistent worlds over time.

    7
  • Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning

    Qingdong He, Xueqin Chen, Chao-Yi Wang, Yanjie Pan, Xiao-Bin Hu, Z. Gan, Yabiao Wang, Chengjie Wang, +2 more

    arXiv · 2025

    This work proposes Reason50K, a large-scale dataset specifically curated for training and evaluating hypothetical instruction reasoning image editing, along with ReasonBrain, a novel framework designed to reason over and execute implicit hypothetical instructions across diverse scenarios.

    7
  • CyberV: Cybernetics for Test-time Scaling in Video Understanding

    Jia-Hao Meng, Shuyang Sun, Yue Tan, Lu Qi, Yun-Hai Tong, Xiang-Tai Li, Longyin Wen

    arXiv · 2025

    This work proposes a novel framework inspired by cybernetic principles, redesigning video MLLMs as adaptive systems capable of self-monitoring, self-correction, and dynamic resource allocation during inference, and demonstrates consistent gains on general-purpose benchmarks, such as VideoMME and WorldSense.

    7
  • PVUW 2025 Challenge Report: Advances in Pixel-Level Understanding of Complex Videos in the Wild

    Henghui Ding, Chang Liu, Nikhila Ravi, Shuting He, Yunchao Wei, Song Bai, P. Torr, Kehuan Song, +22 more

    IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2025

    6
  • DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation

    Qingdong He, Jinlong Peng, Pengcheng Xu, Bo-Yuan Jiang, Xiao-Bin Hu, Donghao Luo, Yong Liu, Yabiao Wang, +3 more

    arXiv · 2024

    This work proposes a novel framework, DynamicControl, which supports dynamic combinations of diverse control signals, allowing adaptive selection of different numbers and types of conditions, and demonstrates its superiority over existing methods in terms of controllability, generation quality and composability under various conditional controls.

    6
  • Is Your Driving World Model an All-Around Player?

    Ling-Dong Kong, Ao Liang, Tian-Yi Yan, Hong-Si Liu, Wesley Yang, Zi-Qi Huang, Xianbing Sun, Wei Yin, +15 more

    arXiv · 2026

    WorldLens is introduced, a unified benchmark that measures world-model fidelity across the full spectrum, from pixel quality and 4D geometry to closed-loop driving and human perceptual alignment, through five complementary aspects and 24 standardized dimensions and forms a unified ecosystem for assessing generated worlds not merely by visual appeal, but by physical and behavioral fidelity.

    5
  • EMOv2: Pushing 5M Vision Model Frontier

    Jiang-Ning Zhang, Teng Hu, Hao-Yang He, Zhu-Cun Xue, Yabiao Wang, Chengjie Wang, Yong Liu, Xiang-Tai Li, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025

    5
  • RecTok: Reconstruction Distillation along Rectified Flow

    Qingyu Shi, Size Wu, Jinbin Bai, Kai-Dong Yu, Yu-Jing Wang, Yun-Hai Tong, Xiang-Tai Li, Xue-Long Li

    arXiv · 2025

    This work proposes RecTok, which overcomes the limitations of high-dimensional visual tokenizers through two key innovations: flow semantic distillation and reconstruction--alignment distillation, to make the forward flow in flow matching semantically rich, rather than focusing on the latent space as in previous works.

    5
  • UMC: Unified Resilient Controller for Legged Robots with Joint Malfunctions

    Yu-Heng Qiu, Xin Lin, Jingbo Wang, Xiang-Tai Li, Lu Qi, Ming-Hsuan Yang

    arXiv · 2025

    A novel, model-free, two-stage training framework, Unified Malfunction Controller (UMC), incorporating a masking mechanism to enhance damage resilience, is proposed, incorporating a masking mechanism to enhance damage resilience.

    4
  • Auto Cherry-Picker : Learning from High-quality Generative Data Driven by Language

    Yicheng Chen, Xiang-Tai Li, Yining Li, Yanhong Zeng, Jianzong Wu, Xiangyu Zhao, Kai Chen

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work presents Auto Cherry-Picker (ACP), a novel framework that generates high-quality cross-modality training samples at scale to augment perception and multi-modal training and finds a positive correlation between CLIS and performance gains in downstream tasks.

    4
  • PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model

    Hao Yang, Qian-Yu Zhou, Haijia Sun, Xiang-Tai Li, Fengqi Liu, Xue-Quan Lu, Li-Zhuang Ma, Shui-Cheng Yan

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    This paper presents the first work that studies the generalizability of state space models (SSMs) in DG PCC and finds that directly applying SSMs into DG PCC will encounter several challenges.

    4
  • Masked Generative Transformer Is What You Need for Image Editing

    Wei Chow, Lin-Feng Li, Xianbing Sun, Ling-Dong Kong, Zefeng Li, Qi Xu, Hang Song, Tian Ye, +9 more

    arXiv · 2026

    This work presents EditMGT, an MGT-based editing framework that is the first of its kind and achieves state-of-the-art image similarity on multiple benchmarks while delivering 6x faster editing, demonstrating that MGTs offer a compelling alternative to diffusion-based editing.

    3
  • MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control

    Hua-Dai Liu, Jia-Lei Wang, Xiang-Tai Li, Wen Wang, Qian Chen, Rongjie Huang, Yang Liu, Jia-Yang Xu, +2 more

    ACM International Conference on Multimedia (ACM MM) · 2025

    MelodyEdit is proposed, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion, which outperforms state-of-the-art inversion techniques in editing fidelity and content preservation.

    3
  • 3
  • 3
  • CPT-Interp: Continuous sPatial and Temporal Motion Modeling for 4D Medical Image Interpolation

    Xia Li, Runzhao Yang, Xiang-Tai Li, A. Lomax, Ye Zhang, J. Buhmann

    arXiv · 2024

    This study draws inspiration from fluid mechanics to propose a novel approach for continuously modeling patient anatomic motion using implicit neural representation that ensures both spatial and temporal continuity, effectively bridging Eulerian and Lagrangian specifications together to naturally facilitate continuous frame interpolation.

    3
  • Flow2Seg: Motion-Aided Semantic Segmentation

    Xiang-Tai Li, Jian-Gang Bai, Kuiyuan Yang, Yun-Hai Tong

    Lecture notes in computer science · 2019

    This paper leverages motion information densely represented by optical flow to assist the semantic segmentation task and finds that optical flow improves image-based segmentation on object boundaries especially on small thin objects.

    3
  • 4th PVUW MeViS 3rd Place Report: Sa2VA

    Haobo Yuan, Tao Zhang, Xiang-Tai Li, Lu Qi, Zi-Long Huang, Shilin Xu, Jia-Shi Feng, Ming-Hsuan Yang

    arXiv · 2025

    This report shows that with a simple modification to the test time inference method on stronger MLLMs, it can lead to stronger results on MeVIS, and adopts the recent method Sa2VA, a unified model for dense grounded understanding of both images and videos.

    2
  • PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification

    Hao Yang, Qian-Yu Zhou, Haijia Sun, Xiang-Tai Li, Xue-Quan Lu, Li-Zhuang Ma, Shui-Cheng Yan

    arXiv · 2025

    This paper proposes PointDGRWKV, the first RWKV-based framework tailored for DG PCC, which introduces two key modules to enhance spatial modeling and cross-domain robustness, while maintaining RWKV's linear efficiency.

    2
  • MotionBooth: Motion-Aware Customized Text-to-Video Generation

    neural information processing systems · 2024

    2
  • MIMAFace: Face Animation via Motion-Identity Modulated Appearance Feature Learning

    Yue Han, Jun-Wei Zhu, Yu-Xiang Feng, Xiaozhong Ji, Keke He, Xiang-Tai Li, Zhu-Cun Xue, Yong Liu

    arXiv · 2024

    This work meticulously examines the essential appearance features in the facial animation tasks, and introduces a Motion-Identity Modulated Appearance Learning Module (MIA) that modulates CLIP features at both motion and identity levels, and designs an Inter-clip Affinity Learning Module (ICA) to model temporal relationships across clips.

    2
  • Fine-Grained Multimodal Alignment for Image-Text Retrieval via Graph Learning

    Mao Chen, Xiangkai Zhang, Lu Qi, Xiang-Tai Li, Xu Yang, Steven C. H. Hoi, Zhiyong Liu, Ming-Hsuan Yang

    International Journal of Computer Vision · 2026

    This work proposes a Graph-based Fine-Grained multimodal Alignment (GFGA) method, which adopts a concept-based fusion module in stage I to complement typical fragment-based fusion, which lacks comprehensive interactions.

    1
  • Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking

    Sheng-Qiong Wu, Bo-Bo Li, Xinkai Wang, Xiang-Tai Li, Lei Cui, Fu-Ru Wei, Shui-Cheng Yan, Hao Fei, +1 more

    arXiv · 2026

    The interleaved Analyzing-Drafting problem-solving loop (AD-Loop), a new think paradigm that dynamically alternates between analytic and drafting operations, is introduced, highlighting AD-Loop as a principled and broadly applicable strategy for synergizing comprehension and creation.

    1
  • 1
  • Bridge Feature Matching and Cross-Modal Alignment with Mutual-Filtering for Zero-Shot Anomaly Detection

    Yuhu Bai, Jiangning Zhang, Yunkang Cao, Guangyuan Lu, Qingdong He, Xiang-Tai Li, Guan-Zhong Tian

    IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) · 2025

    FiSeCLIP for ZSAD with training-free CLIP is introduced, combining the feature matching with the cross-modal alignment, and CLIP's inherent potential to restore its local semantic correlation is explored, adapting it for fine-grained anomaly detection tasks to enable a more accurate filtering process.

    1
  • 1
  • 1
  • Effective Adapter for Face Recognition in the Wild

    Yunhao Liu, Lu Qi, Yu-Ju Tsai, Xiang-Tai Li, Kelvin C. K. Chan, Ming-Hsuan Yang

    arXiv · 2023

    This paper proposes an effective adapter for augmenting existing face recognition models trained on high-quality facial datasets using two similar structures, one fixed and the other trainable, to process both the unrefined and enhanced images using two similar structures.

    1
  • Dynamic Dual Sampling Module For Fine-Grained Semantic Segmentation

    Chen Shi, Xiang-Tai Li, Yanran Wu, Yun-Hai Tong, Yi Xu

    Proceedings - International Conference on Image Processing · 2021

    A Dynamic Dual Sampling Module (DDSM) is proposed to conduct dynamic affinity modeling and propagate semantic context to local details, which yields a more discriminative representation.

    1
  • –
  • UniVR: Thinking in Visual Space for Unified Visual Reasoning

    Zhong-Wei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, An-Ran Wang, Xiang-Tai Li, Xiao-Jie Jin

    arXiv · 2026

    UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.

    –
  • From bits to atoms: A survey of cross-layer safety in Embodied AI

    Tuo Feng, Yue Zhang, Shi-Ji Zhou, Zhao-Xin Fan, Wen-Guan Wang, Xiao-Wei Chi, Tian-Yu Shen, Xiang-Tai Li, +22 more

    AI Plus · 2026

    This survey organizes embodied AI safety through a four-layer reference architecture and advocates for unifying software alignment with rigorous hardware constraints to build trustworthy systems capable of operating safely within human-centric environments.

    –
  • –
  • BACON: Bayesian Optimal Condensation Framework for Dataset Distillation

    Zheng Zhou, Hong Zhao, Guang-Liang Cheng, Xiang-Tai Li, Shuchang Lyu, Wen-Quan Feng, Qi Zhao

    arXiv · 2024

    The BAyesian optimal CONdensation framework (BACON) is proposed, which is the first work to introduce the Bayesian theoretical framework to the literature of DD and provides theoretical support for enhancing the performance of DD.

    –
  • –
  • –
  • Query Learning of Both Thing and Stuff for Panoptic Segmentation

    Shilin Xu, Xiang-Tai Li, Yi-Bo Yang, Hongyang Li, Guangliang Cheng, Yu Tong

    Proceedings - International Conference on Image Processing · 2022

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.