Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Mike Zheng Shou

    Possible advisor · last author on 6 of their early first-author papers, 2024–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works33 from public data

TitleCited by
  • Hallucination of Multimodal Large Language Models: A Survey

    Zechen Bai, Pichao Wang, Tianjun Xiao, Tong He, Zongbo Han, Zheng Zhang, Mike Zheng Shou

    arXiv · 2024

    A comprehensive analysis of the phenomenon of hallucination in multimodal large language models (MLLMs), also known as Large Vision-Language Models (LVLMs), offering a detailed overview of the underlying causes, evaluation benchmarks, metrics, and strategies developed to address this issue.

    466
  • ShowUI: One Vision-Language-Action Model for GUI Visual Agent

    Kevin Qinghong Lin, Linjie Li, Difei Gao, Zhengyuan Yang, Shiwei Wu, Zechen Bai, Stan Weixian Lei, Lijuan Wang, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work develops a vision-language-action model in digital world, namely ShowUI, which features the following innovations: UI-Guided Visual Token Selection to reduce computational costs by formulating screenshots as an UI connected graph, adaptively identifying their redundant relationship and serve as the criteria for token selection during self-attention blocks.

    278
  • Unsupervised Multi-Source Domain Adaptation for Person Re-Identification

    Zechen Bai, Zhigang Wang, Jian Wang, Di Hu, Errui Ding

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    The proposed method outperforms state-of-the-art UDA person re-ID methods by a large margin, and even achieves comparable performance to the supervised approaches without any post-processing techniques.

    96
  • Show, Recall, and Tell: Image Captioning with Recall Mechanism

    Li Wang, Zechen Bai, Yonghua Zhang, Hongtao Lu

    Proceedings of the AAAI Conference on Artificial Intelligence · 2020

    Inspired by pointing mechanism in text summarization, this paper adopts a soft switch to balance the generated-word probabilities between SG and RWS in the CIDEr optimization step, and introduces an individual recalled-word reward (WR) to boost training.

    74
  • Going Beyond Real Data: A Robust Visual Representation for Vehicle Re-identification

    Zhedong Zheng, Minyue Jiang, Zhigang Wang, Jian Wang, Zechen Bai, Xuanmeng Zhang, Xin Feng Yu, Xiao Tan, +3 more

    IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2020

    63
  • Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation

    Zechen Bai, Yuta Nakashima, Noa García

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings by modules the generated sentences according to three artistic topics and enhances each description with external knowledge.

    58
  • ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation

    Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Mao, Dongxing, Qinchen Wu, Weichen Zhang, +5 more

    arXiv · 2023

    An advanced Actor-Critic Embodied Agent framework is proposed, which incorporates a sophisticated GUI parser driven by an LLM-agent and an enhanced reasoning mechanism adept at handling lengthy procedural tasks that outshine existing methods in performance.

    49
  • Adaptive Slot Attention: Object Discovery with Dynamic Slot Number

    Ke Fan, Zechen Bai, Tianjun Xiao, Tong He, Max Horn, Yanwei Fu, Francesco Locatello, Zheng Zhang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    An adaptive slot attention (AdaSlot) mecha-nism that dynamically determines the optimal number of slots based on the content of the data is introduced that exhibits the capability to dynamically adapt the slot number according to each instance's complexity.

    38
  • Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

    Yiqi Lin, Guoqiang Liang, Ziyun Zeng, Zechen Bai, Yanzhe Chen, Mike Zheng Shou

    arXiv · 2026

    A scalable data generation pipeline is introduced that transforms existing video editing pairs into high-fidelity training quadruplets, leveraging image generative models to create synthesized reference scaffolds, and a unified editing architecture is proposed, Kiwi-Edit, that synergizes learnable queries and latent visual features for reference semantic guidance.

    36
  • World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy

    Xiaokang Liu, Zechen Bai, Hai Ci, Kevin Ma, Mike Zheng Shou

    arXiv · 2026

    World-VLA-Loop substantially improves VLA performance while reducing reliance on costly physical interaction and feeding rollouts from each improved policy back to augment and fine-tune the world model.

    36
  • AssistGUI: Task-Oriented PC Graphical User Interface Automation

    Difei Gao, Lei Ji, Zechen Bai, Mingyu Ouyang, Peiran Li, Donzxing Mao, Qinchen Wu, Weichen Zhang, +5 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    32
  • One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos

    Zechen Bai, Tong He, Haiyang Mei, Pichao Wang, Ziteng Gao, Joya Chen, Lei Liu, Zheng Zhang, +1 more

    neural information processing systems · 2024

    31
  • Robust Vehicle Re-identification via Rigid Structure Prior

    Minyue Jiang, Xuanmeng Zhang, Yue Yu, Zechen Bai, Zhedong Zheng, Zhigang Wang, Jian Wang, Xiao Tan, +3 more

    IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) · 2021

    23
  • Enhancing Emotional Experience by Building Emotional Virtual Characters in VR Volleyball Games

    Zechen Bai, Naiming Yao, Nidhi Mishra, Hui Chen, Hongan Wang, Nadia Magnenat‐Thalmann

    Computer Animation and Virtual Worlds · 2021

    This article proposes to enhance emotional experience by building emotional virtual characters in VR volleyball games that cannot only arouse their emotions but also express their facial expressions according to the game situation.

    22
  • EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models

    Zechen Bai, Gao, Chen, Mike Zheng Shou

    arXiv · 2025

    EVOLVE-VLA is introduced, a test-time training framework enabling VLAs to continuously adapt through environment interaction with minimal or zero task-specific demonstrations, and represents a critical step toward VLAs that truly learn and adapt, moving beyond static imitation toward continuous self-improvements.

    20
  • Object-Centric Multiple Object Tracking

    Zixu Zhao, Jiaze Wang, Max Horn, Yizhuo Ding, Tong He, Zechen Bai, Dominik Zietlow, Carl-Johann Simon-Gabriel, +8 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    This paper proposes a video object-centric model for MOT that consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions.

    15
  • Unsupervised Open-Vocabulary Object Localization in Videos

    Ke Fan, Zechen Bai, Tianjun Xiao, Dominik Zietlow, Max Horn, Zixu Zhao, Carl-Johann Simon-Gabriel, Mike Zheng Shou, +6 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    A method that first localizes objects in videos via a slot attention approach and then assigns text to the obtained slots is proposed, which is effectively the first unsupervised approach that yields good results on regular video benchmarks.

    15
  • Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters

    Zechen Bai, Peng Chen, Xiaolan Peng, Lu Liu, Naiming Yao, Hui Chen

    IEEE Conference on Virtual Reality and 3D User Interfaces (VR) · 2024

    A deep learning model was first trained to retarget the facial expression from input face images to virtual human faces by estimating the blendshape coefficients and a practical toolkit was developed using Unity 3D, making it compatible with the most popular VR applications.

    14
  • DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting

    Binqian Xu, Xiangbo Shu, Haiyang Mei, Zechen Bai, Basura Fernando, Mike Zheng Shou, Jinhui Tang

    neural information processing systems · 2024

    Experimental results on diverse datasets consistently demonstrate that DoFIT excels in cross-domain collaborative training and exhibits significant advantages over conventional FIT methods in alleviating catastrophic forgetting.

    13
  • Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models

    Zongbo Han, Zechen Bai, Haiyang Mei, Qianli Xu, Changqing Zhang, Mike Zheng Shou

    arXiv · 2024

    A new perspective is proposed, suggesting that the inherent biases in LVLMs might be a key factor in hallucinations, and a simple method is proposed to effectively mitigate the hallucination of LVLMs by skipping the output of '\n'.

    11
  • Factorized Visual Tokenization and Generation

    Zechen Bai, Jianxiong Gao, Ziteng Gao, Pichao Wang, Zheng Zhang, Tong He, Mike Zheng Shou

    arXiv · 2024

    Factorized Quantization (FQ) is introduced, a novel approach that revitalizes VQ-based tokenizers by decomposing a large codebook into multiple independent sub-codebooks, enabling more efficient and scalable visual tokenization.

    11
  • Play with Emotional Characters: Improving User Emotional Experience by A Data-driven Approach in VR Volleyball Games

    Zechen Bai, Naiming Yao, Nidhi Mishra, Hui Chen, Hongan Wang, Nadia Magnenat‐Thalmann

    IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW) · 2021

    9
  • Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation

    Yanzhe Chen, Kevin Ma, Qi Lv, Yiqi Lin, Zechen Bai, Chen Gao, Mike Zheng Shou

    arXiv · 2026

    This work identifies a critical diversity trap and proposes Anchor-Centric Adaptation (ACA), a two-stage framework that first stabilizes a policy skeleton through repeated demonstrations at core anchors, then selectively expands coverage to high-risk boundaries via teacher-forced error mining and constrained residual updates.

    5
  • AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation

    Difei Gao, Siyuan Hu, Zechen Bai, Kevin Qinghong Lin, Mike Zheng Shou

    ACM International Conference on Multimedia (ACM MM) · 2024

    A novel PC-Copilot, AssistEditor, that focuses on automating the video editing workflow, and significantly streamlines the video editing process, making advanced editing accessible to users with varying levels of expertise.

    5
  • SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

    Xiaoyan Zhang, Zechen Bai, Haofan Wang, Yiren Song

    arXiv · 2026

    SIGMA (Selective-Interleaved Generation with Multi-Attribute Tokens), a unified post-training framework that enables interleaved multi-condition generation within diffusion transformers, introduces selective multi-attribute tokens, which allow the model to interpret and compose multiple visual conditions in an interleaved text-image sequence.

    4

Show all 33 works

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.