Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Mike Zheng Shou

    Possible advisor · last author on 4 of their early first-author papers, 2023–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works15 from public data

TitleCited by
  • Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

    Jinheng Xie, Weijia Mao, Zechen Bai, David Junhao Zhang, Weihao Wang, Kevin Qinghong Lin, Yuchao Gu, Zhijie Chen, +2 more

    arXiv · 2024

    Across various benchmarks, Show-o demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation, which significantly highlights its potential as a next-generation foundation model.

    792
  • Long-Context Autoregressive Video Modeling with Next-Frame Prediction

    Yuchao Gu, Weijia Mao, Mike Zheng Shou

    arXiv · 2025

    This paper proposes the long short-term context modeling using asymmetric patchify kernels, which apply large kernels to distant frames to reduce redundant tokens, and standard kernels to local frames to preserve fine-grained detail, providing an effective baseline for long-context autoregressive video modeling.

    126
  • Exocentric-to-Egocentric Video Generation

    Jiawei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, Mike Zheng Shou

    neural information processing systems · 2024

    This work designs an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model, and introduces the temporal attention layers into the egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames.

    28
  • DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing

    Jiawei Liu, Yan‐Pei Cao, Jay Zhangjie Wu, Weijia Mao, Yuchao Gu, Rui Zhao, Jussi Keppo, Ying Shan, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work proposes the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer.

    28
  • UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

    Weijia Mao, Zhenheng Yang, Mike Zheng Shou

    arXiv · 2025

    UniRL, a self-improving post-training approach that enables the model to generate images from prompts and use them as training data in each iteration, without relying on any external image data, is introduced.

    23
  • Mitty: Diffusion-based Human-to-Robot Video Generation

    Yiren Song, Cheng Liu, Weijia Mao, Mike Zheng Shou

    arXiv · 2025

    Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation and develops an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets.

    13
  • The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation

    Weijia Mao, Hao Chen, Zhenheng Yang, Mike Zheng Shou

    arXiv · 2025

    Adv-GRPO is introduced, an RL framework with an adversarial reward that iteratively updates both the reward model and the generator, and it is shown that combining reference samples with foundation-model rewards enables distribution transfer and flexible style customization.

    10
  • ShowRoom3D: Text to High-Quality 3D Room Generation Using 3D Priors

    Weijia Mao, Yan‐Pei Cao, Jiawei Liu, Zhongcong Xu, Mike Zheng Shou

    arXiv · 2023

    ShowRoom3D enables the generation of rooms with improved structural integrity, enhanced clarity from any view, reduced content repetition, and higher consistency across different perspectives, and significantly outperforms state-of-the-art approaches by a large margin in terms of user study.

    8
  • Novel View Synthesis for High-fidelity Headshot Scenes

    Satoshi Tsutsui, Weijia Mao, Sijing Lin, Yunyi Zhu, Ma, Murong, Mike Zheng Shou

    arXiv · 2022

    This work learns a Generative Adversarial Network to mix a NeRF-synthesized image and a 3DMM-rendered image and produces a photorealistic scene with a face preserving the skin details, and experiments with various real-world scenes demonstrate the effectiveness of this approach.

    7
  • BitDance: Scaling Autoregressive Generative Models with Binary Tokens

    Yuang Ai, Jiaming Han, Shaobin Zhuang, Weijia Mao, Xuefeng Hu, Ziyan Yang, Zhenheng Yang, Yali Wang, +3 more

    arXiv · 2026

    BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices, and proposes next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference.

    6
  • DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

    Rui Zhao, Weijia Mao, Mike Zheng Shou

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    DoraCycle is proposed, which integrates two multimodal cycles: text-to-image-to-text and image-to-text-to-image and is optimized through cross-entropy loss computed at the cycle endpoints, where both endpoints share the same modality.

    6
  • UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

    Weijia Mao, Zhenheng Yang, Mike Zheng Shou

    arXiv · 2025

    UniMoD is introduced, a task-aware token pruning method that employs a separate router for each task to determine which tokens should be pruned, and reveals that token redundancy is primarily influenced by different tasks and layers.

    5
  • DeVRF: Fast Deformable Voxel Radiance Fields for Dynamic Scenes

    Jia-Wei Liu, Yan-Pei Cao, Weijia Mao, Wenqiao Zhang, David Junhao Zhang, Jussi Keppo, Ying Shan, Xiaohu Qie, +1 more

    neural information processing systems · 2022

    4
  • UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model

    Shaobin Zhuang, Yuang Ai, Jiaming Han, Weijia Mao, Xiaohui Li, Fangyikang Wang, Xiao Yi Wang, Yan Li, +7 more

    arXiv · 2026

    This paper introduces UniWeTok, a unified discrete tokenizer designed to bridge this gap using a massive binary codebook and introduces Pre-Post Distillation and a Generative-Aware Prior to enhance the semantic extraction and generative prior of the discrete tokens.

    –
  • AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

    Yuchao Gu, Guian Fang, Yuxin Jiang, Weijia Mao, Song Han, Han Cai, Mike Zheng Shou

    Lecture notes in computer science · 2026

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.