Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Li Yuan

    Possible advisor · last author on 6 of their early first-author papers, 2022–2026

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works17 from public data

TitleCited by
  • Masked Autoencoders for Point Cloud Self-supervised Learning

    Yatian Pang, Wenxiao Wang, Francis E. H. Tay, Wei Liu, Yonghong Tian, Li Yuan

    Lecture notes in computer science · 2022

    560
  • LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

    Bin Zhu, Bin Lin, Munan Ning, Yan Qin Yang, Cui, Jiaxi, HongFa Wang, Yatian Pang, Wenhao Jiang, +6 more

    arXiv · 2023

    This work proposes LanguageBind, taking the language as the bind across different modalities because the language modality is well-explored and contains rich semantics, and freezes the language encoder acquired by VL pretraining, then train encoders for other modalities with contrastive learning.

    490
  • MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

    Bin Lin, Zhenyu Tang, Yang Ye, Jinfa Huang, Junwu Zhang, Yatian Pang, Peng Jin, Munan Ning, +2 more

    IEEE Transactions on Multimedia · 2026

    A baseline for sparse LVLMs is established and empirical guidelines for exploring the sparse LVLMs are provided, which uniquely activates only the top-$k$ experts through routers during deployment, keeping the remaining experts inactive.

    390
  • Open-Sora Plan: Open-Source Large Video Generation Model

    Bin Lin, Yunyang Ge, Xinhua Cheng, Zongjian Li, Bin Benjamin Zhu, Shaodong Wang, He, Xianyi, Ye Yang, +16 more

    arXiv · 2024

    This work introduces Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs, and hopes it can inspire the video generation research community.

    325
  • Cosmos 3: Omnimodal World Models for Physical AI

    NVIDIA, Niket Agarwal, Aditi, Niket Agarwal, Arslan Ali, Jon Allen, Martin Antolini, Adeline Aubame, +22 more

    arXiv · 2026

    This report introduces Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture and demonstrates omnimodal world models as scalable, general-purpose backbones for embodied agents.

    150
  • Masked Autoencoders for 3D Point Cloud Self-supervised Learning

    Yatian Pang, Eng Hock Tay, Yuan Li, Zhenghua Chen

    World Scientific Annual Review of Artificial Intelligence · 2024

    92
  • UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

    Bin Lin, Zongjian Li, Xinhua Cheng, Niu, Yuwei, Ye Yang, Xianyi He, Shenghai Yuan, Wangbo Yu, +4 more

    arXiv · 2025

    This work proposes UniWorld-V1, a unified generative framework built upon semantic features extracted from powerful multimodal large language models and contrastive semantic encoders, which achieves impressive performance across diverse tasks, including image understanding, generation, manipulation, and perception.

    51
  • VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention

    Mingzhe Zheng, Xu, Yongqi, Haojian Huang, Xiaoliang Ma, Yexin Liu, Shu, Wenjie, Yatian Pang, Tang, Feilong, +3 more

    arXiv · 2025

    VideoGen-of-Thought is introduced, a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: Narrative fragmentation, identity-aware cross-shot propagation, and transition artifacts.

    41
  • Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting

    Junwu Zhang, Zhenyu Tang, Yatian Pang, Xinhua Cheng, Jin, Peng, Yida Wei, Munan Ning, Yuan, Li

    arXiv · 2023

    The core idea is to combine the powerful image generation capability of the 2D diffusion model and the texture alignment ability of the repainting strategy for generating high-quality multi-view images with consistency to alleviate multi-view bias as well as texture degradation and speed up the generation process.

    28
  • DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses

    Yatian Pang, Bin B. Zhu, Bin Bin, Mingzhe Zheng, Francis E. H. Tay, Ser-Nam Lim, Harry Yang, Yuan Li

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    This work presents DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs, and introduces a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance.

    20
  • Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle

    Zhenyu Tang, Junwu Zhang, Xinhua Cheng, Wangbo Yu, Chaoran Feng, Yatian Pang, Bin Lin, Yuan Li

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    18
  • Envision3D: One Image to 3D with Anchor Views Interpolation

    Yatian Pang, Jia, Tanghui, Yujun Shi, Zhenyu Tang, Junwu Zhang, Xinhua Cheng, Xing Zhou, Francis E. H. Tay, +1 more

    arXiv · 2024

    A novel cascade diffusion framework is proposed, which decomposes the challenging dense views generation task into two tractable stages, namely anchor views generation and anchor views interpolation, and yields dense, multi-view consistent images, providing comprehensive 3D information.

    14
  • E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras

    Chaoran Feng, Zhenyu Tang, Wangbo Yu, Yatian Pang, Yian Zhao, Jianbin Zhao, Li Yuan, Yonghong Tian

    ACM International Conference on Multimedia (ACM MM) · 2025

    This work proposes E-4DGS, the first event-driven dynamic Gaussian Splatting approach, for novel view synthesis from multi-view event streams with fast-moving cameras, and introduces an event-based initialization scheme to ensure stable training and proposes event-adaptive slicing splatting for time-aware reconstruction.

    12
  • Repaint123: Fast and High-Quality One Image to 3D Generation with Progressive Controllable Repainting

    Junwu Zhang, Zhenyu Tang, Yatian Pang, Xinhua Cheng, Peng Jin, Yida Wei, Xing Zhou, Munan Ning, +1 more

    Lecture notes in computer science · 2024

    12
  • Abnormal Wedge Bond Detection Using Convolutional Autoencoders in Industrial Vision Systems

    Ji-Yan Wu, Yatian Pang, Xiang Li, Wen Feng Lu

    International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME) · 2022

    1
  • Next Patch Prediction for AutoRegressive Visual Generation

    Yatian Pang, Peng JIN, Shuo Yang, Bin Benjamin Zhu, Bin Lin, Chaoran Feng, Zhenyu Tang, Liuhan Chen, +4 more

    Proceedings of the AAAI Conference on Artificial Intelligence · 2026

    –
  • SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video

    Chengshu Zhao, Yunyang Ge, Xinhua Cheng, Bin B. Zhu, Yatian Pang, Lin, Bin, Fan Yang, Feng Gao, +1 more

    arXiv · 2025

    An end-to-end model named SwapAnyone is introduced, treating video body-swapping as a video inpainting task with reference fidelity and motion control, and introducing a novel EnvHarmony strategy for training the authors' model progressively to improve the ability to maintain environmental harmony.

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.