Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Li YuanSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works17 from public data
- 560
- LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment490
This work proposes LanguageBind, taking the language as the bind across different modalities because the language modality is well-explored and contains rich semantics, and freezes the language encoder acquired by VL pretraining, then train encoders for other modalities with contrastive learning.
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models390
A baseline for sparse LVLMs is established and empirical guidelines for exploring the sparse LVLMs are provided, which uniquely activates only the top-$k$ experts through routers during deployment, keeping the remaining experts inactive.
- Open-Sora Plan: Open-Source Large Video Generation Model325
This work introduces Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs, and hopes it can inspire the video generation research community.
- Cosmos 3: Omnimodal World Models for Physical AI150
This report introduces Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture and demonstrates omnimodal world models as scalable, general-purpose backbones for embodied agents.
- 92
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation51
This work proposes UniWorld-V1, a unified generative framework built upon semantic features extracted from powerful multimodal large language models and contrastive semantic encoders, which achieves impressive performance across diverse tasks, including image understanding, generation, manipulation, and perception.
- VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention41
VideoGen-of-Thought is introduced, a step-by-step framework that automates multi-shot video synthesis from a single sentence by systematically addressing three core challenges: Narrative fragmentation, identity-aware cross-shot propagation, and transition artifacts.
- Repaint123: Fast and High-quality One Image to 3D Generation with Progressive Controllable 2D Repainting28
The core idea is to combine the powerful image generation capability of the 2D diffusion model and the texture alignment ability of the repainting strategy for generating high-quality multi-view images with consistency to alleviate multi-view bias as well as texture degradation and speed up the generation process.
- DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses20
This work presents DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs, and introduces a Mutually Aligned Geometry Diffusion Model to generate fine-grained depth and normal maps for enriched guidance.
- 18
- Envision3D: One Image to 3D with Anchor Views Interpolation14
A novel cascade diffusion framework is proposed, which decomposes the challenging dense views generation task into two tractable stages, namely anchor views generation and anchor views interpolation, and yields dense, multi-view consistent images, providing comprehensive 3D information.
- E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras12
This work proposes E-4DGS, the first event-driven dynamic Gaussian Splatting approach, for novel view synthesis from multi-view event streams with fast-moving cameras, and introduces an event-based initialization scheme to ensure stable training and proposes event-adaptive slicing splatting for time-aware reconstruction.
- 12
- 1
- –
- SwapAnyone: Consistent and Realistic Video Synthesis for Swapping Any Person into Any Video–
An end-to-end model named SwapAnyone is introduced, treating video body-swapping as a video inpainting task with reference fidelity and motion control, and introducing a novel EnvHarmony strategy for training the authors' model progressively to improve the ability to maintain environmental harmony.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.