Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Suggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works15 from public data
- Show-o: One Single Transformer to Unify Multimodal Understanding and Generation792
Across various benchmarks, Show-o demonstrates comparable or superior performance to existing individual models with an equivalent or larger number of parameters tailored for understanding or generation, which significantly highlights its potential as a next-generation foundation model.
- Long-Context Autoregressive Video Modeling with Next-Frame Prediction126
This paper proposes the long short-term context modeling using asymmetric patchify kernels, which apply large kernels to distant frames to reduce redundant tokens, and standard kernels to local frames to preserve fine-grained detail, providing an effective baseline for long-context autoregressive video modeling.
- Exocentric-to-Egocentric Video Generation28
This work designs an exocentric-to-egocentric view translation prior to provide spatially aligned egocentric features as a concatenation guidance for the input of egocentric video diffusion model, and introduces the temporal attention layers into the egocentric video diffusion pipeline to improve the temporal consistency cross egocentric frames.
- DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing28
This work proposes the image-based video-NeRF editing pipeline with a set of innovative designs, including multi-view multi-pose Score Distillation Sampling (SDS) from both the 2D personalized diffusion prior and 3D diffusion prior, reconstruction losses, text-guided local parts super-resolution, and style transfer.
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning23
UniRL, a self-improving post-training approach that enables the model to generate images from prompts and use them as training data in each iteration, without relying on any external image data, is introduced.
- Mitty: Diffusion-based Human-to-Robot Video Generation13
Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation and develops an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets.
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation10
Adv-GRPO is introduced, an RL framework with an adversarial reward that iteratively updates both the reward model and the generator, and it is shown that combining reference samples with foundation-model rewards enables distribution transfer and flexible style customization.
- ShowRoom3D: Text to High-Quality 3D Room Generation Using 3D Priors8
ShowRoom3D enables the generation of rooms with improved structural integrity, enhanced clarity from any view, reduced content repetition, and higher consistency across different perspectives, and significantly outperforms state-of-the-art approaches by a large margin in terms of user study.
- Novel View Synthesis for High-fidelity Headshot Scenes7
This work learns a Generative Adversarial Network to mix a NeRF-synthesized image and a 3DMM-rendered image and produces a photorealistic scene with a face preserving the skin details, and experiments with various real-world scenes demonstrate the effectiveness of this approach.
- BitDance: Scaling Autoregressive Generative Models with Binary Tokens6
BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices, and proposes next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference.
- DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles6
DoraCycle is proposed, which integrates two multimodal cycles: text-to-image-to-text and image-to-text-to-image and is optimized through cross-entropy loss computed at the cycle endpoints, where both endpoints share the same modality.
- UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths5
UniMoD is introduced, a task-aware token pruning method that employs a separate router for each task to determine which tokens should be pruned, and reveals that token redundancy is primarily influenced by different tasks and layers.
- 4
- UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model–
This paper introduces UniWeTok, a unified discrete tokenizer designed to bridge this gap using a massive binary codebook and introduces Pre-Post Distillation and a Generative-Aware Prior to enhance the semantic extraction and generative prior of the discrete tokens.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.