Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Nenghai YuSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works23 from public data
- Multi-attentional Deepfake Detection995
A new multi-attentional deepfake detection network that consists of three key components: multiple spatial attention heads to make the network attend to different local parts, a new regional independence loss and an attention guided data augmentation strategy, and state-of-the-art performance.
- HairCLIP: Design Your Hair by Text and Reference Image147
This paper proposes a new hair editing interaction mode, which enables manipulating hair attributes individually or jointly based on the texts or reference images provided by users, and encode the image and text conditions in a shared embedding space and proposes a unified hair editing framework.
- SimAC: A Simple Anti-Customization Method for Protecting Face Privacy Against Text-to-Image Synthesis of Diffusion Models56
This paper examines the relationship between time step selection and the model's perception in the frequency domain of images and finds that lower time steps can give much more contributions to adversarial noises, and proposes an adaptive greedy search for optimal time steps that seamlessly integrates with existing anti-customization methods.
- Improved Image Matting via Real-time User Clicks and Uncertainty Estimation42
An improved deep image matting framework which is trimap-free and only needs several user click interactions to eliminate the ambiguity is proposed, and a new uncertainty estimation module that can predict which parts need polishing and a following local refinement module is introduced.
- HairCLIPv2: Unifying Hair Editing via Proxy Feature Blending31
Besides the unprecedented user interaction mode support, quantitative and qualitative experiments demonstrate the superiority of HairCLIPv2 in terms of editing effects, irrelevant attribute preservation and visual naturalness.
- FreeFlux: Understanding and Exploiting Layer-Specific Roles in RoPE-Based MMDiT for Versatile Image Editing24
This work presents the first mechanistic analysis of RoPE-based MMDiT models, introducing an automated probing strategy that disentangles positional information versus content dependencies by strategically manipulating RoPE during generation.
- A Geometric Distortion Immunized Deep Watermarking Framework with Robustness Generalizability15
A Swin Transformer and Deformable Convolu-tional Network (DCN)-based watermark model backbone that effectively improves the feature processing flexibility, greatly enhancing the robustness, especially for geometric distortions.
- Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation14
This work designs three loss functions: Block Alignment Loss, Text Encoder Alignment Loss, and Overlap Loss, each tailored to mitigate ambiguities within the MMDiT architecture that cause semantic ambiguity persists when generating multiple similar subjects.
- Bokeh Diffusion: Defocus Blur Control in Text-to-Image Diffusion Models13
This work proposes Bokeh Diffusion, a scene-consistent bokeh control framework that explicitly conditions a diffusion model on a physical defocus blur parameter and introduces a hybrid training pipeline that aligns in-the-wild images with synthetic blur augmentations, providing diverse scenes and subjects as well as supervision to learn the separation of image content from lens blur.
- 13
- UniForensics: Face Forgery Detection via General Facial Representation10
UniForensics is introduced, a novel deepfake detection framework that leverages a transformer-based video classification network, initialized with a meta-functional face encoder for enriched facial representation, that outperforms existing face forgery detection methods in generalization ability and robustness.
- FaceRSA: RSA-Aware Facial Identity Cryptography Framework9
This paper presents the first facial identity cryptography framework with full properties analogous to RSA, and leverages the powerful generative capabilities of StyleGAN to achieve megapixel-level facial identity anonymization and deanonymization.
- 8
- PI-Light: Physics-Inspired Diffusion for Full-Image Relighting5
Experiments demonstrate that $\pi$-Light synthesizes specular highlights and diffuse reflections across a wide variety of materials, achieving superior generalization to real-world scenes compared with prior approaches.
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling4
Self-Adaptive Attention Scaling (SaaS), a method that leverages the consistency of cross-attention between adjacent timesteps to dynamically scale the attention activation for each sub-instruction.
- Rank-Based No-Reference Quality Assessment for Face Swapping1
This work observed that consistency among errors in various facial attributes can indicate overall image quality, and proposed a no-reference image quality assessment method for face swapping that achieved state-of-the-art results in both coarse- and fine-grained tests.
- 1
- –
- –
- –
- PnP-U3D: Plug-and-Play 3D Framework Bridging Autoregression and Diffusion for Unified Understanding and Generation–
This work presents the first unified framework for 3D understanding and generation that combines autoregression with diffusion, and adopts an autoregressive next-token prediction paradigm for 3D understanding, and a continuous diffusion paradigm for 3D generation.
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.