Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Ping LuoSuggested from co-authorship
- Liang LinSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works22 from public data
- MotionCtrl: A Unified and Flexible Motion Controller for Video Generation673
MotionCtrl is presented, a unified and flexible motion controller for video generation designed to effectively and independently control camera and object motion and is a relatively generalizable model that can adapt to a wide array of camera poses and trajectories once trained.
- Multi-label Image Recognition by Recurrently Discovering Attentional Regions338
This work achieves the interpretable and contextualized multi-label image classification by developing a recurrent memorized-attention module that demonstrates superior performances over other existing state-of-the-arts in both accuracy and efficiency.
- Deep Reasoning with Knowledge Graph for Social Relationship Understanding198
This work has found that the interplay between these two factors can be effectively modeled by a novel structured knowledge graph with proper message propagation and attention and can be efficiently integrated into the deep neural network architecture to promote social relationship understanding.
- RestoreFormer: High-Quality Blind Face Restoration from Undegraded Key-Value Pairs177
This work proposes a method, RestoreFormer, which explores fully-spatial attentions to model contextual information and surpasses existing works that use local operators and outperforms advanced state-of-the-art methods on one synthetic dataset and three real-world datasets.
- Recurrent Attentional Reinforcement Learning for Multi-Label Image Recognition176
A recurrent attention reinforcement learning framework to iteratively discover a sequence of attentional and informative regions that are related to different semantic objects and further predict label scores conditioned on these regions to facilitate multi-label recognition.
- LSTM Pose Machines148
It is shown that if the weight sharing scheme is imposed to the multi-stage CNN, it could be re-written as a Recurrent Neural Network (RNN), which decouples the relationship among multiple network stages and results in significantly faster speed in invoking the network for videos.
- FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models61
A tuning-free framework to achieve trajectory-controllable video generation, by imposing guidance on both noise construction and attention computation, and proposes FreeTraj, a tuning-free approach that enables trajectory control by modifying noise sampling and attention mechanisms.
- RestoreFormer++: Towards Real-World Blind Face Restoration From Undegraded Key-Value Pairs56
This work proposes RestoreFormer++, which on the one hand introduces fully-spatial attention mechanisms to model the contextual information and the interplay with the priors, and on the other hand explores an extending degrading model to help generate more realistic degraded face images to alleviate the synthetic-to-real-world gap.
- StyleAdapter: A Unified Stylized Image Generation Model52
The StyleAdapter is proposed, a unified stylized image generation model capable of producing a variety of stylized images that match both the content of a given prompt and the style of reference images, without the need for per-style fine-tuning.
- Diffusion-based Blind Text Image Super-Resolution43
Extensive experiments on synthetic and real-world datasets demonstrate that the Diffusion-based Blind Text Image Super-Resolution (DiffTSR) can restore text images with more accurate text structures as well as more realistic appearances simultaneously.
- Learning a Reinforced Agent for Flexible Exposure Bracketing Selection24
A novel deep neural network to automatically select exposure bracketing, named EBSNet, which is sufficiently flexible without having the above restrictions and can be jointly trained to produce favorable results against recent state-of-the-art approaches.
- RoVer: Robot Reward Model as Test-Time Verifier for Vision-Language-Action Model21
The approach effectively transforms available computing resources into better action decision-making, realizing the benefits of test-time scaling without extra training overhead.
- Multi-label image recognition with attentive transformer-localizer module18
A novel attentive transformer-localizer (ATL) module is applied to progressively localize the attentive regions from the convolutional feature maps in a proposal-free manner, and the LSTM network sequentially predicts label scores for the localized regions and updates the parameters of the ATL module while capturing the global dependencies among these regions.
- 17
- Denoising as Adaptation: Noise-Space Domain Adaptation for Image Restoration17
This paper shows that it is possible to perform domain adaptation via the noise space using diffusion models and derives a meaningful diffusion loss that guides the restoration model in progressively aligning both restored synthetic and real-world outputs with a target clean distribution.
- Novel Antimicrobial Nano Bacteriocin: Lactic Acid Bacteria‐Derived, Self‐Assembled, and Enhanced for Superior Antimicrobial Activity14
The carrier‐free self‐assembly approach overcomes AMP stability and solubility limitations and paves the way for next‐generation antimicrobial therapies.
- Image Conductor: Precision Control for Interactive Video Synthesis11
Image Conductor, a method for precise control of camera transitions and object movements to generate video assets from a single image, is proposed, and a trajectory-oriented video motion data curation pipeline for training is developed.
- 8
- Analysis and Benchmarking of Extending Blind Face Image Restoration to Videos7
A Temporal Consistency Network (TCN) cooperated with alignment smoothing to reduce jitters and flickers in restored videos is proposed, a flexible component that can be seamlessly plugged into the most advanced face image restoration algorithms, ensuring the quality of image-based restoration is maintained as closely as possible.
- Image Deblurring Aided by Low-Resolution Events6
An alternately performed model is proposed in this paper to deblur high-resolution images with the help of low-resolution events and enhances the quality of events with EventSRNet by extracting the structure information in the generated sharp image.
- 1
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.