Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Tao XiangSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works35 from public data
- Simpler is Better: Few-shot Semantic Segmentation with Classifier Weight Transformer243
This work proposes to simplify the meta-learning task by focusing solely on the simplest component – the classifier, whilst leaving the en-coder and decoder to pre-training, and introduces a Classifier Weight Transformer (CWT) designed to dynamically adapt the support-set trained classifier’s weights to each query image in an inductive way.
- Stochastic Classifiers for Unsupervised Domain Adaptation185
This paper introduces a novel method called STochastic clAssifieRs (STAR) for addressing the problem of misaligned local regions between source and target domain, which finds that using more classifiers leads to better performance, but also introduces more model parameters, therefore risking overfitting.
- Task Residual for Tuning Vision-Language Models182
A new efficient tuning approach for VLMs named Task Residual Tuning (TaskRes), which performs directly on the text-based classifier and explicitly decouples the prior knowledge of the pre-trained models and new knowledge regarding a target task.
- Geometry Guided Adversarial Facial Expression Synthesis165
A Geometry-Guided Generative Adversarial Network (G2-GAN) for continuously-adjusting and identity-preserving facial expression synthesis and can generate compelling perceptual results on different expression editing tasks.
- GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph138
An effective adapter-style tuning strategy, dubbed GraphAdapter, which performs the textual adapter by explicitly modeling the dual-modality structure knowledge with a dual knowledge graph, which significantly outperforms previous adapter-based methods.
- Can SAM Boost Video Super-Resolution?39
This paper investigates a more robust and semantic-aware prior for enhanced VSR by utilizing the Segment Anything Model (SAM), a powerful foundational model that is less susceptible to image degradation.
- Uncertainty-Aware Source-Free Domain Adaptive Semantic Segmentation35
This paper proposes to use Bayesian Neural Network (BNN) to improve the target self- training by better estimating and exploiting pseudo-label uncertainty and introduces two novel self-training based components: Uncertainty-aware Online Teacher-Student Learning (UOTSL) and Uncertainly-aware FeatureMix (UFM).
- Prediction Calibration for Generalized Few-Shot Semantic Segmentation35
To make the proposed cross-attention module training tractable at the pixel level, this module is designed based on feature-score cross-covariance and episodically trained to be generalizable at inference time.
- Recent Progress of Face Image Synthesis35
A comprehensive review of typical face synthesis works that involve traditional methods as well as advanced deep learning approaches is provided, particularly, Generative Adversarial Net (GAN) is highlighted to generate photo-realistic and identity preserving results.
- P-BiC: Ultra-High-Definition Image Moiré Patterns Removal via Patch Bilateral Compensation32
A novel patch bilateral compensation network (P-BiC) for the demoire pattern removal in UHD images, which is memory-efficient and prior-knowledge-based, and quantitatively and qualitatively evaluates the effectiveness of P-BiC with extensive experiments.
- A Dive into SAM Prior in Image Restoration32
This paper uses the prior knowledge of the state-of-the-art segment anything model (SAM) to boost the performance of existing IR networks in an parameter-efficient tuning manner and proposes a lightweight SAM prior tuning (SPT) unit, which allows us to effectively integrate semantic priors into existingIR networks, resulting in significant improvements in restoration quality.
- Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation32
This study first explores the significance of high-frequency information in VLN and provides evidence that it is instrumental in bolstering visual-textual matching processes, and proposes a sophisticated and versatile Frequency-enhanced Data Augmentation technique to improve the VLN model’s capability of capturing critical high-frequency information.
- Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation30
This paper proposes Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics, and introduces contact-gated visuotactile fusion to selectively integrate tactile signals and a contact-aware attention bias to better regulate cross-modal interactions during manipulation.
- 28
- Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models24
This work introduces three customized ensemble strategies, each tailored to one specific scenario, and introduces the zero-shot ensemble, automatically adjusting the logits of different models based on their confidence when only pre-trained VLMs are available.
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination19
Efficient-WAM is introduced, a World-Action Model motivated by the idea that a compact future branch can still support effective action generation, even at lower visual fidelity, and achieves 98 ms per action chunk during physical deployment, over 30x faster than the Motus baseline with comparable task success.
- Conditional Expression Synthesis with Face Parsing Transformation19
A Couple-Agent Face Parsing based Generative Adversarial Network (CAFP-GAN) that unites the knowledge of facial semantic regions and controllable expression signals is introduced that has the compelling ability on continuous expression synthesis.
- 17
- 11
- Data Pyramid for Embodied Manipulation: A Survey8
This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity.
- Unraveling Motion Uncertainty for Local Motion Deblurring8
This work proposes a novel method named Motion-Uncertainty-Guided Network (MUGNet), which harnesses a probabilistic representational model to explicitly address the intricacies stemming from motion uncertainties and demonstrates the superiority of the MUGNet with extensive experiments.
- Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models7
Token Expand-and-Merge-VLA (TEAM-VLA), a training-free token compression framework that accelerates VLA inference while preserving task performance, and achieves a balanced trade-off between efficiency and effectiveness.
- 4
- Self-Evolved Imitation Learning in Simulated World3
Self-Evolved Imitation Learning (SEIL), a framework that progressively improves a few-shot model through simulator interactions, and introduces a lightweight selector that filters complementary and informative trajectories from the generated pool to ensure demonstration quality.
- COLA: Context-Aware Language-Driven Test-Time Adaptation3
This paper investigates a more general source model capable of adaptation to multiple target domains without needing shared labels, achieved by using a pre-trained vision-language model (VLM) that can recognize images through matching with class descriptions.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.