Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Tao Xiang

    Possible advisor · last author on 3 of their early first-author papers, 2020–2021

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works35 from public data

TitleCited by
  • Simpler is Better: Few-shot Semantic Segmentation with Classifier Weight Transformer

    Zhihe Lu, Sen He, Xiatian Zhu, Li Zhang, Yi-Zhe Song, Tao Xiang

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    This work proposes to simplify the meta-learning task by focusing solely on the simplest component – the classifier, whilst leaving the en-coder and decoder to pre-training, and introduces a Classifier Weight Transformer (CWT) designed to dynamically adapt the support-set trained classifier’s weights to each query image in an inductive way.

    243
  • Stochastic Classifiers for Unsupervised Domain Adaptation

    Zhihe Lu, Yongxin Yang, Xiatian Zhu, Cong Liu, Yi-Zhe Song, Tao Xiang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2020

    This paper introduces a novel method called STochastic clAssifieRs (STAR) for addressing the problem of misaligned local regions between source and target domain, which finds that using more classifiers leads to better performance, but also introduces more model parameters, therefore risking overfitting.

    185
  • Task Residual for Tuning Vision-Language Models

    Tao Yu, Zhihe Lu, Xin Jin, Zhibo Chen, Xinchao Wang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    A new efficient tuning approach for VLMs named Task Residual Tuning (TaskRes), which performs directly on the text-based classifier and explicitly decouples the prior knowledge of the pre-trained models and new knowledge regarding a target task.

    182
  • Geometry Guided Adversarial Facial Expression Synthesis

    Lingxiao Song, Zhihe Lu, Ran He, Zhenan Sun, Tieniu Tan

    ACM International Conference on Multimedia (ACM MM) · 2018

    A Geometry-Guided Generative Adversarial Network (G2-GAN) for continuously-adjusting and identity-preserving facial expression synthesis and can generate compelling perceptual results on different expression editing tasks.

    165
  • GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph

    Xin Li, Dongze Lian, Zhihe Lu, Jiawang Bai, Zhibo Chen, Xinchao Wang

    arXiv · 2023

    An effective adapter-style tuning strategy, dubbed GraphAdapter, which performs the textual adapter by explicitly modeling the dual-modality structure knowledge with a dual knowledge graph, which significantly outperforms previous adapter-based methods.

    138
  • Can SAM Boost Video Super-Resolution?

    Zhihe Lu, Zeyu Xiao, Jiawang Bai, Zhiwei Xiong, Xinchao Wang

    arXiv · 2023

    This paper investigates a more robust and semantic-aware prior for enhanced VSR by utilizing the Segment Anything Model (SAM), a powerful foundational model that is less susceptible to image degradation.

    39
  • Uncertainty-Aware Source-Free Domain Adaptive Semantic Segmentation

    Zhihe Lu, Da Li, Yi-Zhe Song, Tao Xiang, Timothy M. Hospedales

    IEEE Transactions on Image Processing · 2023

    This paper proposes to use Bayesian Neural Network (BNN) to improve the target self- training by better estimating and exploiting pseudo-label uncertainty and introduces two novel self-training based components: Uncertainty-aware Online Teacher-Student Learning (UOTSL) and Uncertainly-aware FeatureMix (UFM).

    35
  • Prediction Calibration for Generalized Few-Shot Semantic Segmentation

    Zhihe Lu, Sen He, Da Li, Yi-Zhe Song, Tao Xiang

    IEEE Transactions on Image Processing · 2023

    To make the proposed cross-attention module training tractable at the pixel level, this module is designed based on feature-score cross-covariance and episodically trained to be generalizable at inference time.

    35
  • Recent Progress of Face Image Synthesis

    Zhihe Lu, Zhihang Li, Jie Cao, Ran He, Zhenan Sun

    IAPR Asian Conference on Pattern Recognition (ACPR) · 2017

    A comprehensive review of typical face synthesis works that involve traditional methods as well as advanced deep learning approaches is provided, particularly, Generative Adversarial Net (GAN) is highlighted to generate photo-realistic and identity preserving results.

    35
  • P-BiC: Ultra-High-Definition Image Moiré Patterns Removal via Patch Bilateral Compensation

    Zeyu Xiao, Zhihe Lu, Xinchao Wang

    ACM International Conference on Multimedia (ACM MM) · 2024

    A novel patch bilateral compensation network (P-BiC) for the demoire pattern removal in UHD images, which is memory-efficient and prior-knowledge-based, and quantitatively and qualitatively evaluates the effectiveness of P-BiC with extensive experiments.

    32
  • A Dive into SAM Prior in Image Restoration

    Zeyu Xiao, Jiawang Bai, Zhihe Lu, Zhiwei Xiong

    arXiv · 2023

    This paper uses the prior knowledge of the state-of-the-art segment anything model (SAM) to boost the performance of existing IR networks in an parameter-efficient tuning manner and proposes a lightweight SAM prior tuning (SPT) unit, which allows us to effectively integrate semantic priors into existingIR networks, resulting in significant improvements in restoration quality.

    32
  • Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation

    Keji He, Chenyang Si, Zhihe Lu, Yan Huang, Liang Wang, Xinchao Wang

    neural information processing systems · 2023

    This study first explores the significance of high-frequency information in VLN and provides evidence that it is instrumental in bolstering visual-textual matching processes, and proposes a sophisticated and versatile Frequency-enhanced Data Augmentation technique to improve the VLN model’s capability of capturing critical high-frequency information.

    32
  • Dream-Tac: A Unified Tactile World Action Model for Contact-Rich Robot Manipulation

    Yunfan Lou, Yifan Ye, Yankai Fu, Jun Cen, Xiaowei Chi, Yaoxu Lyu, Peidong Jia, Sirui Han, +2 more

    arXiv · 2026

    This paper proposes Dream-Tac, a unified Tactile-World Action Model that jointly models actions, future visual observations, and tactile dynamics, and introduces contact-gated visuotactile fusion to selectively integrate tactile signals and a contact-aware attention bias to better regulate cross-modal interactions during manipulation.

    30
  • Memory-Adaptive Vision-and-Language Navigation

    Keji He, Ya Jing, Yan Huang, Zhihe Lu, Dong Un An, Liang Wang

    Pattern Recognition · 2024

    28
  • Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models

    Zhihe Lu, Jiawang Bai, Xin Li, Zeyu Xiao, Xinchao Wang

    arXiv · 2023

    This work introduces three customized ensemble strategies, each tailored to one specific scenario, and introduces the zero-shot ensemble, automatically adjusting the logits of different models based on their confidence when only pre-trained VLMs are available.

    24
  • Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

    Jiajun Li, Tiecheng Guo, Yifan Ye, Rongyu Zhang, Xiaowei Chi, Qianpu Sun, Ying Wai Li, Yunfan Lou, +4 more

    arXiv · 2026

    Efficient-WAM is introduced, a World-Action Model motivated by the idea that a compact future branch can still support effective action generation, even at lower visual fidelity, and achieves 98 ms per action chunk during physical deployment, over 30x faster than the Motus baseline with comparable task success.

    19
  • Conditional Expression Synthesis with Face Parsing Transformation

    Zhihe Lu, Tanhao Hu, Lingxiao Song, Zhaoxiang Zhang, Ran He

    ACM International Conference on Multimedia (ACM MM) · 2018

    A Couple-Agent Face Parsing based Generative Adversarial Network (CAFP-GAN) that unites the knowledge of facial semantic regions and controllable expression signals is introduced that has the compelling ability on continuous expression synthesis.

    19
  • Task-to-Instance Prompt Learning for Vision-Language Models at Test Time

    Zhihe Lu, Jiawang Bai, Xin Li, Zeyu Xiao, Xinchao Wang

    IEEE Transactions on Image Processing · 2025

    17
  • Person identification from lip texture analysis

    Zhihe Lu, Xiang Hu Wu, Ran He

    IEEE International Conference on Digital Signal Processing (DSP) · 2016

    11
  • Data Pyramid for Embodied Manipulation: A Survey

    Yifan Ye, Yankai Fu, Yaoxu Lv, Bohan Hou, Jun Cen, Lingdong Kong, Zheng, Duo, Tianxing Chen, +21 more

    arXiv · 2026

    This work organizes the embodied data ecosystem as a pyramidspanning five complementary sources: real-robot data, UMI-style data, egocentric and exocentric data, simulation data, and general vision-language data, and further characterize each source in terms of data quality, diversity, reusability, and physical fidelity.

    8
  • Unraveling Motion Uncertainty for Local Motion Deblurring

    Zeyu Xiao, Zhihe Lu, Michael Bi Mi, Zhiwei Xiong, Xinchao Wang

    ACM International Conference on Multimedia (ACM MM) · 2024

    This work proposes a novel method named Motion-Uncertainty-Guided Network (MUGNet), which harnesses a probabilistic representational model to explicitly address the intricacies stemming from motion uncertainties and demonstrates the superiority of the MUGNet with extensive experiments.

    8
  • Token Expand-Merge: Training-Free Token Compression for Vision-Language-Action Models

    Yifan Ye, Jiaqi Ma, Jun Cen, Zhihe Lu

    IEEE Robotics and Automation Letters · 2026

    Token Expand-and-Merge-VLA (TEAM-VLA), a training-free token compression framework that accelerates VLA inference while preserving task performance, and achieves a balanced trade-off between efficiency and effectiveness.

    7
  • Parameter-Efficient and Memory-Efficient Tuning for Vision Transformer: A Disentangled Approach

    Taolin Zhang, Jiawang Bai, Zhihe Lu, Dongze Lian, Genping Wang, Xinchao Wang, Shu‐Tao Xia

    Lecture notes in computer science · 2024

    4
  • Self-Evolved Imitation Learning in Simulated World

    Yifan Ye, Jun Cen, Jing Chen, Zhihe Lu

    IEEE Robotics and Automation Letters · 2026

    Self-Evolved Imitation Learning (SEIL), a framework that progressively improves a few-shot model through simulator interactions, and introduces a lightweight selector that filters complementary and informative trajectories from the generated pool to ensure demonstration quality.

    3
  • COLA: Context-Aware Language-Driven Test-Time Adaptation

    Aiming Zhang, Tianyuan Yu, Liang Bai, Jun Tang, Yanming Guo, Yirun Ruan, Yun Zhou, Zhihe Lu

    IEEE Transactions on Image Processing · 2025

    This paper investigates a more general source model capable of adaptation to multiple target domains without needing shared labels, achieved by using a pre-trained vision-language model (VLM) that can recognize images through matching with class descriptions.

    3

Show all 35 works

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.