Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Alex Chichung KotSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works22 from public data
- Towards Photo-Realistic Virtual Try-On by Adaptively Generating↔Preserving Image Content347
This work proposes a novel visual try-on network, namely Adaptive Content Generating and Preserving Network (ACGPN), which can generate photo-realistic images with much better perceptual quality and richer fine-details.
- 239
- DualCheXNet: dual asymmetric feature learning for thoracic disease classification in chest X-rays104
A novel dual asymmetric feature learning network named DualCheXNet is presented for multi-label thoracic disease classification in CXRs and an iterative training strategy is designed to integrate the loss contribution of the involved classifiers into a unified loss, and optimize the process of complementary features learning in an alternative way.
- 64
- Audio-Visual Deception Detection: DOLOS Dataset and Parameter-Efficient Crossmodal Learning44
This work introduces DOLOS 1, the largest gameshow deception detection dataset with rich deceptive conversations, and proposes Parameter-Efficient Crossmodal Learning (PECL), where a Uniform Temporal Adapter (UT-Adapter) explores temporal attention in transformer-based architectures, and a crossmodal fusion module, Plug-in Audio-Visual Fusion (PAVF), combines crossmodal information from audio-visual features.
- 20
- Benchmarking Cross-Domain Audio-Visual Deception Detection16
This work presents the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios and proposes an algorithm to enhance the generalization performance by maximizing the gradient inner products between modality encoders, named “MM-IDGM”.
- 15
- Flexible-Modal Deception Detection with Audio-Visual Adapter12
An advanced Transformer-based framework complemented by an Audio-Visual Adapter (AVA) integrating temporal features from both audio and visual modalities and an innovative multi-modal contrastive learning method that is designed to enhance the correlation between uni-modal features and their integrated counterparts within a consistent feature space are introduced.
- 11
- 7
- Improving Concept Alignment in Vision-Language Concept Bottleneck Models5
This work proposes a novel Contrastive Semi-Supervised (CSS) learning method that leverages a few labeled concept samples to activate truthful visual concepts and improve concept alignment in the CLIP model to improve the classification performance.
- Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion4
Experimental results on MOSI and MOSEI datasets illustrated that the proposed UCRN outperforms recent state-of-the-art techniques and its robustness is highly preferred in real multimodal sequence fusion scenarios.
- 4
- 3
- SparseMamba-PCL: Scribble-Supervised Medical Image Segmentation via SAM-Guided Progressive Collaborative Learning2
A Progressive Collaborative Learning framework that leverages novel algorithms and the Med-SAM foundation model to enhance information quality during training and introduces a Sparse Mamba network, which is highly capable of capturing local and global dependencies.
- MIMIC: Mask Image Pre-training with Mix Contrastive Fine-tuning for Facial Expression Recognition2
This work introduces a novel FER training paradigm named Mask Image pre-training with MIx Contrastive fine-tuning (MIMIC), and demonstrates that the MIMIC outperforms the previous training paradigm, showing its capability to learn better representations.
- –
- –
- –
- –
- Single Image Reflection Removal Based on Deep Residual Learning–
A novel generative adversarial framework is proposed, where the generator is embedded with the deep residual learning, significantly boosting the performance without impairing the intactness of the background by adversarial training.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.