Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Tat‐Seng Chua

    Possible advisor · last author on 9 of their early first-author papers, 2023–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works27 from public data

TitleCited by
  • Dynamic Modality Interaction Modeling for Image-Text Retrieval

    Leigang Qu, Meng Liu, Jianlong Wu, Zan Gao, Liqiang Nie

    International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) · 2021

    A novel modality interaction modeling network based upon the routing mechanism, which is the first unified and dynamic multimodal interaction framework towards image-text retrieval and demonstrates superiority compared with several state-of-the-art baselines.

    230
  • Context-Aware Multi-View Summarization Network for Image-Text Matching

    Leigang Qu, Meng Liu, Da Cao, Liqiang Nie, Qi Tian

    ACM International Conference on Multimedia (ACM MM) · 2020

    A novel context-aware multi-view summarization network to summarize context-enhanced visual region information from multiple views and designs an adaptive gating self-attention module to extract representations of visual regions and words.

    165
  • LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation

    Leigang Qu, Shengqiong Wu, Hao Fei, Liqiang Nie, Tat‐Seng Chua

    ACM International Conference on Multimedia (ACM MM) · 2023

    This work strives to synthesize high-fidelity images that are semantically aligned with a given textual prompt without any guidance and proposes a coarse-to-fine paradigm to achieve layout planning and image generation.

    152
  • Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization

    Yiyang Chen, Zhedong Zheng, Wei Ji, Leigang Qu, Tat‐Seng Chua

    arXiv · 2022

    A unified learning approach to simultaneously modeling the coarse- and fine-grained retrieval by considering the multi-grained uncertainty is introduced, which prevents the model from pushing away potential candidates in the early stage, and thus improves the recall rate.

    94
  • Search-oriented Micro-video Captioning

    Liqiang Nie, Leigang Qu, Dai Meng, Min Zhang, Qi Tian, Alberto Del Bimbo

    ACM International Conference on Multimedia (ACM MM) · 2022

    A large-scale multimodal pre-training network regularized by five tasks to strengthen the downstream video representation and a flow-based diverse captioning model to generate different captions from consumers' search demand is presented.

    42
  • Learnable Pillar-based Re-ranking for Image-Text Retrieval

    Leigang Qu, Meng Liu, Wenjie Wang, Zhedong Zheng, Liqiang Nie, Tat‐Seng Chua

    International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) · 2023

    This paper designs a neighbor-aware graph reasoning module to flexibly exploit the relations and excavate the sparse positive items within a neighborhood, and presents a structure alignment constraint to promote crossmodal collaboration and align the asymmetric modalities.

    25
  • SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation

    Leigang Qu, Haochuan Li, Wenjie Wang, Xiang Liu, Juncheng Li, Liqiang Nie, Tat‐Seng Chua

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work introduces a model-agnostic iterative self-improvement framework (SILMM) that can enable LMMs to provide helpful and scalable self-feedback and optimize text-image alignment via Direct Preference Optimization (DPO).

    21
  • Popularity-aware Distributionally Robust Optimization for Recommendation System

    Jujia Zhao, Wenjie Wang, Xinyu Lin, Leigang Qu, Jizhi Zhang, Tat‐Seng Chua

    ACM International Conference on Information and Knowledge Management (CIKM) · 2023

    This work proposes a novel Popularity- aware Distributionally Robust Optimization (PDRO) framework, which emphasizes the optimization of sparse users/items, while incorporating item popularity to preserve the performance of popular items through two modules.

    19
  • VINCIE: Unlocking In-context Image Editing from Video

    Leigang Qu, Cheng, Feng, Ziyan Yang, Qi Zhao, Shanchuan Lin, Yichun Shi, Yicong Li, Wenjie Wang, +2 more

    arXiv · 2025

    This work introduces a scalable approach to annotate videos as interleaved multimodal sequences and designs a block-causal diffusion transformer trained on three proxy tasks: next-image prediction, current segmentation prediction, and next-segmentation prediction.

    18
  • Discriminative Probing and Tuning for Text-to-Image Generation

    Leigang Qu, Wenjie Wang, Yongqi Li, Hanwang Zhang, Liqiang Nie, Tat‐Seng Chua

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    A discriminative adapter built on T2I models is presented to probe their discriminative abilities on two representative tasks and leverage discriminative fine-tuning to improve their text-image alignment.

    18
  • Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

    Thong Thanh Nguyen, Yi Bin, Junbin Xiao, Leigang Qu, Yicong Li, Jay Zhangjie Wu, Cong-Duy T Nguyen, See-Kiong Ng, +1 more

    Findings of the Association for Computational Linguistics ACL 2024 · 2024

    14
  • Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond

    Yongqi Li, Wenjie Wang, Leigang Qu, Liqiang Nie, Wenjie Li, Tat‐Seng Chua

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2024

    12
  • DanceOPD: On-Policy Generative Field Distillation

    Wei Zhou, Xiongwei Zhu, Zhewei Xu, Bo Dong, Lixue Gong, Y Liang, Meng Chu, Leigang Qu, +3 more

    arXiv · 2026

    DanceOPD is introduced, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective, establishing a practical route for generative field distillation in flow-matching models.

    11
  • Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation

    Yongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu, Yinwei Wei, Wenjie Li, Liqiang Nie, Tat‐Seng Chua

    International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) · 2025

    This study proposes AVG, which discretizes images into vokens while aligning with both the visual information and high-level semantics, and incorporates discriminative training to modify the learning direction during token-to-voken training.

    11
  • TTOM: Test-Time Optimization and Memorization for Compositional Video Generation

    Leigang Qu, Ziyang Wang, Na Zheng, Wenjie Wang, Liqiang Nie, Tat‐Seng Chua

    arXiv · 2025

    This work introduces Test-Time Optimization and Memorization (TTOM), a training-free framework that aligns VFM outputs with spatiotemporal layouts during inference for better text-image alignment in compositional scenarios.

    7
  • GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?

    Ruihang Li, Leigang Qu, Jingxu Zhang, Dongnan Gui, Mengde Xu, Xiaosong Zhang, Han Hu, Wenjie Wang, +1 more

    arXiv · 2026

    GenArena is introduced, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation, and uncovers a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models.

    6
  • WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval

    Tianyue Wang, Leigang Qu, Tianyu Yang, Xiangzhao Hao, Yifan Xu, Haiyun Guo, Jinqiao Wang

    arXiv · 2026

    WISER is a training-free framework that unifies T2I and I2I via a "retrieve-verify-refine" pipeline, explicitly modeling intent awareness and uncertainty awareness, and significantly outperforms previous methods across multiple benchmarks.

    6
  • ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval

    Tianyu Yang, ChenWei He, Xiangzhao Hao, Tianyue Wang, Jiarui Guo, Haiyun Guo, Leigang Qu, Jinqiao Wang, +1 more

    arXiv · 2026

    ReCALL is proposed, a model-agnostic framework that follows a diagnose-generate-refine pipeline that consistently recalibrates degraded capabilities and achieves state-of-the-art performance.

    3
  • AUHead: Realistic Emotional Talking Head Generation via Action Units Control

    Jiayi Lyu, Leigang Qu, Wenjing Zhang, Hanyu Jiang, Kai Liu, Zhenglin Zhou, Xiaobo Xia, Jian Xue, +1 more

    arXiv · 2026

    This work introduces a novel two-stage method (AUHead) to disentangle fine-grained emotion control, i.e, Action Units (AUs), from audio and achieve controllable generation, and proposes an AU-driven controllable diffusion model that synthesizes realistic talking-head videos conditioned on AU sequences.

    2
  • Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos

    H. Gao, Liang Pang, Shicheng Xu, Leigang Qu, Tat‐Seng Chua, Huawei Shen, Xueqi Cheng

    ACM International Conference on Multimedia (ACM MM) · 2025

    Unlike the preference observed in image modalities, it is found that video retrieval bias arises from both unseen visual and temporal information, making the root causes of video bias a complex interplay of these two factors.

    2
  • Optimizing Visual Generative Models via Distribution-wise Rewards

    R K Li, Mengde Xu, Shuyang Gu, Leigang Qu, Fuli Feng, Han Hu, Wenjie Wang

    arXiv · 2026

    A novel framework that finetunes generative models using distribution-wise rewards, ensuring better alignment with real-world data distributions is presented, and a subset-replace strategy that efficiently provides reward signals by updating only a small subset of a generated reference set is introduced.

    1
  • Enhanced Structured Lasso Pruning with Class-wise Information

    Xiang Liu, Mingchen Li, Xia Li, Leigang Qu, Wang, Guansu, Zifan Peng, Yijun Song, Zemin Liu, +2 more

    arXiv · 2025

    1
  • On-Policy Self-Distillation in Diffusion Models

    Wei Zhou, Xiongwei Zhu, Lingdong Kong, Bo Chen, Lei Zhang, Yongyuan Liang, Xiaoxia Hou, Ye Tian, +9 more

    arXiv · 2026

    –
  • OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026

    Yisen Feng, Leigang Qu, Haoyu Zhang, Qiaohui Chu, Meng Liu, Xuemeng Song, Weili Guan, Liqiang Nie

    arXiv · 2026

    This report proposes a reranking-based framework that effectively leverages the strong video-language reasoning capability of multimodal large language model (MLLM) while preserving the efficiency and candidate recall of conventional localization pipelines.

    –
  • Routing Before Looking: Query-Adaptive Evidence Acquisition for Long-form Video Understanding

    Tianyue Wang, Xuying Wu, Yuxiang Ma, Ruiming Liang, Jiaxuan Kang, Yanchao Hao, Zheng Wei, Leigang Qu, +2 more

    arXiv · 2026

    –
  • Visual Content Generation in the Era of Large Foundation Models

    Leigang Qu, Fei Shen, Zhenglin Zhou, Jiayi Lyu, Wenjie Wang, Lu Jiang

    International Conference on Multimedia Retrieval · 2025

    This tutorial will provide an in-depth exploration of the state-of-the-art techniques and methodologies used in visual content generation, emphasizing the role of large-scale generative models.

    –
  • TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models

    Leigang Qu, Haochuan Li, Tan Wang, Wenjie Wang, Yongqi Li, Liqiang Nie, Tat‐Seng Chua

    arXiv · 2024

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it.

Sign in to report