Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Tat‐Seng ChuaSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works27 from public data
- Dynamic Modality Interaction Modeling for Image-Text Retrieval230
A novel modality interaction modeling network based upon the routing mechanism, which is the first unified and dynamic multimodal interaction framework towards image-text retrieval and demonstrates superiority compared with several state-of-the-art baselines.
- Context-Aware Multi-View Summarization Network for Image-Text Matching165
A novel context-aware multi-view summarization network to summarize context-enhanced visual region information from multiple views and designs an adaptive gating self-attention module to extract representations of visual regions and words.
- LayoutLLM-T2I: Eliciting Layout Guidance from LLM for Text-to-Image Generation152
This work strives to synthesize high-fidelity images that are semantically aligned with a given textual prompt without any guidance and proposes a coarse-to-fine paradigm to achieve layout planning and image generation.
- Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization94
A unified learning approach to simultaneously modeling the coarse- and fine-grained retrieval by considering the multi-grained uncertainty is introduced, which prevents the model from pushing away potential candidates in the early stage, and thus improves the recall rate.
- Search-oriented Micro-video Captioning42
A large-scale multimodal pre-training network regularized by five tasks to strengthen the downstream video representation and a flow-based diverse captioning model to generate different captions from consumers' search demand is presented.
- Learnable Pillar-based Re-ranking for Image-Text Retrieval25
This paper designs a neighbor-aware graph reasoning module to flexibly exploit the relations and excavate the sparse positive items within a neighborhood, and presents a structure alignment constraint to promote crossmodal collaboration and align the asymmetric modalities.
- SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation21
This work introduces a model-agnostic iterative self-improvement framework (SILMM) that can enable LMMs to provide helpful and scalable self-feedback and optimize text-image alignment via Direct Preference Optimization (DPO).
- Popularity-aware Distributionally Robust Optimization for Recommendation System19
This work proposes a novel Popularity- aware Distributionally Robust Optimization (PDRO) framework, which emphasizes the optimization of sparse users/items, while incorporating item popularity to preserve the performance of popular items through two modules.
- VINCIE: Unlocking In-context Image Editing from Video18
This work introduces a scalable approach to annotate videos as interleaved multimodal sequences and designs a block-causal diffusion transformer trained on three proxy tasks: next-image prediction, current segmentation prediction, and next-segmentation prediction.
- Discriminative Probing and Tuning for Text-to-Image Generation18
A discriminative adapter built on T2I models is presented to probe their discriminative abilities on two representative tasks and leverage discriminative fine-tuning to improve their text-image alignment.
- 14
- 12
- DanceOPD: On-Policy Generative Field Distillation11
DanceOPD is introduced, an on-policy generative field distillation framework for flow-matching models that routes each sample to one capability field, queries one low-noise student-induced state, and trains with a simple velocity MSE objective, establishing a practical route for generative field distillation in flow-matching models.
- Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation11
This study proposes AVG, which discretizes images into vokens while aligning with both the visual information and high-level semantics, and incorporates discriminative training to modify the learning direction during token-to-voken training.
- TTOM: Test-Time Optimization and Memorization for Compositional Video Generation7
This work introduces Test-Time Optimization and Memorization (TTOM), a training-free framework that aligns VFM outputs with spatiotemporal layouts during inference for better text-image alignment in compositional scenarios.
- GenArena: How Can We Achieve Human-Aligned Evaluation for Visual Generation Tasks?6
GenArena is introduced, a unified evaluation framework that leverages a pairwise comparison paradigm to ensure stable and human-aligned evaluation, and uncovers a transformative finding that simply adopting this pairwise protocol enables off-the-shelf open-source models to outperform top-tier proprietary models.
- WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval6
WISER is a training-free framework that unifies T2I and I2I via a "retrieve-verify-refine" pipeline, explicitly modeling intent awareness and uncertainty awareness, and significantly outperforms previous methods across multiple benchmarks.
- ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval3
ReCALL is proposed, a model-agnostic framework that follows a diagnose-generate-refine pipeline that consistently recalibrates degraded capabilities and achieves state-of-the-art performance.
- AUHead: Realistic Emotional Talking Head Generation via Action Units Control2
This work introduces a novel two-stage method (AUHead) to disentangle fine-grained emotion control, i.e, Action Units (AUs), from audio and achieve controllable generation, and proposes an AU-driven controllable diffusion model that synthesizes realistic talking-head videos conditioned on AU sequences.
- Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos2
Unlike the preference observed in image modalities, it is found that video retrieval bias arises from both unseen visual and temporal information, making the root causes of video bias a complex interplay of these two factors.
- Optimizing Visual Generative Models via Distribution-wise Rewards1
A novel framework that finetunes generative models using distribution-wise rewards, ensuring better alignment with real-world data distributions is presented, and a subset-replace strategy that efficiently provides reward signals by updating only a small subset of a generated reference set is introduced.
- 1
- –
- OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026–
This report proposes a reranking-based framework that effectively leverages the strong video-language reasoning capability of multimodal large language model (MLLM) while preserving the efficiency and candidate recall of conventional localization pipelines.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.