Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Alvin ChanSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works19 from public data
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using Language53
This paper designs a novel clip search model that learns to identify promising video regions to search conditioned on the language query and introduces a set of low-cost semantic indexing features to capture the context of objects and interactions that suggest where to search the query-relevant moment.
- Not All Inputs Are Valid: Towards Open-Set Video Moment Retrieval using Language49
This paper proposes a novel model OpenVMR, which first distinguishes ID and OOD queries based on the normalizing flow technology, and then conducts moment retrieval based on ID queries, and first distinguishes ID and OOD queries based on the normalizing flow technology, and then conducts moment retrieval based on ID queries.
- Rethinking Weakly-Supervised Video Temporal Grounding From a Game Perspective42
This is the first attempt to tackle the challenging task of weakly-supervised video temporal grounding from a novel game perspective, which effectively learns the uncertain relationship between each vision-language pair with diverse granularity and flexible combination for multi-level cross-modal interaction.
- Annotations Are Not All You Need: A Cross-modal Knowledge Transfer Network for Unsupervised Temporal Sentence Grounding38
By modulating and transferring both appearance and action knowledge into the authors' challenging unsupervised task, the model can directly utilize this general knowledge to correlate videos and queries, and accurately retrieve the relevant segment without training.
- Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security29
The Adversarial Prompt Disentanglement framework is proposed, a novel defense mechanism that proactively identifies and neutralizes malicious components in input prompts before they are processed by the LLM, and offers a scalable, ethically grounded defense against prompt-based adversarial threats.
- Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval26
Reaction-Diffusion Multimodal Fusion (RDMF) is proposed, a novel framework that reimagines video-language alignment as a reaction-diffusion process, drawing on the principles of pattern formation introduced by Alan Turing, and represents a pioneering interdisciplinary approach.
- Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal Optimization25
This paper introduces Multi-Modal Adversarial Synergy (MMAS), a groundbreaking framework that crafts universal, black-box multi-modal attacks against LVLMs, and generates a texture scale-constrained Universal Adversarial Perturbation (UAP) for images and a learnable prompt perturbation for text, optimized jointly using only model queries.
- Hierarchical Semantic-Augmented Navigation: Optimal Transport and Graph-Driven Reasoning for Vision-Language Navigation24
Hierarchical Semantic-Augmented Navigation achieves state-of-the-art performance, with significant improvements in navigation success and generalization to unseen environments, by integrating spectral graph theory, optimal transport, and advanced multi-modal learning.
- Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition20
Vision-language analysis broadly corroborates previously reported decision-level PID patterns across tasks, models, interventions, and layers, and introduces Sensory PID, a conditional formulation that conditions on language and separates unique, redundant, and synergistic contributions from video and audio.
- Towards Unified Vision-Language Models with Incomplete Multi-Modal Inputs20
This work makes the first attempt to propose a unified incomplete video-language model to process the incomplete multi-modal inputs and shows that this method can serve as a plug-and-play module for previous works to improve their performance in various multi-modal tasks.
- Towards Robust Temporal Activity Localization Learning with Noisy Labels18
A novel method called Co-Teaching Regularizer (CTR) is proposed, which utilizes two models sharing the same framework to teach each other for robust learning and is significantly more robust to the noisy training data compared to the existing methods.
- Rethinking Video-Language Model from the Language Input Perspective17
A novel plug-and-play framework for various VLM-based methods to fully bridge videos and texts that generates positive and negative texts from the original ones to target specific text components and proposes an attribute-based text reasoning strategy to mine fine-grained textual semantics of generated texts.
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning17
CogniVerse is introduced, a novel MMRAG framework that addresses challenges through a cognitive-inspired, mathematically rigorous approach and significantly outperforms state-of-the-art systems in both accuracy and coherence, while reducing retrieval latency.
- Immuno-VLM: Immunizing Large Vision-Language Models via Generative Semantic Antibodies for Open-World Trustworthiness17
Departing from traditional Open-Set Recognition methods that rely on passive density estimation or inefficient pixel-space outlier generation, Immuno-VLM leverages the generative reasoning of Large Language Models to actively hallucinate ``Semantic Antibodies'', textual descriptions of near-distribution outliers that effectively bound the decision space of known classes.
- SLAP: The Semantic Least Action Principle for Variational Video-Language Modeling17
This work proposes a paradigm shift from probabilistic generation to variational mechanics with the Semantic Least Action Principle (SLAP), a rigorous isomorphism between classical mechanics and semantic dynamics, and naturally enforces object persistence without pixel-level rendering.
- 4
- 1
- How Creative Are Large Language Models in Generating Molecules?–
This work is the first to reframe the abilities required for molecule generation as creativity, providing a systematic understanding of creativity in LLM-based molecular generation and clarifying the appropriate use of LLMs in molecular discovery pipelines.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.