Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Anh Tuan LuuSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works36 from public data
- Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness Predictions20
This work proposes Multi-modal Contrastive Learning for Multimodal Review Helpfulness Prediction (MRHP) problem, concentrating on mutual information between input modalities to explicitly elaborate cross-modAL relations, and introduces Adaptive Weighting scheme for the contrastive learning approach in order to increase flexibility in optimization.
- 18
- 14
- MAMA: Meta-optimized Angular Margin Contrastive Framework for Video-Language Representation Learning12
To adapt to the non-uniform concept distribution, MAMA utilizes a multi-layer perceptron (MLP)-parameterized weighting function that maps loss values to sample weights which enable dynamic adjustment of the model's focus throughout the training.
- DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding12
An energy-based model framework to explicitly learn moment- query distributions is proposed and DemaFormer, a novel Transformer-based architecture that utilizes exponential moving average with a learnable damping factor to effectively encode moment-query inputs is proposed.
- 10
- 10
- Enhancing Multimodal Entity Linking with Jaccard Distance-based Conditional Contrastive Learning and Contextual Visual Augmentation8
JD-CCL (Jaccard Distance-based Conditional Contrastive Learning), a novel approach designed to enhance the ability to match multimodal entity linking models, and CVaCPT (Contextual Visual-aid Controllable Patch Transform), a novel method that enhances visual representations by incorporating multi-view synthetic images and contextual textual representations to scale and shift patch representations.
- Curriculum Demonstration Selection for In-Context Learning7
Curriculum Demonstration Selection (CDS), a novel demonstration selection method for ICL, instead of merely using similarity, additionally partitions samples by their complexity measurements, enabling LLMs to learn from varied complexities within the training set.
- Expand BERT Representation with Visual Information via Grounded Language Learning with Multimodal Partial Alignment7
The proposed GroundedBERT is a grounded language learning method that enhances the BERT representation with visually grounded information and significantly outperforms the baseline language models on various language tasks of the GLUE and SQuAD datasets.
- 6
- CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base4
Comprehensive evaluations on benchmark datasets, including POPE and R-Bench, demonstrate that CutPaste\&Find achieves competitive hallucination detection performance while being significantly more efficient and cost-effective than previous methods.
- READ-PVLA: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling4
A novel REcurrent ADapter (READ) that employs recurrent computation to enable temporal modeling capability and Partial Video-Language Alignment (PVLA) objective via the use of partial optimal transport to maintain task-related information flowing into the authors' READ modules is proposed.
- 4
- 3
- A Comparative Analysis of Contextual Representation Flow in State-Space and Transformer Architectures3
This work presents the first unified, token- and layer-wise analysis of representation propagation in SSMs and TBMs, finding a key divergence: TBMs rapidly homogenize token representations, with diversity reemerging only in later layers, while SSMs preserve token uniqueness early but converge to homogenization deeper.
- 3
- From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models2
The results show that raw hidden states and attention features are difficult to outperform in-domain, however, under distribution shift, structured and compressed features are more robust, suggesting that in-domain performance alone is insufficient to measure progress.
- More Bias, Less Bias: BiasPrompting for Enhanced Multiple-Choice Question Answering2
A novel inference framework that guides LLMs to generate and critically evaluate reasoning across all plausible answer options before reaching a final prediction, and demonstrates significant improvements in five widely used multiple-choice question answering benchmarks.
- TIGER: Text-Conditioned Visual Gated Routing with Acceptance Alignment for Multimodal Speculative Decoding2
TIGER, a Text-conditioned vIsual GatEd Routing framework for multimodal speculative decoding, is proposed, which yields consistent gains in accepted prefix length and speculative speedup under exact verifier-side speculative decoding, while achieving favorable quality-latency trade-offs with comparable downstream accuracy in visual-routing analyses.
- Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding2
This work conducts a thorough empirical study to demystify crucial components that influence the temporal understanding of LVLMs and proposes a temporal-oriented recipe that encompasses temporal-oriented training schemes and an upscaled interface.
- Vision-and-Language Pretraining2
This article categorize and delineate pretraining approaches, along with the summary of state-of-the-art vision-and-language pretrained models, and a list of training datasets and downstream tasks is supplied to further polish the perspective into V\&L pretraining.
- Don't Read Everything: A Curvature-Conditioned Query for Linear Attention1
Curvature-Conditioned Query modifies only the read step and is composable with any linear-attention backbone, and improves perplexity, zero-shot downstream accuracy, S-NIAH retrieval at and beyond the training context, length-extrapolation perplexity from 4K to 20K, and LongBench accuracy.
- 1
- When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for Vision-Language Spatial Reasoning–
This work proposes Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.