Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks11 from public data
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models318
This work comprehensively analyzed multiple state-of-the-art VLA models and revealed consistent brittleness beneath apparent competence, challenging the assumption that high benchmark scores equate to true competency and highlighting the need for evaluation practices that assess reliability under realistic variation.
- World Action Models: The Next Frontier in Embodied AI52
Overall, this survey provides the first systematic account of the WAMs landscape, clarifies key architectural paradigms and their trade-offs, and identifies open challenges and future opportunities for this rapidly evolving field.
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding38
HerMES, a novel training-free architecture for real-time and accurate understanding of video streams that requires no auxiliary computations upon the arrival of user queries, guaranteeing real-time responses for continuous video stream interactions, achieves 10$\times$ faster TTFT compared to prior SOTA.
- CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs37
This work introduces a visual preference optimization module within the DPO framework, enabling MLLMs to learn from both textual and visual preferences simultaneously, and proposes a hierarchical textual preference optimization module that allows the model to capture preferences at multiple granular levels, including response, segment, and token levels.
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models17
Extensive experiments on benchmarks such as Video Hallucination, Video QA, and Captioning performance tasks demonstrate that VistaDPO significantly improves the performance of existing LVMs, effectively mitigating video-language misalignment and hallucination.
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs16
A training-free pruning framework that significantly outperforms existing token pruning methods and generalizes across diverse MLLMs is proposed, and insights further provide actionable guidelines for training efficient MLLMs by aligning model architecture with its intrinsic layer-wise processing dynamics.
- Hint-before-Solving Prompting: Guiding LLMs to Effectively Utilize Encoded Knowledge16
Hint-before-Solving Prompting (HSP), which guides the model to generate hints for solving the problem and then generate solutions containing intermediate reasoning steps to effectively improve the accuracy of reasoning tasks, is introduced.
- Analyzing Reasoning Consistency in Large Multimodal Models under Cross-Modal Conflicts6
This paper proposes the LogicGraph Perturbation Protocol, a structurally injects perturbations into the reasoning chains of diverse LMMs spanning both native reasoning architectures and prompt-driven paradigms to evaluate their self-reflection capabilities.
- 3
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.