Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks12 from public data
- MLVU: Benchmarking Multi-task Long Video Understanding294
A new benchmark called MLVU (Multitask Long Video Understanding Benchmark) is proposed for the comprehensive and in-depth evaluation of LVU, which suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements.
- Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding288
This work proposes Video-XL, a novel approach that leverages MLLMs’ inherent key-value (KV) sparsification capacity to condense the visual input, and introduces a new special token, the Visual Summarization Token (VST), for each interval of the video, which summarizes the visual information within the interval as its associated KV.
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification43
Video-XL-2 is proposed, a novel MLLM that delivers superior cost-effectiveness for long-video understanding based on task-aware KV sparsification and achieves state-of-the-art performance on various long video understanding benchmarks, outperforming existing open-source lightweight models.
- Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly34
This investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that "directly" relate to the visual content, leading to a bias in MLLMs’ responses to other indirect questions, and 2) MLLMs’ attention to visual tokens is notably lower than to system and question tokens.
- Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search30
A novel self-supervised multi-grained multi-modal knowledge graph contrastive hashing method for cross-modal search that outperforms the state-of-the-art methods and fuses the global coarse-grained and local fine-grained embeddings by multihead attention mechanism for inter-modal and intra-modal contrastive learning.
- MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval15
The results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning.
- Video-Browser: Towards Agentic Open-web Video Browsing4
Video-BrowseComp is presented, a challenging benchmark comprising 210 questions tailored for open-web agentic video reasoning that advances the field beyond passive perception toward proactive video reasoning and reveals a critical bottleneck in processing the web’s most dynamic modality: video.
- 4
- 3
- Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning3
This work proposes a novel dynamic self-adaptive multiscale distillation from pre-trained multimodal large model for efficient cross-modal representation learning for the first time, and utilizes only image-level information to achieve state-of-the-art performance on cross-modal retrieval tasks.
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning1
DyCo-RL, which integrates dynamic cross-modal coordination into RLVR optimization and uses the Fisher-Rao geodesic distance to measure within-modality attention shifts, assigning tokens to either visually-oriented or text-oriented functional roles and evaluates the alignment between a token's actual attention allocation and its assigned role.
- VideoCreator: An Agentic System for Multi-turn Video Production–
VideoCreator is a unified video agent that integrates generation and understanding with a project-level memory system and uses persistent memory to retain and reuse prior context across turns, enabling continuous multi-round creation with consistency throughout the project.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.