Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works12 from public data

TitleCited by
  • MLVU: Benchmarking Multi-task Long Video Understanding

    Junjie Zhou, Yan Shu, Bo Zhao, Boya Wu, Zhengyang Liang, Shitao Xiao, Minghao Qin, Xi Yang, +4 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    A new benchmark called MLVU (Multitask Long Video Understanding Benchmark) is proposed for the comprehensive and in-depth evaluation of LVU, which suggests that factors such as context length, image-understanding ability, and the choice of LLM backbone can play critical roles in future advancements.

    294
  • Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

    Yan Shu, Zheng Liu, Peitian Zhang, Minghao Qin, Junjie Zhou, Zhengyang Liang, Tiejun Huang, Bo Ya Zhao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This work proposes Video-XL, a novel approach that leverages MLLMs’ inherent key-value (KV) sparsification capacity to condense the visual input, and introduces a new special token, the Visual Summarization Token (VST), for each interval of the video, which summarizes the visual information within the interval as its associated KV.

    288
  • Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

    Qin, Minghao, Xiangrui Liu, Zhengyang Liang, Yan Shu, Yuan, Huaying, Zhou, Juenjie, Shitao Xiao, Bo Zhao, +1 more

    arXiv · 2025

    Video-XL-2 is proposed, a novel MLLM that delivers superior cost-effectiveness for long-video understanding based on task-aware KV sparsification and achieves state-of-the-art performance on various long video understanding benchmarks, outperforming existing open-source lightweight models.

    43
  • Unveiling the Ignorance of MLLMs: Seeing Clearly, Answering Incorrectly

    Yexin Liu, Zhengyang Liang, Yueze Wang, Xianfeng Wu, Feilong Tang, Muyang He, Jian Li, Zheng Liu, +3 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    This investigation identifies two primary issues: 1) most instruction tuning datasets predominantly feature questions that "directly" relate to the visual content, leading to a bias in MLLMs’ responses to other indirect questions, and 2) MLLMs’ attention to visual tokens is notably lower than to system and question tokens.

    34
  • Self-Supervised Multi-Modal Knowledge Graph Contrastive Hashing for Cross-Modal Search

    Meiyu Liang, Junping Du, Zhengyang Liang, Yongwang Xing, Wei Huang, Zhe Xue

    Proceedings of the AAAI Conference on Artificial Intelligence · 2024

    A novel self-supervised multi-grained multi-modal knowledge graph contrastive hashing method for cross-modal search that outperforms the state-of-the-art methods and fuses the global coarse-grained and local fine-grained embeddings by multihead attention mechanism for inter-modal and intra-modal contrastive learning.

    30
  • MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval

    Huaying Yuan, Jian Ni, Zheng Liu, Yueze Wang, Junjie Zhou, Zhengyang Liang, Bo Zhao, Zhao Cao, +2 more

    neural information processing systems · 2025

    The results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning.

    15
  • Video-Browser: Towards Agentic Open-web Video Browsing

    Zhengyang Liang, Yan Shu, Xiangrui Liu, Minghao Qin, Kaixin Liang, Nicu Sebe, Zheng Liu, Lizi Liao

    arXiv · 2025

    Video-BrowseComp is presented, a challenging benchmark comprising 210 questions tailored for open-web agentic video reasoning that advances the field beyond passive perception toward proactive video reasoning and reveals a critical bottleneck in processing the web’s most dynamic modality: video.

    4
  • A Hypothesis for the Aesthetic Appreciation in Neural Networks

    Cheng Xu, Xin Wang, Haotian Xue, Zhengyang Liang, Quanshi Zhang

    arXiv · 2021

    4
  • Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval

    Zheng Liu, Ze Liu, Zhengyang Liang, Junjie Zhou, Shitao Xiao, Chao Gao, Chen Zhang, Defu Lian

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2025

    3
  • Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning

    Zhengyang Liang, Meiyu Liang, Wei Huang, Yawen Li, Zhe Xue

    arXiv · 2024

    This work proposes a novel dynamic self-adaptive multiscale distillation from pre-trained multimodal large model for efficient cross-modal representation learning for the first time, and utilizes only image-level information to achieve state-of-the-art performance on cross-modal retrieval tasks.

    3
  • DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning

    Huiyan Lin, Yan Shu, Zhengyang Liang, Chi Liu, Xiangrui Liu, Minghao Qin, Derek Li, Bryan Dai, +3 more

    arXiv · 2026

    DyCo-RL, which integrates dynamic cross-modal coordination into RLVR optimization and uses the Fisher-Rao geodesic distance to measure within-modality attention shifts, assigning tokens to either visually-oriented or text-oriented functional roles and evaluates the alignment between a token's actual attention allocation and its assigned role.

    1
  • VideoCreator: An Agentic System for Multi-turn Video Production

    Zhengyang Liang, Yan Shu, Cathal G. Gurrin, Nicu Sebe, Lizi Liao

    International Conference on Multimedia Retrieval · 2026

    VideoCreator is a unified video agent that integrates generation and understanding with a project-level memory system and uses persistent memory to retain and reuse prior context across turns, enabling continuous multi-round creation with consistency throughout the project.

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.