Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Huchuan Lu

    Possible advisor · last author on 7 of their early first-author papers, 2021–2024

    Suggested from co-authorship
  • Long Chen

    Possible advisor · last author on 4 of their early first-author papers, 2023–2024

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works32 from public data

TitleCited by
  • Similarity Reasoning and Filtration for Image-Text Matching

    Haiwen Diao, Ying Zhang, Lin Ma, Huchuan Lu

    Proceedings of the AAAI Conference on Artificial Intelligence · 2021

    The superiority of the proposed Similarity Graph Reasoning and Attention Filtration network with achieving state-of-the-art performances on the Flickr30K and MSCOCO datasets is demonstrated, and the good interpretability of SGR and SAF with extensive qualitative experiments and analyses are demonstrated.

    445
  • Autoregressive Video Generation without Vector Quantization

    Haoge Deng, Pan, Ting, Haiwen Diao, Zhengxiong Luo, Yufeng Cui, Huchuan Lu, Shiguang Shan, Yonggang Qi, +1 more

    arXiv · 2024

    This paper proposes to reformulate the video generation problem as a non-quantized autoregressive modeling of temporal frame-by-frame prediction and spatial set-by-set prediction, and trains a novel video autoregressive model without vector quantization, termed NOVA.

    198
  • Plug-and-Play Regulators for Image-Text Matching

    Haiwen Diao, Ying Zhang, Wei Liu, Xiang Ruan, Huchuan Lu

    IEEE Transactions on Image Processing · 2023

    Two simple but quite effective regulators are developed which efficiently encode the message output to automatically contextualize and aggregate cross-modal representations and can bring an impressive and consistent R@1 gain on multiple models, confirming the general effectiveness and generalization ability of the proposed methods.

    47
  • DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

    Haozhe Xie, Beichen Wen, Jiarui Zheng, Zhaoxi Chen, Fangzhou Hong, Haiwen Diao, Ziwei Liu

    arXiv · 2026

    The Dynamic Object Manipulation (DOM) benchmark is introduced, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation.

    45
  • EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

    Haiwen Diao, Xiaotong Li, Yufeng Cui, Yueze Wang, Haoge Deng, Ting Pan, Wenxuan Wang, Huchuan Lu, +1 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    This work systematically clarify the performance gap between VLMs using pre-trained vision encoders, discrete tokenizers, and minimalist visual layers from scratch, and develops efficient strategies for encoder-free VLMs that rival mainstream encoder-based ones.

    41
  • Exploring Dynamic Transformer for Efficient Object Tracking

    Jiawen Zhu, Xin Yu Chen, Haiwen Diao, Shuai Li, Jun-Yan He, Chenyang Li, Bin Luo, Dong Wang, +1 more

    IEEE Transactions on Neural Networks and Learning Systems · 2025

    This article proposes DyTrack, a dynamic transformer framework for efficient tracking that automatically learns to configure proper reasoning routes for different inputs, thereby improving the utilization of the available computational budget and achieving higher performance at the same running speed.

    36
  • UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory

    Haiwen Diao, Bo Wan, Ying Zhang, Xu Jia, Huchuan Lu, Long Chen

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    It is argued that the scalability, adaptability, and generalizability of state-of-the-art methods are hindered by structural dependency and pertinency on specific pretrained backbones, and a new memoryefficient PETL strategy, Universal Parallel Tuning (UniPT), is proposed to mitigate these weaknesses.

    35
  • SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

    Haiwen Diao, Penghao Wu, Hanming Deng, Jiahao Wang, Shihao Bai, Silei Wu, Fan, Weichen, Wenjie Ye, +22 more

    arXiv · 2026

    This work introduces SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process that points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner.

    28
  • Visual Jigsaw Post-Training Improves MLLMs

    Penghao Wu, Yushan Zhang, Haiwen Diao, Bo Li, Lu, Lewei, Ziwei Liu

    arXiv · 2025

    This work introduces Visual Jigsaw, a generic self-supervised post-training framework designed to strengthen visual understanding in MLLMs that naturally aligns with reinforcement learning from verifiable rewards (RLVR), requires no additional visual generative components, and derives its supervisory signal automatically without any annotations.

    25
  • MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models

    Xiaomin Li, Xu Jia, Qinghe Wang, Haiwen Diao, Mengmeng Ge, Pengxiang Li, You He, Huchuan Lu

    ACM International Conference on Multimedia (ACM MM) · 2024

    This work introduces a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process.

    15
  • Deep Boosting Learning: A Brand-New Cooperative Approach for Image-Text Matching

    Haiwen Diao, Ying Zhang, Shang Gao, Xiang Ruan, Huchuan Lu

    IEEE Transactions on Image Processing · 2024

    This paper proposes a brand-new Deep Boosting Learning (DBL) algorithm, where an anchor branch is first trained to provide insights into the data properties, with a target branch gaining more advanced knowledge to develop optimal features and distance metrics.

    14
  • LLMs Can Evolve Continually on Modality for X-Modal Reasoning

    Jiazuo Yu, Xiong, Haomiao, Zhang, Lu, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang, Huchuan Lu, +2 more

    arXiv · 2024

    PathWeave is proposed, a flexible and scalable framework with modal-Path sWitching and ExpAnsion abilities that enables MLLMs to continually EVolve on modalities for $\mathbb{X}$-modal reasoning.

    13
  • GSSF: Generalized Structural Sparse Function for Deep Cross-Modal Metric Learning

    Haiwen Diao, Ying Zhang, Shang Gao, Jiawen Zhu, Long Chen, Huchuan Lu

    IEEE Transactions on Image Processing · 2024

    A Generalized Structural Sparse Function is proposed to dynamically capture thorough and powerful relationships across modalities for pair-wise similarity learning while remaining concise but efficient and reaches a sweet spot between model complexity and capability.

    10
  • VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

    Qing'an Liu, Juntong Feng, Yuhao Wang, Xinzhe Han, Yujie Cheng, Yue Zhu, Haiwen Diao, Yunzhi Zhuge, +1 more

    arXiv · 2026

    VISTA-Bench is introduced, a systematic benchmark from multimodal perception, reasoning, to unimodal understanding domains that evaluates visualized text understanding by contrasting pure-text and visualized-text questions under controlled rendering conditions.

    8
  • From Pixels to Words -- Towards Native One-Vision Models at Scale

    Haiwen Diao, Jiahao Wang, Penghao Wu, Yuhao Dong, Yuwei Niu, Yue Zhu, Zhongang Cai, Weichen Fan, +13 more

    arXiv · 2026

    NEO-ov is introduced, a native foundation model that learns cross-frame and pixel-word correspondence end-to-end without any external encoders, auxiliary adapters, or post-hoc fusion, validating that native"one-vision"architectures are not only feasible but competitive at scale.

    6
  • Vision as Unified Multimodal Generation

    Xiaoyang Han, Jianhua Li, Kewang Deng, Zukai Chen, Xuanke Shi, Sihan Wang, Boxuan Li, Linyan Wang, +9 more

    arXiv · 2026

    Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.

    5
  • Unveiling Encoder-Free Vision-Language Models

    Haiwen Diao, Yufeng Cui, Xiaotong Li, Yueze Wang, Huchuan Lu, Yueze Wang

    neural information processing systems · 2024

    5
  • SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning

    Haiwen Diao, Bo Wan, Xu Jia, Yunzhi Zhuge, Ying Zhang, Huchuan Lu, Long Chen

    Lecture notes in computer science · 2024

    4
  • ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs

    Yuhao Wang, Mu Qiao, Haiwen Diao, Yunzhi Zhuge, P Zhang, Xindong Zhang, Lei Zhang, Hüseyin Ozan Ҫirkinoğlu

    arXiv · 2026

    ERA establishes logit-preserving visual token pruning as a principled framework for efficient MLLMs, unifying theoretical foundation, algorithmic design, and practical deployment.

    2
  • KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification

    Yue Zhu, Haiwen Diao, Shang Gao, Long Chen, Huchuan Lu

    IEEE International Conference on Acoustics Speech and Signal Processing · 2025

    This work introduces an innovative Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission (KARST), which outperforms other PEFT counterparts across model types and data domains, but also surpasses full fine-tuning with a negligible inference cost due to its re-parameterization characteristics.

    2
  • DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

    Xiaotong Li, Fan Zhang, Haiwen Diao, Xinlong Wang, Xinlong Wang, Ling‐Yu Duan

    neural information processing systems · 2024

    2
  • Regularizing Subspace Redundancy of Low-Rank Adaptation

    Yue Zhu, Haiwen Diao, Shang Gao, Jiazuo Yu, Jiawen Zhu, Yunzhi Zhuge, Shuai Hao, Xu Jia, +3 more

    ACM International Conference on Multimedia (ACM MM) · 2025

    ReSoRA theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections and systematically applies de-redundancy constraints to the feature distributions across different projections.

    1
  • LLMs Can Evolve Continually on Modality for $\mathbb{X}$-Modal Reasoning

    Jiazuo Yu, Haomiao Xiong, Lu Zhang, Haiwen Diao, Yunzhi Zhuge, Lanqing Hong, Dong Wang, Jiazuo Yu, +2 more

    neural information processing systems · 2024

    1
  • VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    Zhongbo Zhang, Jiayi Jin, Yifan Wang, Zaibin Zhang, Haiwen Diao, Lijun Wang, Huchuan Lu

    arXiv · 2026

    –
  • SenseNova-U1.5: Towards Native Unified Visual Intelligence

    Haiwen Diao, Jiahao Wang, Chenjing Ding, Hanming Deng, Jiangnan Chen, Ruixi Zhang, Ruohui Wang, Wenwen Tong, +22 more

    arXiv · 2026

    –

Show all 32 works

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.