Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Yanfeng Wang

    Possible advisor · last author on 3 of their early first-author papers, 2023–2025

    Suggested from co-authorship
  • Weidi Xie

    Possible advisor · last author on 3 of their early first-author papers, 2024–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works33 from public data

TitleCited by
  • Kimi-VL Technical Report

    Kimi Team, Angang Du, Yin, Bohong, Bowei Xing, Qu, Bowen, Bowen Wang, Cheng Chen, Chenlin Zhang, +22 more

    arXiv · 2025

    352
  • Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models

    Chang Liu, Haoning Wu, Yujie Zhong, Xiaoyun Zhang, Yanfeng Wang, Weidi Xie

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work proposes a learning-based auto-regressive im-age generation model, termed as Story Gen, with a novel vision-language context module, that can generalize to unseen characters without any optimization, and generate image sequences with coherent content and consistent character.

    91
  • SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass

    Yanxu Meng, Haoning Wu, Ya Zhang, Weidi Xie

    International Conference on 3D Vision (3DV) · 2026

    This work presents SceneGen, a novel framework that takes a scene image and corresponding object masks as input, simultaneously producing multiple 3D assets with geometry and texture, and introduces a novel feature aggregation module that integrates local and global scene information from visual and geometric encoders within the feature extraction module.

    58
  • LAR-SR: A Local Autoregressive Model for Image Super-Resolution

    Baisong Guo, Xiaoyun Zhang, Haoning Wu, Yu Wang, Ya Zhang, Yanfeng Wang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    46
  • SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

    Haoning Wu, Xiao Huang, Yaohui Chen, Zhang, Ya, Yanfeng Wang, Xie, Weidi

    arXiv · 2025

    This paper proposes SpatialScore, the most comprehensive and diverse multimodal spatial understanding benchmark to date, integrating VGBench with relevant data from the other 11 existing datasets, and develops SpatialAgent, a novel multi-agent system incorporating 9 specialized tools for spatial understanding.

    38
  • BabyVision: Visual Reasoning Beyond Language

    Liang Chen, Weichu Xie, Yiyan Liang, Hongfeng He, Hans Zhao, Zhibo Yang, Zhiqi Huang, Haoning Wu, +22 more

    arXiv · 2026

    This work introduces BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for Multimodal LLMs, and explores solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit.

    35
  • Towards Universal Soccer Video Understanding

    Jiayuan Rao, Haoning Wu, Hao Jiang, Ya Zhang, Yanfeng Wang, Weidi Xie

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    An advanced soccer-specific visual encoder, MatchVision, is presented, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks, which highlights the superiority of the proposed data and model.

    33
  • MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning

    Haoning Wu, Shaocheng Shen, Qiang Hu, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang

    IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) · 2025

    This paper introduces MegaFusion, a novel approach that extends existing diffusion-based text-to-image models towards efficient higher-resolution generation without additional fine-tuning or adaptation, and employs an innovative truncate and relay strategy to bridge the denoising processes across different resolutions.

    33
  • Multi-Agent System for Comprehensive Soccer Understanding

    Jiayuan Rao, Zifeng Li, Haoning Wu, Ya Zhang, Yanfeng Wang, Weidi Xie

    ACM International Conference on Multimedia (ACM MM) · 2025

    This paper constructs SoccerWiki, the first large-scale multimodal soccer knowledge base, integrating rich domain knowledge about players, teams, referees, and venues to enable knowledge-driven reasoning and introduces SoccerAgent, a novel multi-agent system that decomposes complex soccer questions via collaborative reasoning.

    23
  • VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation

    Ziyang Luo, Haoning Wu, Dongxu Li, Jing Ma, Mohan Kankanhalli, Junnan Li

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    VideoAutoArena is introduced, an arena-style benchmark inspired by LMSYS Chatbot Arena’s framework, designed to automatically assess LMMs’ video analysis abilities, and introduces a fault-driven evolution strategy, progressively increasing question complexity to push models toward handling more challenging video analysis scenarios.

    19
  • MatchTime: Towards Automatic Soccer Game Commentary Generation

    J. Shyam Sundar Rao, Haoning Wu, Chang Liu, Yanfeng Wang, Weidi Xie

    Conference on Empirical Methods in Natural Language Processing (EMNLP) · 2024

    19
  • AIM 2020 Challenge on Real Image Super-Resolution: Methods and Results

    Pengxu Wei, Hannan Lu, Radu Timofte, Liang Lin, Wangmeng Zuo, Zhihong Pan, Baopu Li, Teng Xi, +22 more

    arXiv · 2020

    19
  • Dual-Branch Network for Portrait Image Quality Assessment

    Wei Sun, Weixia Zhang, Yanwei Jiang, Haoning Wu, Zicheng Zhang, Jun Jia, Yingjie Zhou, Zhongpeng Ji, +3 more

    arXiv · 2024

    A dual-branch network for portrait image quality assessment (PIQA) is introduced, which can effectively address how the salient person and the background of a portrait image influence its visual quality.

    16
  • VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

    Hanwei Zhu, Haoning Wu, Zicheng Zhang, Lingyu Zhu, Yixuan Li, Peilin Chen, Shiqi Wang, Chris Wei Zhou, +21 more

    IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) · 2025

    A novel benchmark comprising thousands of coarse-to-fine grained visual quality comparison tasks, spanning single images, pairs, and multi-image groups is introduced, which serves as a catalyst for future research on inter-pretable and human-aligned quality evaluation systems.

    10
  • MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models

    Siwei Wu, King Zhu, Yu Bai, Yiming Liang, Yizhi LI, Haoning Wu, Jiaheng Liu, Ruibo Liu, +5 more

    Findings of the Association for Computational Linguistics: EACL · 2026

    These findings indicate that while LVLMs demonstrate a strong capability to perceive image details, enhancing their ability to associate information across multiple images hinges on improving the reasoning capabilities of their language model component.

    9
  • ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts

    Mingxin Wang, Bin Hu, Bin Qian, Kaitao Jiang, Haoning Wu, Feng Yan, Bowen Jing, Ruiyang Hao, +7 more

    arXiv · 2026

    Semantic-Temporal WAM (ST-WAM) is proposed to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics, demonstrating that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.

    9
  • WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models

    Runjie Zhou, Youbo Shao, Haoyu Lu, Bowei Xing, Tongtong Bai, Yujie Chen, Jie Zhao, Lin Sui, +11 more

    arXiv · 2026

    WorldVQA is introduced, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models, thereby establishing a standard for assessing the encyclopedic breadth and hallucination rates of current and next-generation frontier models.

    5
  • Towards Pixel-Level VLM Perception via Simple Points Prediction

    Tianhui Song, Haoyu Lu, Hao Yang, Lin Sui, Haoning Wu, Zaida Zhou, Zhiqi Huang, Yiping Bao, +3 more

    arXiv · 2026

    This work lays out that precise spatial understanding can emerge from simple point prediction, challenging the prevailing need for auxiliary components and paving the way for more unified and capable VLMs.

    5
  • PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

    Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, +22 more

    arXiv · 2026

    PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.

    4
  • Boost Video Frame Interpolation via Motion Adaptation

    Haoning Wu, Xiaoyun Zhang, Weidi Xie, Ya Zhang, Yanfeng Wang

    arXiv · 2023

    This paper proposes a novel optimization-based VFI method that can adapt to unseen motions at test time, based on a cycle-consistency adaptation strategy that leverages the motion characteristics among video frames.

    4
  • MRGen: Segmentation Data Engine for Underrepresented MRI Modalities

    Haoning Wu, Ziheng Zhao, Ya Zhang, Yue Wang, Weidi Xie

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    This paper investigates leveraging generative models to synthesize data, for training segmentation models for underrepresented modalities, particularly on annotation-scarce MRI, and believes that MRGen significantly improves segmentation performance on unannotated modalities by providing high-quality synthetic data.

    3
  • Generative Frame Sampler for Long Video Understanding

    Linli Yao, Haoning Wu, Kun Ouyang, Yuanxing Zhang, Caiming Xiong, Bei Chen, Xu Sun, Junnan Li

    Findings of the Association for Computational Linguistics: ACL 2025 · 2025

    3
  • Kimi K2.5: Visual Agentic Intelligence

    Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S. H. Cai, Yuan Cao, Ziwei Chai, Y. Charles, +22 more

    arXiv · 2026

    2
  • Improving Human Image Animation via Semantic Representation Alignment

    Chang Liu, Mengting Chen, Yixuan Huang, Haoning Wu, Chen Ju, Shuai Xiao, Jinsong Lan, Yanfeng Wang

    arXiv · 2026

    A novel approach named SemanticREPA is introduced that leverages these semantic representations as supervision signals through representation alignment to generate coherent and stable human structures and uses the predicted structure representations to refine identity restoration in relevant regions.

    1
  • Count Anything at Any Granularity

    Chang Liu, Haoning Wu, Weidi Xie

    arXiv · 2026

    This work redefine open-world counting as multi-grained counting, where visual exemplars specify target appearance and fine-grained text, with optional negative prompts, specifies the intended semantic granularity across five explicit levels.

    1
  • Kimi K3: Open Frontier Intelligence

    Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, M. C., Jianfeng Cai, Xinyuan Cai, Peizhou Cao, +22 more

    arXiv · 2026

    1
  • NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth Perception

    Qiuwen Wang, Shuai Guo, Haoning Wu, Rong Xie, Li Song, Wenjun Zhang

    ACM Multimedia Asia (MMAsia) · 2023

    This work proposes a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception, and achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency.

    1
  • PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning

    Shaoxuan Li, Zhixuan Zhao, Hanze Deng, Zirun Ma, Shulin Tian, Zuyan Liu, Yushi Hu, Haoning Wu, +4 more

    Lecture notes in computer science · 2026

    –
  • Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance

    Xunchu Zhou, Xiaohong Liu, Yunlong Dong, Tengchuan Kou, Yixuan Gao, Zicheng Zhang, Chunyi Li, Haoning Wu, +1 more

    International Journal of Computer Vision · 2026

    –
  • BasketEvent: Understanding Who Did What and When in Basketball Videos

    Yu Zhang, Jiayuan Rao, Haoning Wu, Wei Xie

    arXiv · 2026

    PlayNet is proposed, a player-centric reasoning framework that maps basketball videos to player-level event predictions with temporal evidence, proving the superiority of player-centric modeling for fine-grained sports video understanding.

    –
  • OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams

    Yibin Yan, Jilan Xu, Shangzhe Di, Haoning Wu, Weidi Xie

    Lecture notes in computer science · 2026

    –
  • SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

    Kun Ouyang, Yuanxin Liu, Haoning Wu, Yi Liu, Hao Zhou, Jie Zhou, Fandong Meng, Xu Guang Sun

    arXiv · 2025

    –
  • R-Bench: Are your Large Multimodal Model Robust to Real-world Corruptions?

    Li, Chunyi, Jianbo Zhang, Zicheng Zhang, Haoning Wu, Yuan Tian, Wei Sun, Lu Guo, Xiaohong Liu, +3 more

    arXiv · 2024

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it.

Sign in to report