Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works14 from public data

TitleCited by
  • Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined Levels

    Haoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen, Liang Liao, Li, Chunyi, Yixuan Gao, Annan Wang, +6 more

    arXiv · 2023

    The proposed Q-Align achieves state-of-the-art performance on image quality assessment (IQA), image aesthetic assessment (IAA), as well as video quality assessment (VQA) tasks under the original LMM structure and unify the three tasks into one model, termed the OneAlign.

    664
  • Exploring Video Quality Assessment on User Generated Contents from Aesthetic and Technical Perspectives

    Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, +1 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    The Disentangled Objective Video Quality Evaluator (DOVER) is proposed, the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective, and the first approach to provide reliable clear-cut quality evaluations from a single aesthetic or technical perspective.

    463
  • Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision

    Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Chunyi Li, Wenxiu Sun, +3 more

    arXiv · 2023

    Q-Bench, a holistic benchmark crafted to systematically evaluate potential abilities of MLLMs on three realms: low-level visual perception, low-level visual description, and overall visual quality assessment, confirms that MLLMs possess preliminary low-level visual skills.

    309
  • Q-Instruct: Improving Low-Level Visual Abilities for Multi-Modality Foundation Models

    Haoning Wu, Zicheng Zhang, Erli Zhang, Chaofeng Chen, Liang Liao, Annan Wang, Kaixin Xu, Chunyi Li, +6 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    The first dataset consisting of human natural language feedback on low-level vision, and a GPT-participated transformation to convert these feedbacks into a rich set of 200K instruction-response pairs, termed Q-Instruct are collected.

    218
  • Towards Explainable In-the-Wild Video Quality Assessment: A Database and a Language-Prompted Approach

    Haoning Wu, Erli Zhang, Liang Liao, Chaofeng Chen, Jingwen Hou, Annan Wang, Wenxiu Sun, Qiong Yan, +1 more

    ACM International Conference on Multimedia (ACM MM) · 2023

    The MaxVQA is proposed, a language-prompted VQA approach that modifies vision-language foundation model CLIP to better capture important quality issues as observed in the authors' analyses.

    80
  • Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning

    Haofeng Liu, Erli Zhang, Junde Wu, Mingxuan Hong, Yueming Jin

    arXiv · 2024

    This work introduces Surgical SAM 2 (SurgSAM2), an advanced model to utilize SAM2 with an Efficient Frame Pruning (EFP) mechanism, to facilitate real-time surgical video segmentation in resource-constrained environments, and significantly improves both efficiency and segmentation accuracy.

    67
  • SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

    Zhitao Zeng, Zhuo, Zhu, Xiaojun Jia, Erli Zhang, Junde Wu, Zhang, Jiaan, Yuxuan Wang, Chang Han Low, +7 more

    arXiv · 2025

    This work proposes SurgVLM, one of the first large vision-language foundation models for surgical intelligence, where this single universal model can tackle versatile surgical tasks and builds upon Qwen2.5-VL, which is built upon Qwen2.5-VL and undergoes instruction tuning to 10+ surgical tasks.

    44
  • Towards Open-Ended Visual Quality Comparison

    Haoning Wu, Hanwei Zhu, Zicheng Zhang, Erli Zhang, Zhaofeng Chen, Liang Liao, Chunyi Li, Annan Wang, +6 more

    Lecture notes in computer science · 2024

    31
  • Q-Boost: On Visual Quality Assessment Ability of Low-Level Multi-Modality Foundation Models

    Zicheng Zhang, Haoning Wu, Zhongpeng Ji, Chunyi Li, Erli Zhang, Wei Sun, Xiaohong Liu, Xiongkuo Min, +4 more

    IEEE International Conference on Multimedia and Expo workshops · 2024

    Q-Boost is introduced, a novel strategy designed to enhance low-level MLLMs in image quality assessment (IQA) and video quality assessment (VQA) tasks, which is structured around two pivotal components: Triadic-Tone Integration and Multi-Prompt Ensemble.

    30
  • Q-Bench+: A Benchmark for Multi-Modal Foundation Models on Low-Level Vision From Single Images to Pairs

    Zicheng Zhang, Haoning Wu, Erli Zhang, Guangtao Zhai, Weisi Lin

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    It is demonstrated that several MLLMs have decent low-level visual competencies on single images, but only GPT-4V exhibits higher accuracy on pairwise comparisons than single image evaluations than single image evaluations like humans.

    28
  • Exploring Opinion-Unaware Video Quality Assessment with Semantic Affinity Criterion

    Haoning Wu, Liang Liao, Jingwen Hou, Chaofeng Chen, Erli Zhang, Annan Wang, Wenxiu Sun, Qiong Yan, +1 more

    IEEE International Conference on Multimedia and Expo (ICME) · 2023

    26
  • A JND Guided Foveation Video Coding

    Erli Zhang, Debin Zhao, Yongbing Zhang, Hongbin Liu, Siwei Ma, Ronggang Wang

    Lecture notes in computer science · 2008

    A novel just noticeable distortion (JND) guided foveation video coding method, by which the foveated region can be adaptively selected according to the video content, is presented.

    2
  • Abstract PO4-07-03: Machine learning breast cancer risk prediction using sequential past mammograms - a pilot study

    Lester Chee Hao Leong, Mingjie Xu, Weimin Huang, Erli Zhang, Engracia Loh, Sze Yiun Teo, Geok Hoon Lim, Veronique Kiak Mien Tan, +3 more

    Cancer Research · 2024

    The study showed that deep learning breast cancer risk prediction can be further improved by using sequential past mammograms instead of random mammograms for AI model training and can also potentially enhance other AI risk prediction models that employ combined mammogram and traditional breast cancer clinical risk factors.

    1
  • Computer Vision - ECCV 2024 : 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part III

    Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, Gül Varol, Haoning Wu, Hanwei Zhu, +12 more

    DR-NTU (Nanyang Technological University) · 2025

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.