Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Kaipeng Zhang

    Possible advisor · last author on 4 of their early first-author papers, 2024–2025

    Suggested from co-authorship
  • Shuqiang Jiang

    Possible advisor · last author on 3 of their early first-author papers, 2023–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works15 from public data

TitleCited by
  • OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation

    Pengfei Zhou, Xiaopeng Peng, Jiajun Song, Chuanhao Li, Zhaopan Xu, Yue Yang, Ziyao Guo, Hao Zhang, +10 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks, is introduced and IntJudge, a judge model for evaluating open-ended multimodal generation methods, is presented.

    36
  • FoodSky: A food-oriented large language model that can pass the chef and dietetic examinations

    Pengfei Zhou, Weiqing Min, Chaoran Fu, Ying Jin, Mingyu Huang, Xiangyang Li, Shuhuan Mei, Shuqiang Jiang

    Patterns · 2025

    The food-oriented large language model (LLM) FoodSky is introduced, which offers fine-grained perception and reasoning on food data and aims to establish a new benchmark for domain-specific LLMs in addressing real-world food-related challenges.

    29
  • Synthesizing Knowledge-Enhanced Features for Real-World Zero-Shot Food Detection

    Pengfei Zhou, Weiqing Min, Jiajun Song, Yang Zhang, Shuqiang Jiang

    IEEE Transactions on Image Processing · 2024

    A novel framework ZSFDet is proposed to tackle fine-grained problems in Zero-Shot Food Detection by exploiting the interaction between complex attributes by model the correlation between food categories and attributes in ZSFDet by multi-source graphs to provide prior knowledge for distinguishing fine-grained features.

    22
  • SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food Detection

    Pengfei Zhou, Weiqing Min, Yang Zhang, Jiajun Song, Ying Jin, Shuqiang Jiang

    ACM International Conference on Multimedia (ACM MM) · 2023

    The Semantic Separable Diffusion Synthesizer (SeeDS) framework for Zero-Shot Food Detection (ZSFD) is proposed, which learns the disentangled semantic representation for complex food attributes from ingredients and cuisines, and synthesizes discriminative food features via enhanced semantic information.

    19
  • MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models

    Pengfei Zhou, Fanrui Zhang, Xiaopeng Peng, Zhaopan Xu, Ai, Jiaxin, Qiu, Yansheng, Chuanhao Li, Li, Zhen, +12 more

    arXiv · 2025

    This work introduces MDK12-Bench, a multi-disciplinary benchmark assessing the reasoning capabilities of MLLMs via real-world K-12 examinations, and presents a novel dynamic evaluation framework to mitigate data contamination issues by bootstrapping question forms, question types, and image styles during evaluation.

    17
  • DD-Ranking: Rethinking the Evaluation of Dataset Distillation

    Zekai Li, Xinhao Zhong, Samir Khaki, Zhiyuan Liang, Yuhao Zhou, Mingjia Shi, Ziqiao Wang, Xuanlei Zhao, +22 more

    arXiv · 2025

    DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.

    14
  • Multimodal Food Learning

    Weiqing Min, Xingjian Hong, Yuxin Liu, Mingyu Huang, Ying Jin, Pengfei Zhou, Leyi Xu, Yilin Wang, +2 more

    ACM Transactions on Multimedia Computing Communications and Applications · 2025

    12
  • Self-Supervised Enhancement for Named Entity Disambiguation via Multimodal Graph Convolution

    Pengfei Zhou, Kaining Ying, Zhenhua Wang, Dongyan Guo, Cong Bai

    IEEE Transactions on Neural Networks and Learning Systems · 2022

    10
  • ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges

    Jiaxin Ai, Pengfei Zhou, Zhaopan Xu, Ming Li, Fanrui Zhang, Zizhen Li, Jianwen Sun, Yulong Feng, +3 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    ProJudge-173k, a large-scale instruction-tuning dataset, and a Dynamic Dual-Phase fine-tuning strategy that encourages models to explicitly reason through problem-solving before assessing solutions are proposed, which significantly enhance the process evaluation capabilities of open-source models.

    9
  • MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams

    Pengfei Zhou, Peng, Xiaopeng, Fanrui Zhang, Zhaopan Xu, Ai, Jiaxin, Qiu, Yansheng, Chuanhao Li, Li, Zhen, +13 more

    arXiv · 2025

    MDK12-Bench is introduced, a large-scale multidisciplinary benchmark built from real-world K-12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy and a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination.

    4
  • Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans

    Qiu, Yansheng, Xiao Li, Zhaopan Xu, Pengfei Zhou, Zheng Wang, Kaipeng Zhang

    arXiv · 2025

    This paper proposes Human-Aligned Bench, a benchmark for fine-grained alignment of multimodal reasoning with human performance, and reveals notable differences between the performance of current MLLMs in multimodal reasoning and human performance.

    2
  • Exploring Implicit and Explicit Relations with the Dual Relation-Aware Network for Image Captioning

    Zhiwei Zha, Pengfei Zhou, Cong Bai

    Lecture notes in computer science · 2022

    2
  • Neural-Driven Image Editing

    Pengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao, Zilong Ye, Zekai Li, Suorong Yang, Jiadong Pan, +9 more

    neural information processing systems · 2025

    1
  • MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification

    Xin Pan, Pengfei Zhou, Jiaxin Ai, Wangbo Zhao, Kai Wang, Xiaojiang Peng, Wenqi Shao, Hongxun Yao, +1 more

    Findings of the Association for Computational Linguistics: ACL 2025 · 2025

    1
  • Why Game Development Matters for Scaling World Models

    Pengfei Zhou, 王建书 WANG HeXin, Yixing Ma, Zhenglin Wan, Kaipeng Zhang, Wangbo Zhao, Yang You

    Zenodo (CERN European Organization for Nuclear Research) · 2026

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.