Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Kaipeng ZhangSuggested from co-authorship
- Shuqiang JiangSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works15 from public data
- OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation36
OpenING, a comprehensive benchmark comprising 5,400 high-quality human-annotated instances across 56 real-world tasks, is introduced and IntJudge, a judge model for evaluating open-ended multimodal generation methods, is presented.
- FoodSky: A food-oriented large language model that can pass the chef and dietetic examinations29
The food-oriented large language model (LLM) FoodSky is introduced, which offers fine-grained perception and reasoning on food data and aims to establish a new benchmark for domain-specific LLMs in addressing real-world food-related challenges.
- Synthesizing Knowledge-Enhanced Features for Real-World Zero-Shot Food Detection22
A novel framework ZSFDet is proposed to tackle fine-grained problems in Zero-Shot Food Detection by exploiting the interaction between complex attributes by model the correlation between food categories and attributes in ZSFDet by multi-source graphs to provide prior knowledge for distinguishing fine-grained features.
- SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food Detection19
The Semantic Separable Diffusion Synthesizer (SeeDS) framework for Zero-Shot Food Detection (ZSFD) is proposed, which learns the disentangled semantic representation for complex food attributes from ingredients and cuisines, and synthesizes discriminative food features via enhanced semantic information.
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models17
This work introduces MDK12-Bench, a multi-disciplinary benchmark assessing the reasoning capabilities of MLLMs via real-world K-12 examinations, and presents a novel dynamic evaluation framework to mitigate data contamination issues by bootstrapping question forms, question types, and image styles during evaluation.
- DD-Ranking: Rethinking the Evaluation of Dataset Distillation14
DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.
- 12
- 10
- ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges9
ProJudge-173k, a large-scale instruction-tuning dataset, and a Dynamic Dual-Phase fine-tuning strategy that encourages models to explicitly reason through problem-solving before assessing solutions are proposed, which significantly enhance the process evaluation capabilities of open-source models.
- MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams4
MDK12-Bench is introduced, a large-scale multidisciplinary benchmark built from real-world K-12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy and a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination.
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans2
This paper proposes Human-Aligned Bench, a benchmark for fine-grained alignment of multimodal reasoning with human performance, and reveals notable differences between the performance of current MLLMs in multimodal reasoning and human performance.
- 2
- 1
- 1
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.