Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works14 from public data

TitleCited by
  • Otter: A Multi-Modal Model With In-Context Instruction Tuning

    Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Joshua Adrian Cahyono, Jingkang Yang, Chunyuan Li, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025

    Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant, and substantially enhances model convergence and generalization capabilities.

    716
  • LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

    Kaichen Zhang, Bo Li, Peiyuan Zhang, Fanyi Pu, Joshua Adrian Cahyono, Kairui Hu, Shuai Liu, Yuanhan Zhang, +3 more

    Findings of the Association for Computational Linguistics: NAACL · 2025

    This work introduces LMMS-EVAL, a unified and standardized multimodal benchmark framework with over 50 tasks and more than 10 models to promote transparent and reproducible evaluations, and introduces LMMS-EVAL LITE, a pruned evaluation toolkit that emphasizes both coverage and efficiency.

    373
  • MIMIC-IT: Multi-Modal In-Context Instruction Tuning

    Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, Ziwei Liu

    arXiv · 2023

    MultI-Modal In-Context Instruction Tuning (MIMIC-IT), a dataset comprising 2.8 million multimodal instruction-response pairs, with 2.2 million unique instructions derived from images and videos, is presented and a large VLM named Otter is trained.

    316
  • Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

    Kairui Hu, Penghao Wu, Fanyi Pu, Xiao Wang, Yuanhan Zhang, Yue, Xiang, Bo Li, Ziwei Liu

    arXiv · 2025

    Video-MMMU is introduced, a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos, and proposes a proposed knowledge gain metric, {\Delta}knowledge, that quantifies improvement in performance after video viewing.

    241
  • OtterHD: A High-Resolution Multi-modality Model

    Li, Bo, Peiyuan Zhang, Jingkang Yang, Yuanhan Zhang, Fanyi Pu, Ziwei Liu

    arXiv · 2023

    This study highlights the critical role of flexibility and high-resolution input capabilities in large multimodal models and also exemplifies the potential inherent in the Fuyu architecture's simplicity for handling complex visual data.

    87
  • WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning

    Yuanhan Zhang, Kaichen Zhang, Bo Li, Fanyi Pu, Christopher Arif Setiadharma, Jingkang Yang, Ziwei Liu

    arXiv · 2024

    This paper presents WorldQA, a video understanding dataset designed to push the boundaries of multimodal world models with three appealing properties, and introduces WorldRetriever, an agent designed to synthesize expert knowledge into a coherent reasoning chain, thereby facilitating accurate responses to WorldQA queries.

    20
  • Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos

    Kairui Hu, Penghao Wu, Fanyi Pu, Wang Xiao, Xiang Yue, Bo Li, Yuanhan Zhang, Ziwei Liu

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2026

    17
  • Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    Xinjie Zhang, Peng Zhang, Shicheng Zheng, Jinghao Guo, Zhaoyang Jia, Yifei Shen, Xiang Guo, Yuxuan Luo, +16 more

    arXiv · 2026

    The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.

    2
  • Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs

    Trung Nguyen Quang, Yiming Gao, Fanyi Pu, Kaichen Zhang, Shuo Sun, Ziwei Liu

    arXiv · 2026

    IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality and premise condition, is introduced, which lets us measure conflict detection separately from ordinary multimodal comprehension.

    2
  • Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients

    Yezhen Wang, Zhouhao Yang, Brian Chen, Fanyi Pu, Bo Li, Tianyu Gao, Kenji Kawaguchi

    arXiv · 2025

    This work proposes a novel framework, VLoRP, that extends low-rank gradient projection by introducing an additional degree of freedom for controlling the trade-off between memory efficiency and performance, beyond the rank hyper-parameter, and presents ProjFactor, an adaptive memory-efficient optimizer.

    1
  • Unlocking More Granular Control of Memory-Efficient LLM Finetuning

    Yezhen Wang, Zhouhao Yang, Fanyi Pu, Kenji Kawaguchi

    International Joint Conference on Artificial Intelligence · 2026

    This work systematically investigates the impact of the projection unit on LoRP methods, and extends existing LoRP approaches by introducing an additional degree of freedom, projection granularity, beyond the traditional rank hyperparameter, which enables a framework capable of performing Various-grained Low-Rank Projection of gradients, which is named VLoRP.

    –
  • Scaling spatial intelligence with multimodal foundation models

    Fanyi Pu

    DR-NTU (Nanyang Technological University) · 2026

    –
  • VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

    Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, +22 more

    arXiv · 2026

    –
  • Demystifying Video Reasoning

    Ruisi Wang, Zhongang Cai, Fanyi Pu, Junxiang Xu, Wanqi Yin, Maijunxian Wang, Ran Ji, Chenyang Gu, +6 more

    Lecture notes in computer science · 2026

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.