Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks14 from public data
- Otter: A Multi-Modal Model With In-Context Instruction Tuning716
Otter builds upon Flamingo with Perceiver architecture, and has been instruction tuned for general purpose multi-modal assistant, and substantially enhances model convergence and generalization capabilities.
- LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models373
This work introduces LMMS-EVAL, a unified and standardized multimodal benchmark framework with over 50 tasks and more than 10 models to promote transparent and reproducible evaluations, and introduces LMMS-EVAL LITE, a pruned evaluation toolkit that emphasizes both coverage and efficiency.
- MIMIC-IT: Multi-Modal In-Context Instruction Tuning316
MultI-Modal In-Context Instruction Tuning (MIMIC-IT), a dataset comprising 2.8 million multimodal instruction-response pairs, with 2.2 million unique instructions derived from images and videos, is presented and a large VLM named Otter is trained.
- Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos241
Video-MMMU is introduced, a multi-modal, multi-disciplinary benchmark designed to assess LMMs' ability to acquire and utilize knowledge from videos, and proposes a proposed knowledge gain metric, {\Delta}knowledge, that quantifies improvement in performance after video viewing.
- OtterHD: A High-Resolution Multi-modality Model87
This study highlights the critical role of flexibility and high-resolution input capabilities in large multimodal models and also exemplifies the potential inherent in the Fuyu architecture's simplicity for handling complex visual data.
- WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning20
This paper presents WorldQA, a video understanding dataset designed to push the boundaries of multimodal world models with three appealing properties, and introduces WorldRetriever, an agent designed to synthesize expert knowledge into a coherent reasoning chain, thereby facilitating accurate responses to WorldQA queries.
- 17
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing2
The results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
- Senses Wide Shut: A Representation-Action Gap in Omnimodal LLMs2
IMAVB, a curated 500-clip benchmark of long-form movies with a 2x2 design crossing target modality and premise condition, is introduced, which lets us measure conflict detection separately from ordinary multimodal comprehension.
- Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients1
This work proposes a novel framework, VLoRP, that extends low-rank gradient projection by introducing an additional degree of freedom for controlling the trade-off between memory efficiency and performance, beyond the rank hyper-parameter, and presents ProjFactor, an adaptive memory-efficient optimizer.
- Unlocking More Granular Control of Memory-Efficient LLM Finetuning–
This work systematically investigates the impact of the projection unit on LoRP methods, and extends existing LoRP approaches by introducing an additional degree of freedom, projection granularity, beyond the traditional rank hyperparameter, which enables a framework capable of performing Various-grained Low-Rank Projection of gradients, which is named VLoRP.
- –
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.