Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks5 from public data
- Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models153
InsightV, an early effort to scalably produce long and robust reasoning data for complex multi-modal tasks, and an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs), is presented.
- Towards Language-Driven Video Inpainting via Multimodal Large Language Models48
This work introduces a new task - language-driven video inpainting, which uses natural language instructions to guide the inpainting process, integrating Multimodal Large Language Models to understand and execute complex language-based inpaintingrequests effectively.
- Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation36
A novel framework is presented: Pair then Relation (Pair-Net), which uses a Pair Proposal Network (PPN) to learn and filter sparse pair-wise relationships between subjects and objects and achieves over 10% absolute gains compared to the baseline, PSGFormer.
- 10
- UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning4
UNO (UNified Object-centric VidSGG), a single-stage, unified framework that jointly addresses both tasks within an end-to-end architecture, and introduces object temporal consistency learning, which enforces consistent object representations across frames without relying on explicit tracking modules.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.