Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works5 from public data

TitleCited by
  • Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

    Yuhao Dong, Zuyan Liu, Hailong Sun, Jingkang Yang, Winston Hu, Yongming Rao, Ziwei Liu

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    InsightV, an early effort to scalably produce long and robust reasoning data for complex multi-modal tasks, and an effective training pipeline to enhance the reasoning capabilities of multi-modal large language models (MLLMs), is presented.

    153
  • Towards Language-Driven Video Inpainting via Multimodal Large Language Models

    Jianzong Wu, Xiangtai Li, Chenyang Si, Shangchen Zhou, Jingkang Yang, Jiangning Zhang, Yining Li, Kai Chen, +3 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work introduces a new task - language-driven video inpainting, which uses natural language instructions to guide the inpainting process, integrating Multimodal Large Language Models to understand and execute complex language-based inpaintingrequests effectively.

    48
  • Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation

    Jinghao Wang, Zhengyu Wen, Xiangtai Li, Zujin Guo, Jingkang Yang, Ziwei Liu

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2024

    A novel framework is presented: Pair then Relation (Pair-Net), which uses a Pair Proposal Network (PPN) to learn and filter sparse pair-wise relationships between subjects and objects and achieves over 10% absolute gains compared to the baseline, PSGFormer.

    36
  • Ego-R1: Agentic Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

    Shulin Tian, Rui Wang, Hongming Guo, Penghao Wu, Yuhao Dong, Xiuying Wang, Jingkang Yang, H L Zhang, +2 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2026

    10
  • UNO: Unifying One-stage Video Scene Graph Generation via Object-Centric Visual Representation Learning

    Huy Le, Nhat Chung, Tung Kieu, Jingkang Yang, Ngan Le

    IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) · 2026

    UNO (UNified Object-centric VidSGG), a single-stage, unified framework that jointly addresses both tasks within an end-to-end architecture, and introduces object temporal consistency learning, which enforces consistent object representations across frames without relying on explicit tracking modules.

    4

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.