Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Chen Change Loy

    Possible advisor · last author on 9 of their early first-author papers, 2023–2025

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works11 from public data

TitleCited by
  • Aligning Bag of Regions for Open-Vocabulary Object Detection

    Size Wu, Wenwei Zhang, Sheng Jin, Wentao Liu, Chen Change Loy

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    This work proposes to align the embedding of bag of regions beyond individual regions as a bag, and surpasses the previous best results by 4.6 box AP50 and 2.8 mask AP on novel categories of open-vocabulary COCO and LVIS benchmarks, respectively.

    200
  • CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

    Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Xiangtai Li, Wentao Liu, Chen Change Loy

    arXiv · 2023

    An in-depth analysis of the region-language alignment in CLIP models is embarked on, which proposes an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs.

    159
  • OMG-Seg: Is One Model Good Enough for all Segmentation?

    Xiangtai Li, Haobo Yuan, Wei Li, Henghui Ding, Size Wu, Wenwei Zhang, Yining Li, Kai Chen, +1 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    It is shown that OMG-Seg, a transformer-based encoder-decoder architecture with task-specific queries and outputs, can support over ten distinct segmentation tasks and yet significantly reduce computational and parameter overhead across various tasks and datasets.

    140
  • Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

    Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Zhonghua Wu, Qingyi Tao, Wentao Liu, Wei Li, +1 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    Harmon is presented, a unified autoregres-sive framework that harmonizes understanding and generation tasks with a shared MAR encoder and achieves state-of-the-art image generation results on the GenEval, MJHQ30K and WISE benchmarks while matching the performance of methods with dedicated semantic encoders on image understanding benchmarks.

    82
  • OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

    Size Wu, Zhonghua Wu, Zhaopeng Gong, Tao, Qingyi, Jin, Sheng, Li, Qinyue, Wei Li, Chen Change Loy

    arXiv · 2025

    This report presents OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation that bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector.

    50
  • F-LMM: Grounding Frozen Large Multimodal Models

    Size Wu, Sheng Jin, Wenwei Zhang, Lumin Xu, Wentao Liu, Wei Li, Chen Change Loy

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    F-LMM—grounding frozen off-the-shelf LMMs in human-AI conversations—a straightforward yet effective design based on the fact that word-pixel correspondences conducive to visual grounding inherently exist in the attention mechanism of well-trained LMMs.

    38
  • DST-Det: Open-Vocabulary Object Detection via Dynamic Self-Training

    Shilin Xu, Xiangtai Li, Size Wu, Wenwei Zhang, Yunhai Tong, Chen Change Loy

    IEEE Transactions on Circuits and Systems for Video Technology · 2024

    29
  • CLIM: Contrastive Language-Image Mosaic for Region Representation

    Size Wu, Wenwei Zhang, Lumin Xu, Sheng Jin, Wentao Liu, Chen Change Loy

    Proceedings of the AAAI Conference on Artificial Intelligence · 2024

    20
  • DST-Det: Simple Dynamic Self-Training for Open-Vocabulary Object Detection

    Shilin Xu, Xiangtai Li, Size Wu, Wenwei Zhang, Yunhai Tong, Chen Change Loy

    arXiv · 2023

    13
  • UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

    Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, Wei Song, Yibin Wang, Zhixiong Zhang, Tianhang Wang, +3 more

    arXiv · 2026

    12
  • Generative Photographic Control for Scene-Consistent Video Cinematic Editing

    Huiqiang Sun, Liao Shen, Peng, Zhan, Kun Wang, Size Wu, Yuhang Zang, Tianqi Liu, Zihao Huang, +4 more

    arXiv · 2025

    This paper proposes CineCtrl, the first video cinematic editing framework that provides fine control over professional camera parameters (e.g., bokeh, shutter speed), and introduces a decoupled cross-attention mechanism to disentangle camera motion from photographic inputs, allowing fine-grained, independent control without compromising scene consistency.

    3

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.