Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Chen Change LoySuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works11 from public data
- Aligning Bag of Regions for Open-Vocabulary Object Detection200
This work proposes to align the embedding of bag of regions beyond individual regions as a bag, and surpasses the previous best results by 4.6 box AP50 and 2.8 mask AP on novel categories of open-vocabulary COCO and LVIS benchmarks, respectively.
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction159
An in-depth analysis of the region-language alignment in CLIP models is embarked on, which proposes an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs.
- OMG-Seg: Is One Model Good Enough for all Segmentation?140
It is shown that OMG-Seg, a transformer-based encoder-decoder architecture with task-specific queries and outputs, can support over ten distinct segmentation tasks and yet significantly reduce computational and parameter overhead across various tasks and datasets.
- Harmonizing Visual Representations for Unified Multimodal Understanding and Generation82
Harmon is presented, a unified autoregres-sive framework that harmonizes understanding and generation tasks with a shared MAR encoder and achieves state-of-the-art image generation results on the GenEval, MJHQ30K and WISE benchmarks while matching the performance of methods with dedicated semantic encoders on image understanding benchmarks.
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation50
This report presents OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation that bridging the off-the-shelf multimodal large language models (LLMs) and diffusion models through a set of learnable queries and a light-weight transformer-based connector.
- F-LMM: Grounding Frozen Large Multimodal Models38
F-LMM—grounding frozen off-the-shelf LMMs in human-AI conversations—a straightforward yet effective design based on the fact that word-pixel correspondences conducive to visual grounding inherently exist in the attention mechanism of well-trained LMMs.
- 29
- 20
- 13
- 12
- Generative Photographic Control for Scene-Consistent Video Cinematic Editing3
This paper proposes CineCtrl, the first video cinematic editing framework that provides fine control over professional camera parameters (e.g., bokeh, shutter speed), and introduces a decoupled cross-attention mechanism to disentangle camera motion from photographic inputs, allowing fine-grained, independent control without compromising scene consistency.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.