Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Huchuan LuSuggested from co-authorship
- Long ChenSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works32 from public data
- Similarity Reasoning and Filtration for Image-Text Matching445
The superiority of the proposed Similarity Graph Reasoning and Attention Filtration network with achieving state-of-the-art performances on the Flickr30K and MSCOCO datasets is demonstrated, and the good interpretability of SGR and SAF with extensive qualitative experiments and analyses are demonstrated.
- Autoregressive Video Generation without Vector Quantization198
This paper proposes to reformulate the video generation problem as a non-quantized autoregressive modeling of temporal frame-by-frame prediction and spatial set-by-set prediction, and trains a novel video autoregressive model without vector quantization, termed NOVA.
- Plug-and-Play Regulators for Image-Text Matching47
Two simple but quite effective regulators are developed which efficiently encode the message output to automatically contextualize and aggregate cross-modal representations and can bring an impressive and consistent R@1 gain on multiple models, confirming the general effectiveness and generalization ability of the proposed methods.
- DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation45
The Dynamic Object Manipulation (DOM) benchmark is introduced, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation.
- EVEv2: Improved Baselines for Encoder-Free Vision-Language Models41
This work systematically clarify the performance gap between VLMs using pre-trained vision encoders, discrete tokenizers, and minimalist visual layers from scratch, and develops efficient strategies for encoder-free VLMs that rival mainstream encoder-based ones.
- Exploring Dynamic Transformer for Efficient Object Tracking36
This article proposes DyTrack, a dynamic transformer framework for efficient tracking that automatically learns to configure proper reasoning routes for different inputs, thereby improving the utilization of the available computational budget and achieving higher performance at the same running speed.
- UniPT: Universal Parallel Tuning for Transfer Learning with Efficient Parameter and Memory35
It is argued that the scalability, adaptability, and generalizability of state-of-the-art methods are hindered by structural dependency and pertinency on specific pretrained backbones, and a new memoryefficient PETL strategy, Universal Parallel Tuning (UniPT), is proposed to mitigate these weaknesses.
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture28
This work introduces SenseNova-U1, a native unified multimodal paradigm built upon NEO-unify, in which understanding and generation evolve as synergistic views of a single underlying process that points toward a broader roadmap where models do not translate between modalities, but think and act across them in a native manner.
- Visual Jigsaw Post-Training Improves MLLMs25
This work introduces Visual Jigsaw, a generic self-supervised post-training framework designed to strengthen visual understanding in MLLMs that naturally aligns with reinforcement learning from verifiable rewards (RLVR), requires no additional visual generative components, and derives its supervisory signal automatically without any annotations.
- MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models15
This work introduces a multimodal large language model (MLLM)-based recaptioner to expand the initial prompt to focus more on appearance and an appearance injection module to adapt appearance prior from video frames to the motion modeling process.
- Deep Boosting Learning: A Brand-New Cooperative Approach for Image-Text Matching14
This paper proposes a brand-new Deep Boosting Learning (DBL) algorithm, where an anchor branch is first trained to provide insights into the data properties, with a target branch gaining more advanced knowledge to develop optimal features and distance metrics.
- LLMs Can Evolve Continually on Modality for X-Modal Reasoning13
PathWeave is proposed, a flexible and scalable framework with modal-Path sWitching and ExpAnsion abilities that enables MLLMs to continually EVolve on modalities for $\mathbb{X}$-modal reasoning.
- GSSF: Generalized Structural Sparse Function for Deep Cross-Modal Metric Learning10
A Generalized Structural Sparse Function is proposed to dynamically capture thorough and powerful relationships across modalities for pair-wise similarity learning while remaining concise but efficient and reaches a sweet spot between model complexity and capability.
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?8
VISTA-Bench is introduced, a systematic benchmark from multimodal perception, reasoning, to unimodal understanding domains that evaluates visualized text understanding by contrasting pure-text and visualized-text questions under controlled rendering conditions.
- From Pixels to Words -- Towards Native One-Vision Models at Scale6
NEO-ov is introduced, a native foundation model that learns cross-frame and pixel-word correspondence end-to-end without any external encoders, auxiliary adapters, or post-hoc fusion, validating that native"one-vision"architectures are not only feasible but competitive at scale.
- Vision as Unified Multimodal Generation5
Experiments show that a single unified model can match leading task-specialized systems across structured visual understanding, dense geometric prediction, segmentation, and multi-view visual geometry.
- 5
- 4
- ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs2
ERA establishes logit-preserving visual token pruning as a principled framework for efficient MLLMs, unifying theoretical foundation, algorithmic design, and practical deployment.
- KARST: Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission for Visual Classification2
This work introduces an innovative Multi-Kernel Kronecker Adaptation with Re-Scaling Transmission (KARST), which outperforms other PEFT counterparts across model types and data domains, but also surpasses full fine-tuning with a negligible inference cost due to its re-parameterization characteristics.
- 2
- Regularizing Subspace Redundancy of Low-Rank Adaptation1
ReSoRA theoretically decomposes the low-rank submatrices into multiple equivalent subspaces and systematically applies de-redundancy constraints to the feature distributions across different projections and systematically applies de-redundancy constraints to the feature distributions across different projections.
- 1
- –
- –
- –
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it.