Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Xinchao WangSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works26 from public data
- OoD-Bench: Quantifying and Understanding Two Dimensions of Out-of-Distribution Generalization145
This work first identifies and measure two distinct kinds of distribution shifts that are ubiquitous in various datasets, and compares OoD generalization algorithms across two groups of benchmarks, revealing their strengths on one shift as well as limitations on the other shift.
- Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding107
Dimple is proposed, the first Discrete Diffusion Multimodal Large Language Model (DMLLM), which validates the feasibility and advantages of DMLLM and enhances its inference efficiency and controllability.
- Discrete Diffusion in Large Language and Multimodal Models: A Survey66
This work traces the historical development of dLLMs and dMLLMs, formalize the underlying mathematical frameworks, list commonly-used modeling methods, and categorize representative models, and analyzes key techniques for training, inference, quantization.
- 47
- Distribution Shift Inversion for Out-of-Distribution Prediction34
This paper proposes a portable Distribution Shift Inversion (DSI) algorithm, in which, before being fed into the prediction model, the OoD testing samples are first linearly combined with additional Gaussian noise and then transferred back towards the training distribution using a diffusion model trained only on the source distribution.
- Introducing Visual Perception Token into Multimodal Large Language Model28
This work proposes the concept of Visual Perception Token, aiming to empower MLLM with a mechanism to control its visual perception processes, and designs two types of Visual Perception Tokens, termed the Region Selection Token and the Vision Re-Encoding Token.
- Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms19
This survey provides a unified and up-to-date overview of safety in Vision-Language-Action models, defining the scope of VLA safety, distinguishing it from text-only LLM safety and classical robotic safety, and reviewing the foundations of VLA models, including architectures, training paradigms, and inference mechanisms.
- 18
- Revisiting Self-Supervised Heterogeneous Graph Learning from Spectral Clustering Perspective15
This paper theoretically revisiting SHGL from the spectral clustering perspective and introducing a novel framework enhanced by rank and dual consistency constraints that integrates node-level and cluster-level consistency constraints that concurrently capture invariant and clustering information to facilitate learning in downstream tasks.
- Neural Lineage12
This paper introduces a novel task known as neural lineage detection, aiming at discovering lineage relationships between parent and child models, and proposes a learning-free and learning-based methods that out-perform the baseline in various learning settings and are adaptable to a variety of visual models.
- Through the Dual-Prism: A Spectral Perspective on Graph Data Augmentation for Graph Classification7
This work investigates the interplay between graph properties, their augmentation, and their spectral behavior, and found that keeping the low-frequency eigenvalues unchanged can preserve the critical properties at a large scale when generating augmented graphs.
- 6
- 5
- Every Step Counts: Decoding Trajectories as Authorship Fingerprints of dLLMs5
This work shows that the decoding mechanism of dLLMs not only enhances model utility but also can be used as a powerful tool for model attribution, and proposes a novel information extraction scheme called the Directed Decoding Map (DDM), which captures structural relationships between decoding steps and better reveals model-specific behaviors.
- NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors4
A simple and training-free framework, No-Language-Hallucination Decoding, NoLan is proposed, which refines the output distribution by dynamically suppressing language priors, modulated based on the output distribution difference between multimodal and text-only inputs.
- HG-Adapter: Improving Pre-Trained Heterogeneous Graph Neural Networks with Dual Adapters4
A unified framework is proposed that combines two new adapters with potential labeled data extension to improve the generalization of pre-trained HGNN models and designs dual structure-aware adapters to adaptively fit task-related homogeneous and heterogeneous structural information.
- 4
- Refinement Provenance Inference: Detecting LLM-Refined Training Prompts from Model Behavior3
RePro is proposed, a logit-based provenance framework that fuses teacher-forced likelihood features with logit-ranking signals and consistently attains strong performance and transfers well across refiners, suggesting that it exploits refiner-agnostic distribution shifts rather than rewrite-style artifacts.
- 2
- 1
- 1
- Vertical Retargeting for Stereoscopic Images via Stereo Seam Carving1
This article proposes two seam coupling strategies for vertical retargeting, namely, real mapping and virtual mapping, and guarantees valid and geometrically consistent stereo seam pairs to be found in the horizontal direction.
- Multi-Level Collaboration in Model Merging–
This paper theoretically establishes a performance correlation between merging and ensembling and finds that even when previous restrictions are not met, there is still a way for model merging to attain a near-identical and superior performance similar to that of ensembling.
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.