Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Traced back 39 generations →Advisors
- Wang XinchaoFrom a thesis record ↗
Works50 from public data
- DepGraph: Towards Any Structural Pruning566
This work proposes a general and fully automatic method, Dependency Graph (DepGraph), to explicitly model the dependency between layers and comprehensively group coupled parameters for pruning, and demonstrates that, even with a simple norm-based criterion, the proposed method consistently yields gratifying performances.
- DeepCache: Accelerating Diffusion Models for Free464
DeepCache capitalizes on the inherent temporal redundancy observed in the sequential denoising steps of diffusion models, which caches and retrieves features across adjacent denoising stages, thereby curtailing redundant computations.
- Up to 100x Faster Data-Free Knowledge Distillation110
This work introduces an efficacious scheme, termed as FastDFKD, that allows us to accelerate DFKD by a factor of orders of magnitude and proposes to learn a meta-synthesizer that seeks common features in training data as the initialization for the fast data synthesis.
- 102
- Efficient Reasoning Models: A Survey92
This survey aims to provide a comprehensive overview of recent advances in efficient reasoning by categorizing existing works into three key directions: shorter - compressing lengthy CoTs into concise yet effective reasoning chains; smaller - developing compact language models with strong reasoning capabilities through techniques such as knowledge distillation, model compression techniques, and reinforcement learning.
- 72
- Knowledge Amalgamation from Heterogeneous Networks by Common Feature Learning61
This paper proposes a common feature learning scheme, in which the features of all teachers are transformed into a common space and the student is enforced to imitate them all so as to amalgamate the intact knowledge.
- Contrastive Model Invertion for Data-Free Knolwedge Distillation50
Experiments demonstrate that Contrastive Model Inversion not only generates more visually plausible instances than the state of the arts, but also achieves significantly superior performance when the generated data are used for knowledge distillation.
- TinyFusion: Diffusion Transformers Learned Shallow47
TinyFusion, a depth pruning method designed to remove redundant layers from diffusion transformers via end-to-end learning, and explicitly models and optimizes the post-fine-tuning performance of pruned models.
- Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient37
CoDe partition the multi-scale inference process into a seamless collaboration between a large model and a small model, specializing in generating low-frequency content at smaller scales, while the smaller model serves as the ’refiner’, solely focusing on predicting high-frequency details at larger scales.
- 30
- Adversarial Self-Supervised Data-Free Distillation for Text Classification27
This work proposes a novel two-stage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT), and introduces a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge.
- 26
- SparseD: Sparse Attention for Diffusion Language Models19
Experimental results demonstrate that SparseD achieves lossless acceleration, delivering up to $1.50\times over FlashAttention at a 64k context length with 1,024 denoising steps, and establish SparseD as a practical and efficient solution for deploying DLMs in long-context applications.
- 19
- DMax: Aggressive Parallel Decoding for dLLMs18
DMax mitigates error accumulation in parallel decoding, enabling aggressive decoding parallelism while preserving generation quality, and represents each intermediate decoding state as an interpolation between the predicted token embedding and the mask embedding, enabling iterative self-revising in embedding space.
- PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning18
This paper proposes PointLoRA, a simple yet effective method that combines low-rank adaptation (LoRA) with multi-scale token selection to efficiently fine-tune point cloud models, reducing the need for tunable parameters while enhancing global feature capture.
- Knowledge Amalgamation for Object Detection With Transformers18
A more effective KA scheme for Transformer-based object detection models is explored, considering the architecture characteristics of Transformers, and a hint is generated within the sequence-level amalgamation by concatenating teacher sequences instead of redundantly aggregating them to a fixed-size one as previous KA approaches.
- 14
- 11
- 10
- dVoting: Fast Voting for dLLMs8
This work introduces dVoting, a fast voting technique that boosts reasoning capability without training, with only an acceptable extra computational overhead, and leverages the arbitrary-position generation capability of dLLMs to perform iterative refinement by sampling, identifying uncertain tokens via consistency analysis, regenerating them through voting, and repeating this process until convergence.
- 7
- 7
- 6
- A Comprehensive Study of Structural Pruning for Vision Models5
The first comprehensive benchmark, termed Prun-ingBench, for structural pruning is presented, which employs a unified and consistent framework for evaluating the effectiveness of diverse structural pruning techniques and provides easily implementable interfaces to facilitate the implementation of future pruning methods.
- 5
- Federated Selective Aggregation for Knowledge Amalgamation5
Experimental results demonstrate that FedSA effectively amalgamates knowledge from decentralized models and achieves competitive performance to centralized baselines.
- Invisible Safety Threat: Malicious Finetuning for LLM via Steganography4
This paper finetuned the model to understand and apply a steganographic technique, and produces steganographic malicious outputs in response to hidden malicious prompts, while the user interface displays only a fully benign cover interaction.
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs4
This work proposes Mix-Quant, a simple and effective phase-aware quantization framework for fast agentic inference that combines phase-aware algorithmic quantization with hardware-efficient NVFP4 execution to alleviate the inference bottleneck in LLM agents.
- dMoE: dLLMs with Learnable Block Experts4
The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner, thereby mitigating the memory-bound bottleneck.
- 4
- 3
- 3
- Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models2
This work introduces a backdoor defense framework for MDLMs named DiSP (Diffusion Self-Purification), driven by a key observation: selectively masking certain vision tokens at inference time can neutralize a backdoored model's trigger-induced behaviors and restore normal functionality.
- MixReasoning: Switching Modes to Think2
This work proposes MixReasoning, a framework that dynamically adjusts the depth of reasoning within a single response, which shortens reasoning length and substantially improves efficiency without compromising accuracy.
- 2
- Rethinking Token Reduction for Large Vision-Language Models1
This paper begins by formulating token reduction as a learnable compression mapping, unifying existing formats such as pruning and merging into a single learning objective, and introduces a data-efficient training paradigm capable of learning optimal compression mappings with limited computational costs.
- Taming the Phantom: Token-Asymmetric Filtering for Hallucination Mitigation in Large Vision-Language Models1
Token-Asymmetric Filtering (TAF) is introduced—a training-free, plug-and-play method that modulates intermediate attention maps in LVLMs and significantly mitigates hallucinations across a range of state-of-the-art LVLMs.
- 1
- –
- –
- Q-ARVD: Quantizing Autoregressive Video Diffusion Models–
Q-ARVD is proposed, a novel framework for accurate ARVD quantization that incorporates a final-quality aware frame-weighting mechanism into the quantization objective, and introduces an outlier-aware adaptive dual-scale quantization that automatically detects the presence and quantity of outlier channels for an arbitrary layer, and isolates them to protect normal channels.
- –
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.