Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Weiming LüSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works21 from public data
- LLM-Pruner: On the Structural Pruning of Large Language Models1,001
This work explores LLM compression in a task-agnostic manner, which aims to preserve the multi-task solving and language generation ability of the original LLM, and adopts structural pruning that selectively removes non-critical coupled structures based on gradient information, maximally preserving the majority of the LLM's functionality.
- DepGraph: Towards Any Structural Pruning566
This work proposes a general and fully automatic method, Dependency Graph (DepGraph), to explicitly model the dependency between layers and comprehensively group coupled parameters for pruning, and demonstrates that, even with a simple norm-based criterion, the proposed method consistently yields gratifying performances.
- Structural Pruning for Diffusion Models258
The essence of Diff-Pruning is encapsulated in a Taylor expansion over pruned timesteps, a process that disregards non-contributory diffusion steps and ensembles informative gradients to identify important weights.
- Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition229
This work proposes a two-stage entity identifier that effectively utilizes the boundary information of entities and partially matched spans during training and outperforms previous state-of-the-art models.
- A Trigger-Sense Memory Flow Framework for Joint Entity and Relation Extraction73
A Trigger-Sense Memory Flow Framework (TriMF) is presented, which builds a memory module to remember category representations learned in entity recognition and relation extraction tasks and designs a multi-level memory flow attention mechanism to enhance the bi-directional interaction between entity recognitionand relation extraction.
- TinyFusion: Diffusion Transformers Learned Shallow47
TinyFusion, a depth pruning method designed to remove redundant layers from diffusion transformers via end-to-end learning, and explicitly models and optimizes the post-fine-tuning performance of pruned models.
- SlimSAM: 0.1% Data Makes Segment Anything Slim40
SlimSAM is introduced, a novel data-efficient SAM compression method that achieves superior performance with extremely less training data, and is encapsulated in the alternate slimming framework which effectively enhances knowledge inheritance under severely limited training data availability and exceptional pruning ratio.
- Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient37
CoDe partition the multi-scale inference process into a seamless collaboration between a large model and a small model, specializing in generating low-frequency content at smaller scales, while the smaller model serves as the ’refiner’, solely focusing on predicting high-frequency details at larger scales.
- Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional Network31
Inspired by the human reasoning processing, a path-based graph with reasoning paths which extracted from supporting documents is introduced, which contains a new question-aware gating mechanism to regulate the usefulness of information propagating across documents and add question information during reasoning.
- Introducing Visual Perception Token into Multimodal Large Language Model28
This work proposes the concept of Visual Perception Token, aiming to empower MLLM with a mechanism to control its visual perception processes, and designs two types of Visual Perception Tokens, termed the Region Selection Token and the Vision Re-Encoding Token.
- Adversarial Self-Supervised Data-Free Distillation for Text Classification27
This work proposes a novel two-stage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT), and introduces a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge.
- 19
- MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations18
This work proposes Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.
- Enrich cross-lingual entity links for online wikis via multi-modal semantic matching11
This work proposes a novel approach, called MMSM, to enrich cross-lingual links for online Wikis, which jointly trains two novel end-to-end neural matching models, Entity Description Matching Model and Entity Image Matching model, which can utilize entity description and images for the cross-lingsual entity matching.
- 7
- SynET: Synonym Expansion using Transitivity5
A novel approach named SynET is proposed, which considers both the contexts of two given synonym pairs, and introduces an auxiliary task to reduce the impact of noisy sentences, and proposes a Multi-Perspective Entity Matching Network to match entities from multiple perspectives.
- 4
- Boosting Cross-lingual Entity Alignment with Textual Embedding4
This paper proposes two textual embedding models called Cross-TextGCN and Cross- textMatch to embed description for each entity in multilingual knowledge graph to boost the cross-lingual entity alignment model.
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.