Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

Possible advisorsa guess from early papers, not confirmed

  • Weiming Lü

    Possible advisor · last author on 6 of their early first-author papers, 2020–2022

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works21 from public data

TitleCited by
  • LLM-Pruner: On the Structural Pruning of Large Language Models

    Xinyin Ma, Gongfan Fang, Xinchao Wang

    arXiv · 2023

    This work explores LLM compression in a task-agnostic manner, which aims to preserve the multi-task solving and language generation ability of the original LLM, and adopts structural pruning that selectively removes non-critical coupled structures based on gradient information, maximally preserving the majority of the LLM's functionality.

    1,001
  • DepGraph: Towards Any Structural Pruning

    Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, Xinchao Wang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    This work proposes a general and fully automatic method, Dependency Graph (DepGraph), to explicitly model the dependency between layers and comprehensively group coupled parameters for pruning, and demonstrates that, even with a simple norm-based criterion, the proposed method consistently yields gratifying performances.

    566
  • Structural Pruning for Diffusion Models

    Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2023

    The essence of Diff-Pruning is encapsulated in a Taylor expansion over pruned timesteps, a process that disregards non-contributory diffusion steps and ensembles informative gradients to identify important weights.

    258
  • Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition

    Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang, Wen Wang, Weiming Lü

    Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) · 2021

    This work proposes a two-stage entity identifier that effectively utilizes the boundary information of entities and partially matched spans during training and outperforms previous state-of-the-art models.

    229
  • A Trigger-Sense Memory Flow Framework for Joint Entity and Relation Extraction

    Yongliang Shen, Xinyin Ma, Yechun Tang, Weiming Lü

    Proceedings of the Web Conference 2021 · 2021

    A Trigger-Sense Memory Flow Framework (TriMF) is presented, which builds a memory module to remember category representations learned in entity recognition and relation extraction tasks and designs a multi-level memory flow attention mechanism to enhance the bi-directional interaction between entity recognitionand relation extraction.

    73
  • TinyFusion: Diffusion Transformers Learned Shallow

    Gongfan Fang, K. Li, Xinyin Ma, Xinchao Wang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    TinyFusion, a depth pruning method designed to remove redundant layers from diffusion transformers via end-to-end learning, and explicitly models and optimizes the post-fine-tuning performance of pruned models.

    47
  • SlimSAM: 0.1% Data Makes Segment Anything Slim

    Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2023

    SlimSAM is introduced, a novel data-efficient SAM compression method that achieves superior performance with extremely less training data, and is encapsulated in the alternate slimming framework which effectively enhances knowledge inheritance under severely limited training data availability and exceptional pruning ratio.

    40
  • Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient

    Zigeng Chen, Xinyin Ma, Gongfan Fang, Xinchao Wang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    CoDe partition the multi-scale inference process into a seamless collaboration between a large model and a small model, specializing in generating low-frequency content at smaller scales, while the smaller model serves as the ’refiner’, solely focusing on predicting high-frequency details at larger scales.

    37
  • Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional Network

    Zeyun Tang, Yongliang Shen, Xinyin Ma, Wei Hong Xu, Jiale Yu, Weiming Lü

    International Joint Conference on Artificial Intelligence · 2020

    Inspired by the human reasoning processing, a path-based graph with reasoning paths which extracted from supporting documents is introduced, which contains a new question-aware gating mechanism to regulate the usefulness of information propagating across documents and add question information during reasoning.

    31
  • Introducing Visual Perception Token into Multimodal Large Language Model

    Runpeng Yu, Xinyin Ma, Xinchao Wang

    arXiv · 2025

    This work proposes the concept of Visual Perception Token, aiming to empower MLLM with a mechanism to control its visual perception processes, and designs two types of Visual Perception Tokens, termed the Region Selection Token and the Vision Re-Encoding Token.

    28
  • Adversarial Self-Supervised Data-Free Distillation for Text Classification

    Xinyin Ma, Yongliang Shen, Gongfan Fang, Chen Chen, Chenghao Jia, Weiming Lü

    Conference on Empirical Methods in Natural Language Processing (EMNLP) · 2020

    This work proposes a novel two-stage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT), and introduces a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge.

    27
  • Isomorphic Pruning for Vision Models

    Gongfan Fang, Xinyin Ma, Michael Bi Mi, Xinchao Wang

    Lecture notes in computer science · 2024

    19
  • MuVER: Improving First-Stage Entity Retrieval with Multi-View Entity Representations

    Xinyin Ma, Yong Jiang, Nguyễn Bách, Tao Wang, Zhongqiang Huang, Fei Huang, Weiming Lü

    Conference on Empirical Methods in Natural Language Processing (EMNLP) · 2021

    This work proposes Multi-View Entity Representations (MuVER), a novel approach for entity retrieval that constructs multi-view representations for entity descriptions and approximates the optimal view for mentions via a heuristic searching method.

    18
  • Enrich cross-lingual entity links for online wikis via multi-modal semantic matching

    Weiming Lü, Peng Wang, Xinyin Ma, Wei Hong Xu, Chen Chen

    Information Processing & Management · 2020

    This work proposes a novel approach, called MMSM, to enrich cross-lingual links for online Wikis, which jointly trains two novel end-to-end neural matching models, Entity Description Matching Model and Entity Image Matching model, which can utilize entity description and images for the cross-lingsual entity matching.

    11
  • DeepCache: Accelerating Diffusion Models for Free

    Xinyin Ma, Gongfan Fang, Xinchao Wang

    arXiv · 2023

    7
  • SynET: Synonym Expansion using Transitivity

    Jiale Yu, Yongliang Shen, Xinyin Ma, Chenghao Jia, Chen Chen, Weiming Lü

    Findings of the Association for Computational Linguistics: EMNLP 2020 · 2020

    A novel approach named SynET is proposed, which considers both the contexts of two given synonym pairs, and introduces an auxiliary task to reduce the impact of noisy sentences, and proposes a Multi-Perspective Entity Matching Network to match entities from multiple perspectives.

    5
  • Prompting to Distill: Boosting Data-Free Knowledge Distillation via Reinforced Prompt

    Xinyin Ma, Xinchao Wang, Gongfan Fang, Yongliang Shen, Weiming Lü

    International Joint Conference on Artificial Intelligence · 2022

    4
  • Boosting Cross-lingual Entity Alignment with Textual Embedding

    Wei Hong Xu, Chen Chen, Chenghao Jia, Yongliang Shen, Xinyin Ma, Weiming Lü

    Lecture notes in computer science · 2020

    This paper proposes two textual embedding models called Cross-TextGCN and Cross- textMatch to embed description for each entity in multilingual knowledge graph to boost the cross-lingual entity alignment model.

    4
  • Region-Aware Test-Time Scaling for Compositional Image Generation

    Mingzhu Shen, Peng Ye, Xinyin Ma, Gongfan Fang, Christos-Savvas Bouganis, Yiren Zhao, Xinchao Wang

    Lecture notes in computer science · 2026

    –
  • VGEdit: Unlocking Video Generation Priors for Reasoning-Informed Image Editing

    Haiquan Lu, Gongfan Fang, Xinyin Ma, Xinchao Wang

    Lecture notes in computer science · 2026

    –
  • Remix-DiT: Mixing Diffusion Transformers for Multi-Expert Denoising

    Gongfan Fang, Xinyin Ma, Xinchao Wang

    neural information processing systems · 2024

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.