Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works19 from public data

TitleCited by
  • V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models

    Yang Luo, Xuanlei Zhao, Baijiong Lin, Lingting Zhu, Liyao Tang, Yuqi Liu, Ying-Cong Chen, Shengju Qian, +2 more

    arXiv · 2025

    V-ReasonBench is introduced, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics, which offers a unified and reproducible framework for measuring video reasoning.

    21
  • DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers

    Xuanlei Zhao, Shenggan Cheng, Chen, Chang, Zangwei Zheng, Ziming Liu, Zheming Yang, Yang You

    arXiv · 2024

    Dynamic Sequence Parallelism (DSP) is proposed as a novel abstraction of sequence parallelism that dynamically switches the parallel dimension among all sequences according to the computation stage with efficient resharding strategy and offers significant reductions in communication costs, adaptability across modules, and ease of implementation with minimal constraints.

    19
  • Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning

    Shenggan Cheng, Sheng-Jie Lin, Lansong Diao, Hao Wu, Siyu Wang, Chang Si, Ziming Liu, Xuanlei Zhao, +3 more

    ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) · 2025

    The evaluation shows Concerto can match or outperform state-of-the-art parallel frameworks, including Megatron-LM, JAX/XLA, DeepSpeed, and Alpa, all of which include extensive hand-crafted optimization.

    17
  • FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters

    Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Tian Zheng, Ruidong Wu, Xiwen Zhang, Jian Xun Peng, +1 more

    ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming · 2024

    FastFold is presented, an efficient implementation of AlphaFold for both training and inference of Dynamic Axial Parallelism (DAP), and a series of low-level optimizations aimed at reducing communication, computation, and memory costs are implemented.

    17
  • DD-Ranking: Rethinking the Evaluation of Dataset Distillation

    Zekai Li, Xinhao Zhong, Samir Khaki, Zhiyuan Liang, Yuhao Zhou, Mingjia Shi, Ziqiao Wang, Xuanlei Zhao, +22 more

    arXiv · 2025

    DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.

    14
  • HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices

    Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, Yang You

    arXiv · 2024

    A novel approach that presents a principled framework for heterogeneous parallel computing using CPUs and GPUs and employs heterogeneous parallel computing and asynchronous overlap for LLMs to mitigate I/O bottlenecks to achieve low-latency LLMs inference on resource-constrained devices is introduced.

    13
  • Dynamic Vision Mamba

    Mengxuan Wu, Zekai Li, Zhiyuan Liang, Li, Moyang, Xuanlei Zhao, Samir Khaki, Zhu, Zheng, Xiaojiang Peng, +4 more

    arXiv · 2025

    The proposed method, Dynamic Vision Mamba (DyVM), effectively reduces FLOPs with minor performance drops and generalizes well across different Mamba vision model architectures and different vision tasks.

    8
  • HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism

    Geng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu, Yang You

    ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming · 2026

    HelixPipe is a novel pipeline parallelism for long sequence transformer training that introduces attention parallel partition, which schedules attention computations of different micro batches across different pipeline stages in parallel, reducing pipeline bubbles.

    6
  • HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing

    Mengxuan Wu, Xuanlei Zhao, Ziqiao Wang, Ruicheng Feng, Zhangyang Wang, Kai Wang

    arXiv · 2026

    HY-WU (Weight Unleashing), a memory-first adaptation framework that shifts adaptation pressure away from overwriting a single shared parameter point, is proposed, which implements functional (operator-level) memory as a neural module: a generator that synthesizes weight updates on-the-fly from the instance condition, yielding instance-specific operators without test-time optimization.

    5
  • REPA Works Until It Doesn’t: Early-Stopped, Holistic Alignment Supercharges Diffusion Training

    Ziqiao Wang, Wangbo Zhao, Yuhao Zhou, Zekai Li, Zhiyuan Liang, Mingjia Shi, Xuanlei Zhao, Pengfei Zhou, +4 more

    neural information processing systems · 2025

    5
  • Faster Vision Mamba is Rebuilt in Minutes via Merged Token Re-training

    Mingjia Shi, Yuhao Zhou, RQ Yu, Zekai Li, Zhiyuan Liang, Xuanlei Zhao, Xiaojiang Peng, Shanmukha Ramakrishna Vedantam, +3 more

    arXiv · 2024

    This work shows how simple and effective the fast recovery can be achieved at minute-level, in particular, a 35.9% accuracy spike over 3 epochs of training on Vim-Ti, and that a quick round of retraining after token merging yeilds robust results across various compression ratios.

    5
  • StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training

    Ziming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao, Kai Wang, Xuanlei Zhao, J. Demmel, Yang You

    neural information processing systems · 2025

    StarTrail is proposed, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements, which significantly surpasses state-of-the-art methods that support Long sequence lengths.

    4
  • AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference

    Xuanlei Zhao, Shenggan Cheng, Guangyang Lu, Jiarui Fang, Haotian Zhou, Bin Jia, Ziming Liu, Yang You

    arXiv · 2024

    AutoChunk is proposed, an automatic and adaptive compiler system that efficiently reduces activation memory for long sequence inference by chunk strategies and can outperform state-of-the-art methods by a large margin.

    4
  • BESIII production with distributed computing

    Xianmei Zhang, T Yan, Xuanlei Zhao, Z T Ma, Xiaolang Yan, Tao Lin, Z. Y. Deng, W D Li, +4 more

    Journal of Physics Conference Series · 2015

    A dataset-based data transfer system has been developed to support data movements among sites and the experience to cope with lack of grid experience and low manpower among the BESIII community is shown.

    4
  • Multi-VO support in IHEP's distributed computing environment

    T Yan, B. Suo, Xuanlei Zhao, Xianmei Zhang, Z T Ma, Xiaolang Yan, Tao Lin, Z. Y. Deng, +5 more

    Journal of Physics Conference Series · 2015

    The BESDIRAC platform is extended to support multi-VO scenario, instead of setting up a self-contained distributed computing environment for each VO, which makes DIRAC as a service for the community of those experiments.

    4
  • ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows

    Penghao Wang, Yuhao Zhou, Ming-Ta Wu, Ziheng Qin, Zhu, Bangyuan, Huang, Shengbin, Xuanlei Zhao, Panpan Zhang, +7 more

    arXiv · 2025

    This work contributes CS-54k, a high-quality corpus of scientific Q&A pairs in computer science, built from 14k CC-licensed papers, and derives two complementary subsets: CS-4k, a carefully curated benchmark for evaluating AI's ability to assist scientific research, and CS-50k, a large-scale training dataset.

    2
  • Training Variable Long Sequences with Data-Centric Parallel

    Geng Zhang, Xuanlei Zhao, Kai Wang, Yang You

    arXiv · 2026

    –
  • CDIO: Cross-Domain Inference Optimization with Resource Preference Prediction for Edge-Cloud Collaboration

    Zheming Yang, Wen Ji, Qi Guo, Dieli Hu, Chang Zhao, Xiaowei Li, Xuanlei Zhao, Zhao Yi, +2 more

    arXiv · 2025

    CDIO, a cross-domain inference optimization framework designed for edge-cloud collaboration, can predict resource preference types by analyzing spatial complexity and processing requirements of the task and guide resource allocation in the edge-cloud system.

    –
  • Resources monitoring and automatic management system for multi-VO distributed computing system

    J. Chen, I. Pelevanyuk, Y Sun, Alexey Zhemchugov, T Yan, Xuanlei Zhao, X M Zhang

    Journal of Physics Conference Series · 2017

    A resources monitoring and automatic management system based on Resource Status System of DIRAC, composed of three parts: information collection, status decision and automatic control, and information display is presented.

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.