Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks19 from public data
- V-ReasonBench: Toward Unified Reasoning Benchmark Suite for Video Generation Models21
V-ReasonBench is introduced, a benchmark designed to assess video reasoning across four key dimensions: structured problem-solving, spatial cognition, pattern-based inference, and physical dynamics, which offers a unified and reproducible framework for measuring video reasoning.
- DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers19
Dynamic Sequence Parallelism (DSP) is proposed as a novel abstraction of sequence parallelism that dynamically switches the parallel dimension among all sequences according to the computation stage with efficient resharding strategy and offers significant reductions in communication costs, adaptability across modules, and ease of implementation with minimal constraints.
- Concerto: Automatic Communication Optimization and Scheduling for Large-Scale Deep Learning17
The evaluation shows Concerto can match or outperform state-of-the-art parallel frameworks, including Megatron-LM, JAX/XLA, DeepSpeed, and Alpa, all of which include extensive hand-crafted optimization.
- FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters17
FastFold is presented, an efficient implementation of AlphaFold for both training and inference of Dynamic Axial Parallelism (DAP), and a series of low-level optimizations aimed at reducing communication, computation, and memory costs are implemented.
- DD-Ranking: Rethinking the Evaluation of Dataset Distillation14
DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.
- HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices13
A novel approach that presents a principled framework for heterogeneous parallel computing using CPUs and GPUs and employs heterogeneous parallel computing and asynchronous overlap for LLMs to mitigate I/O bottlenecks to achieve low-latency LLMs inference on resource-constrained devices is introduced.
- Dynamic Vision Mamba8
The proposed method, Dynamic Vision Mamba (DyVM), effectively reduces FLOPs with minor performance drops and generalizes well across different Mamba vision model architectures and different vision tasks.
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline Parallelism6
HelixPipe is a novel pipeline parallelism for long sequence transformer training that introduces attention parallel partition, which schedules attention computations of different micro batches across different pipeline stages in parallel, reducing pipeline bubbles.
- HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing5
HY-WU (Weight Unleashing), a memory-first adaptation framework that shifts adaptation pressure away from overwriting a single shared parameter point, is proposed, which implements functional (operator-level) memory as a neural module: a generator that synthesizes weight updates on-the-fly from the instance condition, yielding instance-specific operators without test-time optimization.
- 5
- Faster Vision Mamba is Rebuilt in Minutes via Merged Token Re-training5
This work shows how simple and effective the fast recovery can be achieved at minute-level, in particular, a 35.9% accuracy spike over 3 epochs of training on Vim-Ti, and that a quick round of retraining after token merging yeilds robust results across various compression ratios.
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training4
StarTrail is proposed, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements, which significantly surpasses state-of-the-art methods that support Long sequence lengths.
- AutoChunk: Automated Activation Chunk for Memory-Efficient Long Sequence Inference4
AutoChunk is proposed, an automatic and adaptive compiler system that efficiently reduces activation memory for long sequence inference by chunk strategies and can outperform state-of-the-art methods by a large margin.
- BESIII production with distributed computing4
A dataset-based data transfer system has been developed to support data movements among sites and the experience to cope with lack of grid experience and low manpower among the BESIII community is shown.
- Multi-VO support in IHEP's distributed computing environment4
The BESDIRAC platform is extended to support multi-VO scenario, instead of setting up a self-contained distributed computing environment for each VO, which makes DIRAC as a service for the community of those experiments.
- ResearchGPT: Benchmarking and Training LLMs for End-to-End Computer Science Research Workflows2
This work contributes CS-54k, a high-quality corpus of scientific Q&A pairs in computer science, built from 14k CC-licensed papers, and derives two complementary subsets: CS-4k, a carefully curated benchmark for evaluating AI's ability to assist scientific research, and CS-50k, a large-scale training dataset.
- –
- CDIO: Cross-Domain Inference Optimization with Resource Preference Prediction for Edge-Cloud Collaboration–
CDIO, a cross-domain inference optimization framework designed for edge-cloud collaboration, can predict resource preference types by analyzing spatial complexity and processing requirements of the task and guide resource allocation in the edge-cloud system.
- Resources monitoring and automatic management system for multi-VO distributed computing system–
A resources monitoring and automatic management system based on Resource Status System of DIRAC, composed of three parts: information collection, status decision and automatic control, and information display is presented.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.