Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works16 from public data

TitleCited by
  • Efficient Reasoning Models: A Survey

    Sicheng Feng, Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2025

    This survey aims to provide a comprehensive overview of recent advances in efficient reasoning by categorizing existing works into three key directions: shorter - compressing lengthy CoTs into concise yet effective reasoning chains; smaller - developing compact language models with strong reasoning capabilities through techniques such as knowledge distillation, model compression techniques, and reinforcement learning.

    92
  • DMax: Aggressive Parallel Decoding for dLLMs

    Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu, Xinchao Wang

    arXiv · 2026

    DMax mitigates error accumulation in parallel decoding, enabling aggressive decoding parallelism while preserving generation quality, and represents each intermediate decoding state as an interpolation between the predicted token embedding and the mask embedding, enabling iterative self-revising in embedding space.

    18
  • Sponge Tool Attack: Stealthy Denial-of-Efficiency against Tool-Augmented Agentic Reasoning

    Qi Li, Xinchao Wang

    arXiv · 2026

    This work designs STA as an iterative, multi-agent collaborative framework with explicit rewritten policy control, and generates benign-looking prompt rewrites from the original one with high semantic fidelity, which results in substantial computational overhead while remaining stealthy by preserving the original task semantics and user intent.

    13
  • dVoting: Fast Voting for dLLMs

    Sicheng Feng, Zigeng Chen, Xinyin Ma, Gongfan Fang, Xinchao Wang

    arXiv · 2026

    This work introduces dVoting, a fast voting technique that boosts reasoning capability without training, with only an acceptable extra computational overhead, and leverages the arbitrary-position generation capability of dLLMs to perform iterative refinement by sampling, identifying uncertain tokens via consistency analysis, regenerating them through voting, and repeating this process until convergence.

    8
  • Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model

    Xinyin Ma, Julius Berner, Chao Liu, Arash Vahdat, Weili Nie, Xinchao Wang

    arXiv · 2026

    Flex-Forcing is introduced, a unified training and inference framework that enables a video diffusion model to seamlessly operate under both bidirectional and autoregressive generation regimes, and achieves consistently better video quality, long-video stability than strong baselines with a rigid inference schedule.

    5
  • BadWAM: When World-Action Models Dream Right but Act Wrong

    Qi Li, Xingyi Yang, Xinchao Wang

    arXiv · 2026

    BadWAM is introduced, a unified framework for modeling and evaluating World-Action Drift Attacks: a new class of WAM-specific adversarial attacks that use small visual perturbations to break the alignment between what a WAM imagines and what it executes.

    5
  • dOPSD: On-Policy Self-Distillation for Diffusion Language Models

    Phuong Tuan Dat, Qi Li, Xinchao Wang

    arXiv · 2026

    dOPSD derives the teacher's privilege directly from the student's own denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher's advantage emerges from the model's own decoding process.

    5
  • On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment

    Bo Yin, Qi Li, Xinchao Wang

    arXiv · 2026

    FATE is proposed, an on-policy self-evolving framework that transforms verifier-scored failures into repair supervision without expert demonstrations, and introduces Pareto-Front Policy Optimization (PFPO), combining supervised warmup with Pareto-aware policy optimization to preserve safety-utility trade-offs.

    4
  • Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs

    Haiquan Lu, Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2026

    This work proposes Mix-Quant, a simple and effective phase-aware quantization framework for fast agentic inference that combines phase-aware algorithmic quantization with hardware-efficient NVFP4 execution to alleviate the inference bottleneck in LLM agents.

    4
  • dMoE: dLLMs with Learnable Block Experts

    Sicheng Feng, Zigeng Chen, Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2026

    The central idea of dMoE is to aggregate token-level expert distributions within each block into a unified block-level expert distribution, which is then used to guide expert routing in a more coherent manner, thereby mitigating the memory-bound bottleneck.

    4
  • Encapsulating Knowledge in One Prompt

    Qi Li, Runpeng Yu, Xinchao Wang

    arXiv · 2024

    4
  • Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models

    Guangnian Wan, Qi Li, Gongfan Fang, Xinyin Ma, Xinchao Wang

    arXiv · 2026

    This work introduces a backdoor defense framework for MDLMs named DiSP (Diffusion Self-Purification), driven by a key observation: selectively masking certain vision tokens at inference time can neutralize a backdoored model's trigger-induced behaviors and restore normal functionality.

    2
  • MixReasoning: Switching Modes to Think

    Haiquan Lu, Gongfan Fang, Xinyin Ma, Qi Li, Xinchao Wang

    arXiv · 2025

    This work proposes MixReasoning, a framework that dynamically adjusts the depth of reasoning within a single response, which shortens reasoning length and substantially improves efficiency without compromising accuracy.

    2
  • AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration

    Ruhao Liu, Weiqi Huang, Qi Li, Xinchao Wang

    arXiv · 2026

    This work proposes AutoMIA, an agentic framework that reformulates membership inference as an automated process of self-exploration and strategy evolution, and enables a systematic, model-agnostic traversal of the attack search space.

    –
  • Q-ARVD: Quantizing Autoregressive Video Diffusion Models

    Siao Tang, Xinyin Ma, Gongfan Fang, Xingyi Yang, Xinchao Wang

    arXiv · 2026

    Q-ARVD is proposed, a novel framework for accurate ARVD quantization that incorporates a final-quality aware frame-weighting mechanism into the quantization objective, and introduces an outlier-aware adaptive dual-scale quantization that automatically detects the presence and quantity of outlier channels for an arbitrary layer, and isolates them to protect normal channels.

    –
  • Ungeneralizable Examples

    Jingwen Ye, Xinchao Wang

    arXiv · 2024

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.