Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works31 from public data

TitleCited by
  • Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

    Yuxi Xie, Anirudh Goyal, Wenyue Zheng, Min‐Yen Kan, Timothy Lillicrap, Kenji Kawaguchi, Michael Shieh

    arXiv · 2024

    This work leverages Monte Carlo Tree Search (MCTS) to iteratively collect preference data, utilizing its look-ahead ability to break down instance-level rewards into more granular step-level signals, to enhance consistency in intermediate steps.

    265
  • Diffusion Language Models are Super Data Learners

    Jinjie Ni, Qian Liu, Longxu Dou, Du, Chao, Zili Wang, Hang Yan, Tianyu Pang, Michael Shieh

    arXiv · 2025

    A Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoregressive (AR) models by training for more epochs by attributing the gains to three compounding factors: any-order modeling, super-dense compute from iterative bidirectional denoising, and built-in Monte Carlo augmentation.

    54
  • ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning

    Jiawei Gu, Hao, Yunzhuo, Huichen Will Wang, Linjie Li, Michael Shieh, Yejin Choi, Ranjay Krishna, Yu Cheng

    arXiv · 2025

    ThinkMorph is built, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement that learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic.

    53
  • Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling

    Yiran Zhao, Wenyue Zheng, Tianle Cai, Xuan Long, Kenji Kawaguchi, Anirudh Goyal, Michael Shieh

    neural information processing systems · 2024

    This work studies a new algorithm called probe sampling, a mechanism that dynamically determines how similar a smaller draft model's predictions are to the target model's predictions for prompt candidates, which is able to accelerate other prompt optimization techniques and adversarial methods.

    46
  • MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use

    Zijian Wu, Liu, Xiangyan, Xinyuan Zhang, Chen, Lingjun, Fanqing Meng, Du, Lingxiao, Yiran Zhao, Zhang, Fanshi, +7 more

    arXiv · 2025

    A comprehensive evaluation of cutting-edge LLMs using a minimal agent framework that operates in a tool-calling loop, significantly surpassing those in previous MCP benchmarks and highlighting the stress-testing nature of MCPMark.

    40
  • Reasoning Robustness of LLMs to Adversarial Typographical Errors

    Eric Gan, Yiran Zhao, Liying Cheng, Mao Yancan, Anirudh Goyal, Kenji Kawaguchi, Min‐Yen Kan, Michael Shieh

    Conference on Empirical Methods in Natural Language Processing (EMNLP) · 2024

    An Adversarial Typo Attack algorithm is designed that iteratively samples typos for words that are important to the query and selects the edit that is most likely to succeed in attacking, which shows that LLMs are sensitive to minimal adversarial typographical changes.

    40
  • LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization

    Guanzheng Chen, Xin Li, Michael Shieh, Lidong Bing

    arXiv · 2025

    LongPO-trained models can achieve results on long-context benchmarks comparable to, or even surpassing, those of superior LLMs (e.g., GPT-4-128K) that involve extensive long-context annotation and larger parameter scales.

    26
  • Training Optimal Large Diffusion Language Models

    Jinjie Ni, Qian Liu, Chao Du, Longxu Dou, Hang Yan, Zili Wang, Tianyu Pang, Michael Shieh

    arXiv · 2025

    24
  • Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw

    Zijun Wang, Haoqin Tu, Letian Zhang, Hardy Chen, Juncheng Wu, Xiangyan Liu, Zhenlong Yuan, Tianyu Pang, +6 more

    arXiv · 2026

    The first real-world safety evaluation of OpenClaw is presented and the CIK taxonomy is introduced, which unifies an agent's persistent state into three dimensions, i.e., Capability, Identity, and Knowledge, for safety analysis, showing that the vulnerabilities are inherent to the agent architecture.

    22
  • LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards

    Guanzheng Chen, Michael Shieh, Lidong Bing

    arXiv · 2026

    LongRLVR is introduced to augment the sparse answer reward with a dense and verifiable context reward, demonstrating that explicitly rewarding the grounding process is a critical and effective strategy for unlocking the full reasoning potential of LLMs in long-context applications.

    18
  • Efficient Process Reward Model Training via Active Learning

    Keyu Duan, Z. Liu, Mao, Xin, Tianyu Pang, Changyu Chen, Qiguang Chen, Michael Shieh, Longxu Dou

    arXiv · 2025

    This work proposes an active learning approach, ActPRM, which proactively selects the most uncertain samples for training, substantially reducing labeling costs, and further advances the actively trained PRM by filtering over 1M+ math reasoning trajectories with ActPRM, retaining 60% of the data.

    18
  • Self-Evaluation as a Defense Against Adversarial Attacks on LLMs

    H. Alex Brown, Leon Lin, Kenji Kawaguchi, Michael Shieh

    arXiv · 2024

    This work introduces a defense against adversarial attacks on LLMs utilizing self-evaluation, using pre-trained models to evaluate the inputs and outputs of a generator model, significantly reducing the cost of implementation in comparison to other, finetuning-based methods.

    18
  • RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

    Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao, Haocheng Lu, Mengkang Hu, Michael Shieh

    arXiv · 2026

    It is suggested that current agents can make useful data-centric discoveries but cannot yet translate feedback into consistent improvements, and RSIBench-Data provides a measurable, auditable testbed for the research capabilities required for recursive self-improvement.

    17
  • ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents

    Fanqing Meng, Lingxiao Du, Zijian Wu, Guanzheng Chen, Xiangyan Liu, Jiaqi Liao, Chonghe Jiang, Zhenglin Wan, +22 more

    arXiv · 2026

    The benchmark, evaluation harness, and construction pipeline are released to support reproducible coworker-agent evaluation and turn-level analysis shows that performance drops after the first exogenous environment update, highlighting adaptation to changing state as a key open challenge.

    17
  • Unnatural Languages Are Not Bugs but Features for LLMs

    Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu, Tianle Cai, Longxu Dou, +4 more

    arXiv · 2025

    This work demonstrates that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models, and demonstrates that models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language.

    8
  • Prompt Optimization via Adversarial In-Context Learning

    Xuan Long, Yiran Zhao, H. Alex Brown, Yuxi Xie, James Zhao, Nancy F. Chen, Kenji Kawaguchi, Michael Shieh, +1 more

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2024

    6
  • ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning

    Yiran Zhao, Yaoqi Ye, Xiang Liu, Michael Shieh, Trung Bui

    arXiv · 2026

    This work proposes ImageEdit-R1, a multi-agent framework for intelligent image editing that leverages reinforcement learning to coordinate high-level decision-making across a set of specialized, pretrained vision-language and generative agents.

    5
  • Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models

    Hongfu Liu, Yuxi Xie, Ye Wang, Michael Shieh

    Conference on Empirical Methods in Natural Language Processing (EMNLP) · 2024

    5
  • Scaling GUI Agents with Visual State Transitions

    Xiangyan Liu, Kaixin Li, Haonan Wang, Biao Wu, Meng Fang, Longxu Dou, Chao Du, Michael Shieh, +1 more

    arXiv · 2026

    Empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.

    3
  • Gym-V: A Unified Vision Environment System for Agentic Vision Research

    Fanqing Meng, Lingxiao Du, Jiawei Gu, J G Liao, Linjie Li, Z Z Wu, Xiangyan Liu, Ziqi Zhao, +4 more

    arXiv · 2026

    Gym-V, a unified platform of 179 procedurally generated visual environments across 10 domains with controllable difficulty, is introduced, finding that observation scaffolding is more decisive for training success than the choice of RL algorithm, with captions and game rules determining whether learning succeeds at all.

    3
  • In-Context Reinforcement Learning for Tool Use in Large Language Models

    Yaoqi Ye, Yiran Zhao, Keyu Duan, Zeyu Zheng, Kenji Kawaguchi, Cihang Xie, Michael Shieh

    arXiv · 2026

    In-Context Reinforcement Learning (ICRL), an RL-only framework that eliminates the need for SFT by leveraging few-shot prompting during the rollout stage of RL, is proposed, demonstrating its effectiveness as a scalable, data-efficient alternative to traditional SFT-based pipelines.

    2
  • Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

    Hardy Chen, Nancy Lau, Haoqin Tu, Shuo Yan, Xiangyan Liu, Zijun Wang, Juncheng Wu, Michael Shieh, +3 more

    arXiv · 2026

    This work builds AgentPressureBench, a 34-task machine-learning repository benchmark spanning three input modalities, and collects 1326 multi-round trajectories from 13 coding agents, finding that stronger models have higher exploitation rates and are supported by a significant Spearman rank correlation.

    2
  • Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents

    Zhuoming Chen, Xinrui Zhong, Qilong Feng, Ranajoy Sadhukhan, Yang Zhou, Michael Shieh, Zhihao Jia, Beidi Chen

    arXiv · 2026

    Vortex is a system that combines a Python-embedded frontend language atop a page-centric tensor abstraction for expressing a broad range of sparse attention algorithms, with an efficient backend tightly integrated into modern LLM serving stacks.

    1
  • ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs

    Eric Gan, H. Alex Brown, David Herel, Kenji Kawaguchi, Min‐Yen Kan, Michael Shieh

    Findings of the Association for Computational Linguistics: ACL 2026 · 2026

    It is shown that current MLLMs rely primarily on coarse temporal cues and struggle with fine-grained visual reasoning, revealing a large gap between current models and human-level multimodal understanding in comics.

    1
  • Single Character Perturbations Break LLM Alignment

    Leon Lin, H. Alex Brown, Kenji Kawaguchi, Michael Shieh

    Proceedings of the AAAI Conference on Artificial Intelligence · 2025

    1

Show all 31 works

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.