Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks31 from public data
- Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning265
This work leverages Monte Carlo Tree Search (MCTS) to iteratively collect preference data, utilizing its look-ahead ability to break down instance-level rewards into more granular step-level signals, to enhance consistency in intermediate steps.
- Diffusion Language Models are Super Data Learners54
A Crossover: when unique data is limited, diffusion language models (DLMs) consistently surpass autoregressive (AR) models by training for more epochs by attributing the gains to three compounding factors: any-order modeling, super-dense compute from iterative bidirectional denoising, and built-in Monte Carlo augmentation.
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning53
ThinkMorph is built, a unified model fine-tuned on approximately 24K high-quality interleaved reasoning traces spanning tasks with varying visual engagement that learns to generate progressive text-image reasoning steps that concretely manipulate visual content while maintaining coherent verbal logic.
- Accelerating Greedy Coordinate Gradient and General Prompt Optimization via Probe Sampling46
This work studies a new algorithm called probe sampling, a mechanism that dynamically determines how similar a smaller draft model's predictions are to the target model's predictions for prompt candidates, which is able to accelerate other prompt optimization techniques and adversarial methods.
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use40
A comprehensive evaluation of cutting-edge LLMs using a minimal agent framework that operates in a tool-calling loop, significantly surpassing those in previous MCP benchmarks and highlighting the stress-testing nature of MCPMark.
- Reasoning Robustness of LLMs to Adversarial Typographical Errors40
An Adversarial Typo Attack algorithm is designed that iteratively samples typos for words that are important to the query and selects the edit that is most likely to succeed in attacking, which shows that LLMs are sensitive to minimal adversarial typographical changes.
- LongPO: Long Context Self-Evolution of Large Language Models through Short-to-Long Preference Optimization26
LongPO-trained models can achieve results on long-context benchmarks comparable to, or even surpassing, those of superior LLMs (e.g., GPT-4-128K) that involve extensive long-context annotation and larger parameter scales.
- 24
- Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw22
The first real-world safety evaluation of OpenClaw is presented and the CIK taxonomy is introduced, which unifies an agent's persistent state into three dimensions, i.e., Capability, Identity, and Knowledge, for safety analysis, showing that the vulnerabilities are inherent to the agent architecture.
- LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards18
LongRLVR is introduced to augment the sparse answer reward with a dense and verifiable context reward, demonstrating that explicitly rewarding the grounding process is a critical and effective strategy for unlocking the full reasoning potential of LLMs in long-context applications.
- Efficient Process Reward Model Training via Active Learning18
This work proposes an active learning approach, ActPRM, which proactively selects the most uncertain samples for training, substantially reducing labeling costs, and further advances the actively trained PRM by filtering over 1M+ math reasoning trajectories with ActPRM, retaining 60% of the data.
- Self-Evaluation as a Defense Against Adversarial Attacks on LLMs18
This work introduces a defense against adversarial attacks on LLMs utilizing self-evaluation, using pre-trained models to evaluate the inputs and outputs of a generator model, significantly reducing the cost of implementation in comparison to other, finetuning-based methods.
- RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement17
It is suggested that current agents can make useful data-centric discoveries but cannot yet translate feedback into consistent improvements, and RSIBench-Data provides a measurable, auditable testbed for the research capabilities required for recursive self-improvement.
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agents17
The benchmark, evaluation harness, and construction pipeline are released to support reproducible coworker-agent evaluation and turn-level analysis shows that performance drops after the first exogenous environment update, highlighting adaptation to changing state as a key open challenge.
- Unnatural Languages Are Not Bugs but Features for LLMs8
This work demonstrates that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usable by models, and demonstrates that models fine-tuned on unnatural versions of instruction datasets perform on-par with those trained on natural language.
- 6
- ImageEdit-R1: Boosting Multi-Agent Image Editing via Reinforcement Learning5
This work proposes ImageEdit-R1, a multi-agent framework for intelligent image editing that leverages reinforcement learning to coordinate high-level decision-making across a set of specialized, pretrained vision-language and generative agents.
- 5
- Scaling GUI Agents with Visual State Transitions3
Empirical studies show that joint dynamics optimization yields stable improvements over single-objective training, and downstream performance scales steadily with the volume of transition data.
- Gym-V: A Unified Vision Environment System for Agentic Vision Research3
Gym-V, a unified platform of 179 procedurally generated visual environments across 10 domains with controllable difficulty, is introduced, finding that observation scaffolding is more decisive for training success than the choice of RL algorithm, with captions and game rules determining whether learning succeeds at all.
- In-Context Reinforcement Learning for Tool Use in Large Language Models2
In-Context Reinforcement Learning (ICRL), an RL-only framework that eliminates the need for SFT by leveraging few-shot prompting during the rollout stage of RL, is proposed, demonstrating its effectiveness as a scalable, data-efficient alternative to traditional SFT-based pipelines.
- Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows2
This work builds AgentPressureBench, a 34-task machine-learning repository benchmark spanning three input modalities, and collects 1326 multi-round trajectories from 13 coding agents, finding that stronger models have higher exploitation rates and are supported by a significant Spearman rank correlation.
- Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents1
Vortex is a system that combines a Python-embedded frontend language atop a page-centric tensor abstraction for expressing a broad range of sparse attention algorithms, with an efficient backend tightly integrated into modern LLM serving stacks.
- ComicVQA: A Benchmark for Visual Reasoning in Multimodal LLMs1
It is shown that current MLLMs rely primarily on coarse temporal cues and struggle with fine-grained visual reasoning, revealing a large gap between current models and human-level multimodal understanding in comics.
- 1
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.