Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks7 from public data
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning56
WorldMM is a novel multimodal memory agent that constructs and retrieves from multiple complementary memories, encompassing both textual and visual representations, that significantly outperforms existing baselines across five long video question-answering benchmarks.
- Self-Refining Video Sampling12
This work presents self-refining video sampling, a simple method that uses a pre-trained video generator trained on large-scale datasets as its own self-refiner to enable iterative inner-loop refinement at inference time without any external verifier or additional training.
- 1
- –
- Safe Few-Step Generation via Velocity Editing–
VESFlow is proposed, a training-free safety method tailored to flow matching with extremely few sampling steps that steers the trajectory toward safe outputs while leaving the conditioning prompt unchanged and introduces a risk score-based filtering that bypasses velocity editing to reduce computational cost while preserving benign prompt generation.
- –
- Confidence-Aware Tool Orchestration for Robust Video Understanding–
Robust-TO is proposed, an agentic video understanding framework that explicitly integrates per-frame trustworthiness into every stage of reasoning and defines a confidence-cost GRPO reward that jointly optimizes correctness, evidence reliability, and efficiency.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.