Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Suggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works33 from public data
- Hallucination of Multimodal Large Language Models: A Survey466
A comprehensive analysis of the phenomenon of hallucination in multimodal large language models (MLLMs), also known as Large Vision-Language Models (LVLMs), offering a detailed overview of the underlying causes, evaluation benchmarks, metrics, and strategies developed to address this issue.
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent278
This work develops a vision-language-action model in digital world, namely ShowUI, which features the following innovations: UI-Guided Visual Token Selection to reduce computational costs by formulating screenshots as an UI connected graph, adaptively identifying their redundant relationship and serve as the criteria for token selection during self-attention blocks.
- Unsupervised Multi-Source Domain Adaptation for Person Re-Identification96
The proposed method outperforms state-of-the-art UDA person re-ID methods by a large margin, and even achieves comparable performance to the supervised approaches without any post-processing techniques.
- Show, Recall, and Tell: Image Captioning with Recall Mechanism74
Inspired by pointing mechanism in text summarization, this paper adopts a soft switch to balance the generated-word probabilities between SG and RWS in the CIDEr optimization step, and introduces an individual recalled-word reward (WR) to boost training.
- 63
- Explain Me the Painting: Multi-Topic Knowledgeable Art Description Generation58
This work presents a framework to bring art closer to people by generating comprehensive descriptions of fine-art paintings by modules the generated sentences according to three artistic topics and enhances each description with external knowledge.
- ASSISTGUI: Task-Oriented Desktop Graphical User Interface Automation49
An advanced Actor-Critic Embodied Agent framework is proposed, which incorporates a sophisticated GUI parser driven by an LLM-agent and an enhanced reasoning mechanism adept at handling lengthy procedural tasks that outshine existing methods in performance.
- Adaptive Slot Attention: Object Discovery with Dynamic Slot Number38
An adaptive slot attention (AdaSlot) mecha-nism that dynamically determines the optimal number of slots based on the content of the data is introduced that exhibits the capability to dynamically adapt the slot number according to each instance's complexity.
- Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance36
A scalable data generation pipeline is introduced that transforms existing video editing pairs into high-fidelity training quadruplets, leveraging image generative models to create synthesized reference scaffolds, and a unified editing architecture is proposed, Kiwi-Edit, that synergizes learnable queries and latent visual features for reference semantic guidance.
- World-VLA-Loop: Closed-Loop Learning of Video World Model and VLA Policy36
World-VLA-Loop substantially improves VLA performance while reducing reliance on costly physical interaction and feeding rollouts from each improved policy back to augment and fine-tune the world model.
- 32
- 31
- 23
- Enhancing Emotional Experience by Building Emotional Virtual Characters in VR Volleyball Games22
This article proposes to enhance emotional experience by building emotional virtual characters in VR volleyball games that cannot only arouse their emotions but also express their facial expressions according to the game situation.
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models20
EVOLVE-VLA is introduced, a test-time training framework enabling VLAs to continuously adapt through environment interaction with minimal or zero task-specific demonstrations, and represents a critical step toward VLAs that truly learn and adapt, moving beyond static imitation toward continuous self-improvements.
- Object-Centric Multiple Object Tracking15
This paper proposes a video object-centric model for MOT that consists of an index-merge module that adapts the object-centric slots into detection outputs and an object memory module that builds complete object prototypes to handle occlusions.
- Unsupervised Open-Vocabulary Object Localization in Videos15
A method that first localizes objects in videos via a slot attention approach and then assigns text to the obtained slots is proposed, which is effectively the first unsupervised approach that yields good results on regular video benchmarks.
- Bring Your Own Character: A Holistic Solution for Automatic Facial Animation Generation of Customized Characters14
A deep learning model was first trained to retarget the facial expression from input face images to virtual human faces by estimating the blendshape coefficients and a practical toolkit was developed using Unity 3D, making it compatible with the most popular VR applications.
- DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic Forgetting13
Experimental results on diverse datasets consistently demonstrate that DoFIT excels in cross-domain collaborative training and exhibits significant advantages over conventional FIT methods in alleviating catastrophic forgetting.
- Skip \n: A Simple Method to Reduce Hallucination in Large Vision-Language Models11
A new perspective is proposed, suggesting that the inherent biases in LVLMs might be a key factor in hallucinations, and a simple method is proposed to effectively mitigate the hallucination of LVLMs by skipping the output of '\n'.
- Factorized Visual Tokenization and Generation11
Factorized Quantization (FQ) is introduced, a novel approach that revitalizes VQ-based tokenizers by decomposing a large codebook into multiple independent sub-codebooks, enabling more efficient and scalable visual tokenization.
- 9
- Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation5
This work identifies a critical diversity trap and proposes Anchor-Centric Adaptation (ACA), a two-stage framework that first stabilizes a policy skeleton through repeated demonstrations at core anchors, then selectively expands coverage to high-risk boundaries via teacher-forced error mining and constrained residual updates.
- AssistEditor: Multi-Agent Collaboration for GUI Workflow Automation in Video Creation5
A novel PC-Copilot, AssistEditor, that focuses on automating the video editing workflow, and significantly streamlines the video editing process, making advanced editing accessible to users with varying levels of expertise.
- SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens4
SIGMA (Selective-Interleaved Generation with Multi-Attribute Tokens), a unified post-training framework that enables interleaved multi-condition generation within diffusion transformers, introduces selective multi-attribute tokens, which allow the model to interpret and compose multiple visual conditions in an interleaved text-image sequence.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.