Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks11 from public data
- Interactive Change-Aware Transformer Network for Remote Sensing Image Change Captioning50
An Interactive Change-Aware Transformer Network (ICT-Net) is proposed, able to extract and incorporate the most critical changes of interest in each encoder layer to improve change description generation and achieves a state-of-the-art performance.
- 34
- Toward Attribute-Controlled Fashion Image Captioning16
The results demonstrate that the proposed approach outperforms existing fashion image captioning models as well as conventional captioning methods and further validate the effectiveness of the proposed method on the MSCOCO and Flickr30K captioning datasets and achieve competitive performance.
- PromptSR: Cascade Prompting for Lightweight Image Super-Resolution10
The experimental results demonstrate the superiority of the PromptSR method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations.
- Towards Blind Bitstream-corrupted Video Recovery: A Visual Foundation Model-driven Framework6
This paper proposes the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with recovery model, which is adapted to different types of corruption and bitstream-level prompts and introduces a novel Corruption-aware Feature Completion (CFC) module.
- Boundary Voting Network for Ambiguity-Aware Timestamp-Supervised Action Segmentation5
The boundary voting network is introduced that mitigates feature ambiguity by hierarchically propagating video-level global prior knowledge into local action-transiting regions by generating key action representations as votes throughout the video and targeting action-transiting regions.
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning4
This paper proposes the explainable deepfake video detection (EDVD) task and designs the EDVD-LLaMA multimodal, a large language model (MLLM) reasoning framework, which provides traceable reasoning processes alongside accurate detection results and trustworthy explanations.
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Open-vocabulary Grounded Situation Recognition3
This paper proposes Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations.
- From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition1
This paper proposes Multimodal Interactive Prompt Distillation (MIPD), a novel framework that distills enriched multimodal knowledge from the foundation model, enabling the student Ov-GSR model to recognize unseen situations and be better aware of rare situations.
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.