Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Yadong MuSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works17 from public data
- Tiny hand gesture recognition without localization via a deep convolutional network116
A deep convolutional neural network is proposed to directly classify hand gestures in images without any segmentation or detection stage that could discard the irrelevant not-hand areas.
- MoE-FFD: Mixture of Experts for Generalized and Parameter-Efficient Face Forgery Detection70
Mixture-of-Experts modules for Face Forgery Detection (MoE-FFD) is introduced, a generalized yet parameter-efficient ViT-based approach that achieves state-of-the-art face forgery detection performance with significantly reduced parameter overhead in cross-dataset, cross-manipulation, and robustness evaluations.
- Dense Events Grounding in Video41
This work proposes Dense Events Propagation Network (DepNet), which adaptively aggregates temporal and semantic information of dense events into a compact set through a second-order attention pooling, then selectively propagates the aggregated information to each single event with soft attention.
- Open-Set Deepfake Detection: A Parameter-Efficient Adaptation Method With Forgery Style Mixture27
This work develops a parameter-efficient ViT-based detection model that includes lightweight forgery feature extraction modules and enables the model to extract global and local forgery clues simultaneously, representing an important step toward open-set Deepfake detection in the wild.
- Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video Grounding25
This paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding and proposes a cross-modal mutual learning framework and trains a sophisticated teacher model to learn collaboratively from the multi-modal videos.
- Learning Sample Importance for Cross-Scenario Video Temporal Grounding17
This paper proposes a novel method called Debiased Temporal Language Localizer (Debias-TLL) to prevent the model from naively memorizing the biases and enforce it to ground the query sentence based on true inter-modal relationship.
- Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets Consistency15
This work proposes an omnipotent distillation algorithm with large language models (LLM) to paraphrase the language query and distill the teacher model to a lightweight student model by enforcing the consistency between the localization results of the paraphrased sentence and the original one.
- Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event Localization15
An expectation-maximization algorithm is proposed that optimizes the pseudo-label acquisition and localization model in a coarse-to-fine manner and performs reasonably well compared to the state-of-the-art supervised methods.
- 13
- Vid-Morp: Video Moment Retrieval Pretraining from Unlabeled Videos in the Wild6
The ReCorrect algorithm, which comprises two main phases: semantics-guided refinement and memory-consensus correction, enhances the pseudo labels by leveraging semantic similarity with video frames to clean out unpaired data and make initial adjustments to temporal boundaries.
- Learning 3-D Human Pose Estimation from Catadioptric Videos6
This work explores a novel way of obtaining gigantic 3-D human pose data without manual annotations by jointly harnessing the epipolar geometry and human skeleton priors and can boil down to an optimization problem over two sets of 2-D estimations.
- SimBase: A Simple Baseline for Temporal Video Grounding4
This paper designs SimBase, a network that leverages lightweight, one-dimensional temporal convolutional layers instead of complex temporal structures, and achieves state-of-the-art results on two large-scale datasets.
- ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos3
This work introduces ActivityForensics, the first large-scale benchmark for localizing manipulated activity in videos, and proposes Temporal Artifact Diffuser (TADiff), a simple yet effective baseline that exposes artifact cues through a diffusion-based feature regularizer.
- 3
- SAKED: Mitigating Hallucination in Large Vision-Language Models via Stability-Aware Knowledge Enhanced Decoding1
Stability-Aware Knowledge-Enhanced Decoding (SAKED), which introduces a layer-wise Knowledge Stability Score (KSS) to quantify knowledge stability throughout the model, and achieves state-of-the-art performance for hallucination mitigation on various models, tasks, and benchmarks.
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.