Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Li-Hua XieSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works40 from public data
- MCD: Diverse Large-Scale Multi-Campus Dataset for Robot Perception98
A comprehensive dataset named MCD (Multi-Campus Dataset), featuring a wide range of sensing modalities, high-accuracy ground truth, and diverse challenging environments across three Eurasian university campuses, and introduces semantic annotations of 29 classes over 59k sparse NRE lidar scans across three domains.
- ARID: A New Dataset for Recognizing Action in the Dark98
It is shown that current action recognition models and frame enhancement methods may not be effective solutions for the task of action recognition in dark videos.
- 60
- Aligning Correlation Information for Domain Adaptation in Action Recognition54
This work proposes a novel adversarial correlation adaptation network (ACAN) to align action videos by aligning pixel correlations and introduces a novel HMDB-ARID dataset with a larger domain shift caused by a larger statistical difference between domains.
- Source-Free Video Domain Adaptation by Learning Temporal Consistency for Action Recognition48
A novel Attentive Temporal Consistent Network (ATCoN) is proposed to address SFVDA by learning temporal consistency, guaranteed by two novel consistency objectives, namely feature consistency and source prediction consistency, performed across local temporal features.
- Video Unsupervised Domain Adaptation with Deep Learning: A Comprehensive Survey33
This article surveys recent progress in VUDA with deep learning with the motivation of VUDA, followed by its definition, and recent progress of methods for both closed-set VUDA and VUDA under different scenarios, and current benchmark datasets for VUDA research.
- Multi-Modal Continual Test-Time Adaptation for 3D Semantic Segmentation33
An adaptive dual-stage mechanism to generate reliable cross-modal predictions by attending to the reliable modality based on the class-wise feature-centroid distance in the latent space and design class-wise momentum queues that capture confident target features for adaptation while stochastically restoring pseudo-source features to revisit source knowledge.
- Partial Video Domain Adaptation with Partial Adversarial Temporal Attentive Network33
A novel Partial Adversarial Temporal Attentive Network (PATAN) is proposed to address the PVDA problem by utilizing both spatial and temporal features for filtering source-only classes and constructs effective overall temporal features by attending to local temporal features that contribute more toward the class filtration process.
- Multi-Source Video Domain Adaptation With Temporal Attentive Moment Alignment Network32
A novel Temporal Attentive Moment Alignment Network (TAMAN) which aims for effective feature transfer by dynamically aligning both spatial and temporal feature moments with high local classification confidence and low disparity between global and local feature discrepancies is proposed.
- Segregator: Global Point Cloud Registration with Semantic and Geometric Cues31
Gaussian distribution-based translation and rotation invariant measurements (G-TRIMs) are proposed to conduct the consistency check and further constrain the problem size.
- Outram: One-shot Global Localization via Triangulated Scene Graph and Global Outlier Pruning30
This work proposes a hierarchical one-shot localization algorithm called Outram that leverages substructures of 3D scene graphs for locally consistent correspondence searching and global substructure-wise outlier pruning and demonstrates the capability in a variety of scenarios in multiple large-scale outdoor datasets.
- MoPA: Multi-Modal Prior Aided Domain Adaptation for 3D Semantic Segmentation29
Valid Ground-based Insertion (VGI) is developed to rectify the imbalance supervision signals by inserting prior rare objects collected from the wild while avoiding introducing artificial artifacts that lead to trivial solutions.
- AIR-Embodied: An Efficient Active 3DGS-based Interaction and Reconstruction Framework with Embodied Large Language Model10
AIR-Embodied is a novel framework that integrates embodied AI agents with large-scale pretrained multi-modal language models to improve active 3DGS reconstruction, providing a robust solution to challenges in active 3D reconstruction.
- SGBA: Semantic Gaussian Mixture Model-Based LiDAR Bundle Adjustment9
SGBA, a LiDAR BA scheme that models the environment as a semantic Gaussian mixture model (GMM) without predefined feature types, is proposed, offering a comprehensive and general representation adaptable to various environments.
- Self-Supervised Video Representation Learning by Video Incoherence Detection8
A novel self-supervised method that leverages incoherence detection for video representation learning and introduces intravideo contrastive learning to maximize the mutual information between incoherent clips from the same raw video.
- Effective action recognition with embedded key point shifts8
This paper proposes a novel temporal feature extraction module, named Key Point Shifts Embedding Module ($KPSEM$), to adaptively extract channel-wise key point shifts across video frames without key point annotation for temporal features extraction.
- Going Deeper into Recognizing Actions in Dark Environments: A Comprehensive Benchmark Study7
This work focuses on the task of action recognition in dark environments, which can be applied to fields such as surveillance and autonomous driving at night and evaluates and advances the robustness of AR models in dark environments.
- PNL: Efficient long-range dependencies extraction with pyramid non-local module for action recognition7
Pyramid Non-Local (PNL) module is proposed, which extends the non-local block by incorporating regional correlation at multiple scales through a pyramid structured module, which upscales the effectiveness of non- local operation by attending to the interaction between different regions.
- Enhancing Scene Coordinate Regression With Efficient Keypoint Detection and Sequential Information6
This letter proposes a unified architecture for both scene encoding and salient keypoint detection, allowing the system to prioritize the encoding of informative regions and introduces a mechanism that utilizes sequential information during both mapping and relocalization.
- Interactive Test-Time Adaptation with Reliable Spatial-Temporal Voxels for Multi-Modal Segmentation6
Latte is proposed, an MM-TTA method that leverages reliable cross-modal spatial-temporal correspondences for multi-modal 3D segmentation and achieves state-of-the-art performance on three different MM-TTA benchmarks compared to previous MM-TTA or TTA methods.
- GERA: Geometric Embedding for Efficient Point Registration Analysis5
This paper proposes a novel point cloud registration network that leverages a pure MLP architecture, constructing geometric information offline, and is the first to replace 3D coordinate inputs with offline-constructed geometric encoding, improving generalization and stability.
- 5
- 4
- AIR-Embodied: Active Interactive Reconstruction for 3D Gaussian Splatting with Embodied Multimodal Agents3
AIR-Embodied is a novel framework that integrates embodied AI agents with large-scale pretrained multi-modal language models to improve active 3DGS reconstruction, providing a robust solution to challenges in active 3D reconstruction.
- 2
- TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion2
TGSFormer is proposed, a scalable Temporal Gaussian Splatting framework for embodied SSC, offering superior accuracy and scalability with significantly fewer primitives while maintaining consistent long-term scene integrity.
- UniRiT: Towards Few-Shot Non-Rigid Point Cloud Registration2
UniRiT adopts a two-step registration strategy that first aligns the centroids of the source and target point clouds and then refines the registration with non-rigid transformations, thereby significantly reducing the problem complexity.
- 2
- 2
- 1
- –
- NoisyEQA: benchmarking Embodied Question Answering with imperfect queries from non-expert users–
A NoisyEQA benchmark designed to evaluate the ability of the robot to identify and correct noisy questions and a ‘Self-Correction’ prompting mechanism to enhance EQA against noise robustness and a novel evaluation metric to measure both noise detection capability and answer quality are introduced.
- –
- –
- Gaussian Semantic Field for One-shot LiDAR Global Localization–
A one-shot LiDAR global localization algorithm featuring semantic disambiguation ability based on a lightweight tri-layered scene graph serving as a light-weight yet performant backend for one-shot localization.
- –
- Effective action recognition with fully supervised and self-supervised methods–
A novel Key Point Shift Embedding Module (KPSEM) is proposed to adaptively extract channel-wise key point shifts across video frames without key point annotation for temporal feature extraction for fully-supervised learning.
- –
- Task-based Approach to College English Teaching with Multimedia and Internet Assistance–
This paper probes the application of multimedia-and internet-assisted task-based approach to college English teaching, and discusses its theoretical and practical significance.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.