Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks54 from public data
- Weakly Supervised Semantic Point Cloud Segmentation: Towards 10× Fewer Labels266
This work proposes a weakly supervised point cloud segmentation approach which requires only a tiny fraction of points to be labelled in the training stage, made possible by learning gradient approximation and exploitation of additional spatial and color smoothness constraints.
- 3D AffordanceNet: A Benchmark for Visual Object Affordance Understanding194
Comprehensive results on the contributed dataset show the promise of visual affordance understanding as a valuable yet challenging benchmark.
- PointSAM: Pointly-Supervised Segment Anything Model for Remote Sensing Images99
This work proposes a negative prompt calibration (NPC) method based on the nonoverlapping nature of instance masks, where overlapping masks are used as negative signals to refine segmentation, and presents a novel pointly-supervised SAM (PointSAM).
- 82
- Improving the Generalization of Segmentation Foundation Model under Distribution Shift via Weakly Supervised Adaptation63
This work proposes a weakly supervised self-training architecture with anchor regularization and low-rank finetuning to improve the robustness and computation efficiency of adaptation of Segment-Anything and proves the effectiveness on 5 types of downstream segmentation tasks.
- Transformation-Invariant Network for Few-Shot Object Detection in Remote-Sensing Images48
This work proposes integrating a feature pyramid network (FPN) and utilizing prototype features to enhance query features, thereby improving existing FSOD methods and tackling the issue of spatial misalignment caused by orientation variations between the query and support images by introducing a transformation-invariant network (TINet).
- Weakly Supervised 3D Point Cloud Segmentation via Multi-Prototype Learning47
This work designs a multi-prototype classifier, each prototype serves as the classifier weights for one subclass, and proposes two constraints respectively for updating the prototypes w.r.t. all point features and for encouraging the learning of diverse prototypes.
- 3D Rigid Motion Segmentation with Mixed and Unknown Number of Models46
This work proposes a multi-model spectral clustering framework that synergistically combines multiple models (homography and fundamental matrix) together and shows that the performance can be substantially improved in this way.
- Motion Segmentation by Exploiting Complementary Geometric Models42
A multi-view spectral clustering framework that synergistically combines multiple models together that can be substantially improved on existing motion segmentation datasets, achieving state-of-the-art performance on all of them.
- COFT-AD: COntrastive Fine-Tuning for Few-Shot Anomaly Detection38
This paper proposes a novel methodology to address the challenge of FSAD which incorporates a model pre-trained on a large source dataset to initialize model weights and adopts contrastive training to fine-tune on the few-shot target domain data.
- Exploring Diversity-Based Active Learning for 3D Object Detection in Autonomous Driving37
This work investigates diversity-based active learning (AL) as a potential solution to alleviate the annotation burden, and proposes a novel acquisition function that enforces spatial and temporal diversity in the selected samples.
- A Timely Survey on Vision Transformer for Deepfake Detection33
This survey presents a timely overview of ViT-based deepfake detection models, categorized into standalone, sequential, and parallel architectures, and succinctly delineates the structure and characteristics of each model.
- 28
- Revisiting Realistic Test-Time Training: Sequential Inference and Adaptation by Anchored Clustering Regularized Self-Training27
A more effective TTT approach is introduced by regularizing self-training with anchored clustering, and the improved model is referred to as TTAC++, which is demonstrated that, under all TTT protocols, TTAC++ consistently outperforms the state-of-the-art methods on five TTT datasets.
- SemiCurv: Semi-Supervised Curvilinear Structure Segmentation22
This work proposes SemiCurv, a semi-supervised learning framework for curvilinear structure segmentation that is able to utilize such unlabelled data to reduce the labelling burden, and introduces a geometric transformation as strong data augmentation and then align segmentation predictions via a differentiable inverse transformation to enable the computation of pixel-wise consistency.
- 20
- 20
- Visual aesthetic understanding: Sample-specific aesthetic classification and deep activation map visualization20
An end-to-end convolutional neural network model which simultaneously implements aesthetic classification and understanding is proposed and it is found that dropping out ambiguous image is a special case of the sample-specific method, and also figure out that as the weights of the non-ambiguous images increase, the performance is positively affected.
- 18
- Clip-Guided Source-Free Object Detection in Aerial Images16
This work utilizes Contrastive Language-Image Pre-training (CLIP) to guide the generation of pseudo-labels, termed CLIP-guided Aggregation (CGA), and leverages CLIP’s zero-shot classification capability to aggregate its scores with the original predicted bounding boxes, enabling us to obtain refined scores for the pseudo-labels.
- Exploring Spatial Diversity for Region-Based Active Learning15
This work proposes that enforcing local spatial diversity is beneficial for active learning in this case, and to incorporate spatial diversity along with the traditional active selection criterion, e.g., data sample uncertainty, in a unified optimization framework for region-based active learning.
- Efficient and Context-Aware Label Propagation for Zero-/Few-Shot Training-Free Adaptation of Vision-Language Model11
This work dynamically constructs a graph over text prompts, few-shot examples, and test samples, using label propagation for inference without task-specific tuning, and introduces a context-aware feature re-weighting mechanism to improve task adaptation accuracy.
- 10
- 9
- 6
- 5
- AnalogAgent: Self-Improving Analog Circuit Design Automation with LLM Agents4
This work proposes AnalogAgent, a training-free agentic framework that integrates an LLM-based multi-agent system (MAS) with self-evolving memory (SEM) for analog circuit design automation, and substantially strengthens open-weight models for high-quality analog circuit design automation.
- 4
- 4
- 4
- 3
- Robust Video Background Identification by Dominant Rigid Motion Estimation3
This paper proposes an efficient local-to-global method to identify background, based on the assumption that as long as there is sufficient camera motion, the cumulative background features will have the largest amount of trajectories.
- Underwater 3D images reconstruction via guided diffusion for streak tube imaging LiDAR2
An RGBD diffusion model for denoising reconstructed images via generative refinement achieves higher accuracy and lower errors compared to existing methods, with a depth resolution surpassing 0.55 mm under a 0.14 m water attenuation length, significantly enhancing underwater STIL imaging performance.
- Box-Level Class-Balanced Sampling For Active Object Detection2
This work proposes a class-balanced sampling strategy to select more objects from minority classes for labelling so as to make the final training data, ground truth labels obtained by AL and pseudo labels, more class-balanced to train a better model.
- Diverse and consistent multi-view networks for semi-supervised regression2
Diverse and Consistent Multi-view Networks (DiCoM) is presented—a novel semi-supervised regression technique based on a multi-view learning framework that combines diversity with consistency based on underlying probabilistic graphical assumptions.
- 2
- OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation1
OPD-Evolver internalizes high-value experience and memory management, enabling OPD-Evolver-9B to challenge giant counterparts such as Qwen3.5-397B-A17B and Step-3.5-Flash, pointing beyond memory-augmented agents toward genuinely qualified agent evolvers.
- Improving Adversarial Robustness for 3D Point Cloud Recognition at Test-Time through Purified Self-Training1
The proposed test-time purified self-training strategy is complementary to purification based method in handling continually changing adversarial attacks on the testing data stream and Adaptive thresholding and feature distribution alignment are introduced to improve the robustness of self-training.
- 1
- –
- –
- –
- Exploiting Vision Language Model for Training-Free 3D Point Cloud Understanding via Improved Graph Score Propagation–
GSP++ is presented, a graph-based inference framework that exploits the manifold structure of test-time point clouds to refine VLM scores without additional training, and introduces a self-training strategy that selects high-confidence positive and negative samples and assigns them calibrated pseudo scores to further stabilize propagation.
- –
- –
- –
- –
- –
- Exploiting Vision Language Model for Training-Free 3D Point Cloud OOD Detection via Graph Score Propagation–
A novel Graph Score Propagation (GSP) method that incorporates prompt clustering and self-training negative prompting to improve OOD scoring with VLM is proposed and demonstrated that GSP consistently outperforms state-of-the-art methods across synthetic and realworld datasets for 3D point cloud OOD detection.
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.