Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks16 from public data
- P2T: Pyramid Pooling Transformer for Scene Understanding321
This paper proposes to adapt pyramid pooling to Multi-Head Self-Attention (MHSA) in the vision transformer, simultaneously reducing the sequence length and capturing powerful contextual features, in a universal vision transformer backbone, dubbed Pyramid Pooling Transformer (P2T).
- EDN: Salient Object Detection via Extremely-Downsampled Network284
This work introduces an Extremely-Downsampled Network (EDN), which employs an extreme downsampling technique to effectively learn a global view of the whole image, leading to accurate salient object localization and construct the Scale-Correlated Pyramid Convolution (SCPC) to accomplish better multi-level feature fusion.
- MobileSal: Extremely Efficient RGB-D Salient Object Detection176
This article proposes an implicit depth restoration (IDR) technique to strengthen the mobile networks’ feature representation capability for RGB-D SOD, and proposes compact pyramid refinement (CPR) for efficient multi-level feature aggregation to derive salient objects with clear boundaries.
- 165
- 148
- Vision Transformers with Hierarchical Attention103
This paper proposes hierarchical MHSA (H-MHSA), a novel approach that computes sell-attention in a hierarchical fashion, and builds a family of hierarchical-attention-based transformer networks, namely HAT-Net, which provides a new perspective for vision transformers.
- DOTS: Decoupling Operation and Topology in Differentiable Architecture Search60
The proposed Decouple the Operation and Topology Search (DOTS), which decouples the topology representation from operation weights and makes an explicit topology search, and is an effective solution for differentiable NAS.
- Scoot: A Perceptual Metric for Facial Sketches52
The results suggest that “spatial structure” and “co-occurrence texture” are two generally applicable perceptual features in face sketch synthesis, and the first largest scale human-perception-based sketch database that can evaluate how well a metric consistent with human perception is evaluated.
- Low-Resolution Self-Attention for Semantic Segmentation43
The Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs is introduced and the effectiveness of the approach is demonstrated by building the LRFormer, a vision transformer with an encoder-decoder structure.
- Revisiting Computer-Aided Tuberculosis Diagnosis39
A large-scale dataset, namely the Tuberculosis X-ray (TBX11 K) dataset, is established, which contains 11 200 chest X-ray (CXR) images with corresponding bounding box annotations for TB areas and a strong baseline, SymFormer, is proposed, for simultaneous CXR image classification and TB infection area detection.
- A Survey and Evaluation of Adversarial Attacks in Object Detection32
This article presents a novel taxonomic framework for categorizing adversarial attacks specific to object detection architectures, synthesizes existing robustness metrics, and provides a comprehensive empirical evaluation of state-of-the-art attack methodologies on popular object detection models, including both traditional detectors and modern detectors with vision-language pretraining.
- 25
- 16
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment14
This work proposes a coupled dual-branch offset learning paradigm that explicitly learns feature and class offsets to dynamically refine both class representations and spatial image features and constructs an efficient semantic segmentation network, OffSeg.
- 10
- RefOnce: Distilling References into a Prototype Memory for Referring Camouflaged Object Detection–
A Ref-COD framework that distills references into a class-prototype memory during training and synthesizes a reference vector at inference via a query-conditioned mixture of prototypes, and a bidirectional attention alignment module that adapts both the query features and the class representation.
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.