Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks27 from public data
- MTU-Net: Multilevel TransUNet for Space-Based Infrared Tiny Ship Detection316
A vision Transformer (ViT) convolutional neural network (CNN) hybrid encoder to extract multilevel features and a FocalIoU loss to achieve both target localization and shape description are designed.
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-Resolution175
Upscale-A-Video is introduced, a text-guided latent diffusion framework for video upscaling that surpasses existing methods in both synthetic and real-world benchmarks, as well as in AI-generated videos, showcasing impressive visual realism and temporal consistency.
- Nighttime Smartphone Reflective Flare Removal Using Optical Center Symmetry Prior50
This work proposes an optical center symmetry prior, which suggests that the reflective flare and light source are always symmetrical around the lens's optical center, and creates the first reflective flare removal dataset called BracketFlare, which contains diverse and realistic reflective flare patterns.
- 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere19
This work represents per-view 4D attributes in a minimally factorized form by decomposing them into base geometry and time-dependent relative motion, and introduces a novel encode-once, query-anywhere and anytime paradigm for 4D reconstruction from monocular videos.
- 3DEnhancer: Consistent Multi-View Diffusion for 3D Enhancement18
A novel 3D enhancement pipeline, dubbed 3DENHANCER, which employs a multi-view latent diffusion model to enhance coarse 3D inputs while preserving multi-view consistency and significantly outperforms existing methods.
- 12
- SPTrack: Spectral Similarity Prompt Learning for Hyperspectral Object Tracking12
A spectral similarity prompt learning approach for hyperspectral object tracking (SPTrack) is proposed, which introduces a spectral matching map based on spectral similarity, which converts 3D hyperspectral data with different spectral bands into single-channel hotmaps, thus enabling cross-spectral domain generalization.
- Sparse-Gated RGB-Event Fusion for Small Object Detection in the Wild9
This work proposes a novel RGB-Event fusion framework that leverages the complementary strengths of RGB and event modalities for enhanced small object detection and introduces a Temporal Multi-Scale Attention Fusion (TMAF) module to encode motion cues from event streams at multiple temporal scales, thereby enhancing the saliency of small object features.
- 7
- 7
- Paint Bucket Colorization Using Anime Character Color Design Sheets6
This work introduces inclusion matching, which allows the network to understand the inclusion relationships between segments, rather than relying solely on direct visual correspondences, and significantly improves performance in both keyframe colorization and consecutive frame colorization.
- 5
- 4
- 2
- 2
- 2
- Self-attention-based multidimensional feature aggregation neural network for OFDM channel estimation1
A multi-dimensional feature aggregation channel estimation network (FACENet) based on self-attention based on self-attention to improve pilot-based channel estimation in orthogonal frequency division multiplexing (OFDM) systems is proposed.
- 1
- 1
- 1
- –
- In-Flight Aircraft Detection in Satellite Videos: A Benchmark Dataset and Temporal Snake Mixing Network–
The Temporal Snake Mixing Network (TSMNet) is proposed, which learns continuous spatial offsets across consecutive frames to adaptively adjust sampling positions, achieving temporal feature alignment and suppressing irrelevant responses in aircraft motion.
- Instance Segmentation as Tracking: A New Paradigm for Multi-small-Object Tracking with Event Cameras–
- –
- Accurate 3D vision perception and dynamic interaction technology for collaborative robots based on multimodal fusion–
A multimodal fusion-based 3D vision perception and interaction control framework for collaborative robots based on an attention mechanism based on a combination of model predictive control and adaptive impedance control is proposed, verifying the effectiveness of the framework in achieving accurate perception and compliant interaction in complex dynamic scenes.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.