Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks156 from public data
- Semantic Flow for Fast and Accurate Scene Parsing456
This paper proposes a Flow Alignment Module (FAM) to learn Semantic Flow between feature maps of adjacent levels, and broadcast high-level features to high resolution features effectively and efficiently and exhibits superior performance over other real-time methods even on light-weight backbone networks.
- Involution: Inverting the Inherence of Convolution for Visual Recognition418
The proposed involution operator could be leveraged as fundamental bricks to build the new generation of neural networks for visual recognition, powering different deep learning models on several prevalent benchmarks, including ImageNet classification, COCO detection and segmentation, together with Cityscapes segmentation.
- Transformer-Based Visual Segmentation: A Survey342
This survey provides a thorough overview of transformer-based visual segmentation, summarizing recent advancements and presents a meta-architecture that unifies all recent transformer-based approaches.
- Improving Semantic Segmentation via Decoupled Body and Edge Supervision313
This paper proposes a new paradigm for semantic segmentation that establishes new state of the art while retaining high efficiency in inference and shows that the proposed framework with various baselines or backbone networks leads to better object inner consistency and object boundaries.
- Rethinking Mobile Block for Efficient Attention-based Models296
This work rethinks lightweight infrastructure from efficient IRB and effective components of Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMB) for lightweight model design.
- Towards Open Vocabulary Learning: A Survey285
This paper thoroughly reviews open vocabulary learning, summarizing and analyzing recent developments in the field, and juxtaposing open vocabulary learning with analogous concepts such as zero-shot learning, open-set recognition, and out-of-distribution detection.
- 255
- Gated Fully Fusion for Semantic Segmentation249
This paper proposes a new architecture, named Gated Fully Fusion(GFF), to selectively fuse features from multiple levels using gates in a fully connected way, and achieves the state of the art results on four challenging scene parsing datasets including Cityscapes, Pascal Context, COCO-stuff and ADE20K.
- TransVOD: End-to-End Video Object Detection With Spatial-Temporal Transformers242
TransVOD is presented, the first end-to-end video object detection system based on simple yet effective spatial-temporal Transformer architectures that streamline the pipeline of current VOD, effectively removing the need for many hand-crafted components for feature aggregation.
- Dual Graph Convolutional Network for Semantic Segmentation.196
The Dual Graph Convolutional Network (DGCNet) models the global context of the input feature by modelling two orthogonal graphs in a single framework, which achieves state-of-the-art results on both Cityscapes and Pascal Context datasets.
- Neural Collapse Inspired Feature-Classifier Alignment for Few-Shot Class Incremental Learning191
A neural collapse inspired framework for FSCIL that holds the neural collapse optimality and does not break the feature-classifier alignment in an incremental fashion and experiments demonstrate that the proposed framework outperforms the state-of-the-art performances.
- SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model176
This paper proposes a new image deepfake detection, localization, and explanation framework, named SIDA (Social media Image Detection, localization, and explanation Assistant), which not only discerns the authenticity of images, but also delineates tampered regions through mask prediction and provides textual explanations of the model’s judgment criteria.
- Sa2VA: Marrying SAM2 With MLLM for Dense Grounded Understanding of Images and Videos168
Experiments show that Sa2VA achieves strong performance across multiple tasks, particularly in referring video object segmentation, highlighting its potential for complex real-world applications.
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction159
An in-depth analysis of the region-language alignment in CLIP models is embarked on, which proposes an approach named CLIPSelf, which adapts the image-level recognition ability of CLIP ViT to local image regions without needing any region-text pairs.
- OMG-Seg: Is One Model Good Enough for all Segmentation?140
It is shown that OMG-Seg, a transformer-based encoder-decoder architecture with task-specific queries and outputs, can support over ten distinct segmentation tasks and yet significantly reduce computational and parameter overhead across various tasks and datasets.
- Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation125
Video K-Net is presented, a simple, strong, and unified framework for fully end-to-end video panoptic seg-mentation that achieves state-of-the-art videoPanoptic segmentation results on Citscapes-VPS and KITTI-STEP without bells and whistles and can serve as a new flexible baseline in video segmentation.
- Enhanced Boundary Learning for Glass-like Object Segmentation123
This paper proposes a novel refined differential module that outputs finer boundary cues and introduces an edge-aware point-based graph convolution network module to model the global shape along the boundary.
- End-to-End Video Object Detection with Spatial-Temporal Transformers116
TransVOD is presented, an end-to-end video object detection model based on a spatial-temporal Transformer architecture that streamline the pipeline of VOD, effectively removing the need for many hand-crafted components for feature aggregation, e.g., optical flow, recurrent neural networks, relation networks.
- Open-Vocabulary SAM: Segment and Recognize Twenty-Thousand Classes Interactively114
The Open-Vocabulary SAM is introduced, a SAM-inspired model designed for simultaneous interactive segmentation and recognition, leveraging two unique knowledge transfer modules: SAM2CLIP and CLIP2SAM.
- PointFlow: Flowing Semantics Through Points for Aerial Image Segmentation111
A point-wise affinity propagation module based on the Feature Pyramid Network (FPN) framework, named PointFlow, is proposed, which reduces the noise introduced by the background while keeping efficiency.
- RTMO: Towards High-Performance One-Stage Real-Time Multi-Person Pose Estimation103
RTMO is introduced, a one-stage pose estimation framework that seamlessly inte-grates coordinate classification by representing keypoints using dual I-D heatmaps within the YOLO architecture, achieving accuracy comparable to top-down methods while maintaining high speed.
- Multi-Task Learning With Multi-Query Transformer for Dense Prediction94
This work proposes a simple pipeline named Multi-Query Transformer (MQTransformer) that is equipped with multiple queries from different tasks to facilitate the reasoning among multiple tasks and simplify the cross-task interaction pipeline.
- An Open and Comprehensive Pipeline for Unified Object Grounding and Detection93
MM-Grounding-DINO is presented, an open-source, comprehensive, and user-friendly baseline, which is built with the MMDetection toolbox, and outperforms the Grounding-DINO-Tiny baseline.
- 86
- 83
- Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection73
A novel ViT-based ViTAD structure is instantiated, designed incrementally from both global and local perspectives, and achieves state-of-the-art results and efficiency on MVTec AD, VisA, and Uni-Medical datasets.
- Panoptic Video Scene Graph Generation73
A high-quality PVSG dataset is contributed, which consists of 400 videos (289 third-person + 111 egocentric videos) with totally 150K frames labeled with panoptic segmentation masks as well as fine, temporal scene graphs.
- 72
- EdgeSAM: Prompt-In-the-Loop Distillation for SAM71
The approach involves distilling the original ViT-based SAM image encoder into a purely CNN-based architecture, better suited for edge devices, and incorporates a lightweight module within the encoder, which achieves a 37-fold speed increase compared to the original SAM.
- Towards Semantic Equivalence of Tokenization in Multimodal LLM70
A novel dynamic Semantic-Equivalent Vision Tokenizer (SeTok) is proposed, which groups visual features into semantic units via a dynamic clustering algorithm, flexibly determining the number of tokens based on image complexity.
- Sfnet: Faster and Accurate Semantic Segmentation Via Semantic Flow70
A Flow Alignment Module (FAM) is proposed to learn Semantic Flow between feature maps of adjacent levels and broadcast high-level features to high-resolution features effectively and efficiently and integrating this FAM to a standard feature pyramid structure exhibits superior performance over other real-time methods.
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology66
TreeBench (Traceable Evidence Evaluation Benchmark), a diagnostic benchmark built on three principles: focused visual perception of subtle targets in complex scenes, traceable evidence via bounding box evaluation, and second-order reasoning to test object interactions and spatial hierarchies beyond simple object localization are proposed.
- Towards Robust Referring Image Segmentation66
This work proposes a new formulation of RIS, named Robust Referring Image Segmentation (R-RIS), which considers the negative sentence inputs besides the regular positive text inputs, and proposes a new transformer-based model, called RefSegformer, with a token-based vision and language fusion module.
- MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation61
Experimental results on the challenging LVIS long-tailed and open-vocabulary benchmarks demonstrate that MosaicFusion can significantly improve the performance of existing instance segmentation models, especially for rare and novel categories.
- UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions59
This paper proposes a high-quality open-sourced UHD-4K text-to-video dataset named UltraVideo, and expands Wan to UltraWan-1K/-4K, which can natively generate high-quality 1K/4K videos with more consistent text controllability, demonstrating the effectiveness of the data curation.
- 59
- EATFormer: Improving Vision Transformer Inspired by Evolutionary Algorithm59
A novel pyramid EATFormer backbone that only contains the proposed EA-based transformer (EAT) block, which consists of three residual parts to model multi-scale, interactive, and individual information separately and design a task-related head docked with transformer backbone to complete final information fusion more flexibly.
- 51
- Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook51
This survey presents a comprehensive review of Mamba-based methodologies in remote sensing, systematically analyzing about 120 Mamba-based remote sensing studies to construct a holistic taxonomy of innovations and applications, and establishes Mamba as a transformative framework for remote sensing analysis.
- PolyphonicFormer: Unified Query Learning for Depth-Aware Video Panoptic Segmentation51
PolyphonicFormer, a vision transformer to unify these sub-tasks under the DVPS task and lead to more robust results, achieves state-of-the-art results on two DVPS datasets, and ranks 1st on the ICCV-2021 BMTT Challenge video + depth track.
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence50
This work introduces Open-o3-Video, a non-agent framework that integrates explicit spatio-temporal evidence into video reasoning by highlighting key timestamps, objects, and bounding boxes, making the reasoning process traceable and verifiable.
- Towards Language-Driven Video Inpainting via Multimodal Large Language Models48
This work introduces a new task - language-driven video inpainting, which uses natural language instructions to guide the inpainting process, integrating Multimodal Large Language Models to understand and execute complex language-based inpaintingrequests effectively.
- 47
- Explore In-Context Learning for 3D Point Cloud Understanding47
This work introduces a novel framework, named Point-In-Context, designed especially for in-context learning in 3D point clouds, where both inputs and outputs are modeled as coordinates for each task.
- 45
- Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation44
The SpatialSonic model, a spatial-aware encoders and azimuth state matrices model utilizing spatial-aware encoders and azimuth state matrices, is introduced utilizing spatial-aware encoders and azimuth state matrices to reveal reasonable spatial guidance.
- 41
- Betrayed by Captions: Joint Caption Grounding and Generation for Open Vocabulary Instance Segmentation40
This work devise a joint Caption Grounding and Generation (CGG) framework, which incorporates a novel grounding loss that only focuses on matching object nouns to improve learning efficiency and introduces a caption generation head that enables additional supervision and contextual modeling as a complementation to the grounding loss.
- NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images: Methods and Results37
This paper reviews the NTIRE 2025 Challenge on Day and Night Raindrop Removal for Dual-Focused Images to establish a new and powerful benchmark for the task of removing raindrops under varying lighting and focus conditions.
- Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation37
Tube-Link is a near-online approach that takes a short subclip as input and outputs the corresponding spatial-temporal tube masks, and introduces temporal contrastive learning to instance-wise discriminative features for tube-level association.
- Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model36
Reason3D is introduced, a novel LLM designed for comprehensive 3D understanding that processes point cloud data and text prompts to produce textual responses and segmentation masks, enabling advanced tasks such as 3D reasoning segmentation, hierarchical searching, express referring, and question answering with detailed mask outputs.
- Pair Then Relation: Pair-Net for Panoptic Scene Graph Generation36
A novel framework is presented: Pair then Relation (Pair-Net), which uses a Pair Proposal Network (PPN) to learn and filter sparse pair-wise relationships between subjects and objects and achieves over 10% absolute gains compared to the baseline, PSGFormer.
- Skeleton-in-Context: Unified Skeleton Sequence Modeling with In-Context Learning35
This work proposes Skeleton-in-Context (SiC), an effective framework for in-context skeleton sequence modeling that is able to handle multiple skeleton-based tasks simultaneously after a single training process and accomplish each task from context according to the given prompt.
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer34
Sail is introduced, a single transformer unified multimodal large language model (MLLM) that integrates raw pixel encoding and language decoding within a singular architecture that eliminates the need for a separate vision encoder, presenting a more minimalist architecture design.
- 34
- 4D Panoptic Scene Graph Generation33
To solve PSG-4D, the introduction of 4D Panoptic Scene Graph (PSG-4D), a new representation that bridges the raw visual data perceived in a dynamic 4D world and high-level visual understanding, and a real-world application example to demonstrate how the model can achieve dynamic scene understanding by integrating a large language model into the system.
- DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries33
This work introduces Dynamic Anchor Queries (DAQ) to shorten the transition gap between the anchor and target queries by dynamically generating anchor queries based on the features of potential candidates and introduces a query-level object Emergence and Disappearance Simulation (EDS) strategy, which unleashes DAQ's potential without any additional cost.
- 33
- 33
- 33
- BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation32
This work introduces a comprehensive dataset, benchmark, and baseline model for video forgery detection, and develops BusterX, an MLLM baseline with RL training that outperforms several leading MLLMs in both detection accuracy and rationale quality.
- HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing32
High-quality, human-rewarded dataset specifically designed for instruction-guided image editing, enabling precise and diverse image manipulations through open-form language instructions, set a new versatile benchmark for instructional image editing datasets.
- DGMamba: Domain Generalization via Generalized State Space Model32
This paper proposes a novel framework for DG, named DGMamba, that excels in strong generalizability toward unseen domains and meanwhile has the advantages of global receptive fields, and efficient linear complexity.
- An Empirical Study of GPT-4o Image Generation Capabilities31
An empirical study of GPT-4o's image generation capabilities is conducted, benchmarking it against leading open-source and commercial models and identifying promising directions for future unified generative models, emphasizing the role of architectural design and data scaling.
- Learning Feature Inversion for Multi-class Anomaly Detection under General-purpose COCO-AD Benchmark29
A simple but more powerful InvAD framework to achieve high-quality feature reconstruction and improves the effectiveness of reconstruction-based methods on popular MVTec AD, VisA, and the newly proposed COCO-AD datasets under a multi-class unsupervised setting.
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World29
This work introduces WorldLens, a full-spectrum benchmark evaluating how well a model builds, understands, and behaves within its generated world, and develops WorldLens-Agent, an evaluation model distilled from these annotations to enable scalable, explainable scoring.
- On Path to Multimodal Generalist: General-Level and General-Bench29
This project introduces General-Level, an evaluation framework that defines 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI.
- Panoptic-PartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation29
The first end-to-end unified framework, Panoptic-PartFormer is designed, designing a meta-architecture that decouples part features and things/stuff features, respectively and proposes a new metric Part-Whole Quality (PWQ), better to measure this task from pixel-region and part-whole perspectives.
- 29
- 28
- Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation28
A unified framework, Reference Twice (RefT), is introduced to exploit the relationship between support and query features for FSIS and related tasks, demonstrating that support object queries encode key factors after base training, allowing query features to be enhanced twice at both feature and query levels.
- 27
- DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers26
A novel task: Layer-Wise Decomposition of Alpha-Composited Images, aiming to recover constituent layers from single overlapped images under the condition of semi-transparent/transparent alpha layer non-linear occlusion, and presents DiffDecompose, a diffusion Transformer-based framework that learns the posterior over possible layer decompositions conditioned on the input image, semantic prompts, and blending type.
- LLAVADI: What Matters For Multimodal Large Language Models Distillation26
Results show that joint alignment for both tokens and logit alignment plays critical roles in teacher-student frameworks and even a 2.7B small-scale model can perform on par with larger models with 7B or 13B parameters.
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding25
The Pixel-SAIL, a single transformer for pixel-wise MLLM tasks, is presented and a novel visual prompt injection strategy to enable the single transformer to understand visual prompt inputs and benefit from the early fusion of visual prompt embeddings and vision tokens is proposed.
- BA-SAM: Scalable Bias-Mode Attention Mask for Segment Anything Model25
A Scalable Bias-Mode Attention Mask (BA-SAM) is proposed to enhance SAM's adaptability to varying image resolutions while eliminating the need for structure modifications, and a new scaling factor is introduced to ensure consistent magnitude in the attention layer's dot product values when the token sequence length changes.
- Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model25
This work designs a mixed backbone that contains convolution and RWKV operation, which achieves the best for both accuracy and efficiency and designs an efficient decoder to utilize the multiscale tokens to obtain high-quality masks.
- DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation23
This work proposes a new task: customized manga generation and introduces DiffSensei, an innovative framework specifically designed for generating manga with dynamic multi-character control, marking a significant advancement in manga generation by enabling text- adaptable character customization.
- Global Aggregation Then Local Distribution for Scene Parsing23
A novel local distribution module is designed which models the affinity map between global and local relationship for each pixel adaptively and can be modularized as an end-to-end trainable block and easily plugged into existing semantic segmentation networks, giving rise to the GALD networks.
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models20
A unified yet straightforward framework that contains a mixed reward function design (Mixed-Reward) and a mixed post-training dataset (Mixed-45K) and a new open-ended reward named Bidirectional Max-Average Similarity (BMAS) by leveraging tokenizer embedding matching between the generated response and the ground truth.
- DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency19
This work proposes a new Dual Consistency SAM (DC-SAM), a prompt-tuning framework that adapts SAM and SAM2 for both image and video in-context segmentation and introduces a novel cycle-consistent cross-attention mechanism to enforce alignment between fused features and visual prompts.
- 19
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding19
Experiments show that AdaVideoRAG significantly improves both efficiency and accuracy on long-video QA tasks and can be seamlessly plugged into existing MLLMs through lightweight APIs, establishing a new paradigm for adaptive retrieval-augmented video analysis.
- 19
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs18
This work introduces Grasp Any Region (GAR) for region-level visual understanding, and constructs GAR-Bench, which not only provides a more accurate evaluation of single-region comprehension, but also, more importantly, measures interactions and complex reasoning across multiple regions.
- Neural Collapse Terminus: A Unified Solution for Class Incremental Learning and Its Variants18
A unified solution to the misalignment dilemma in the three tasks of CIL, LTCIL and FSCIL is offered andoretical analysis indicates that the method holds the neural collapse optimality in an incremental fashion regardless of data imbalance or data scarcity.
- D$^2$GS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction17
A unified framework D$^2$GS is proposed, comprising a Depth-and-Density Guided Dropout strategy that suppresses overfitting by adaptively masking redundant Gaussians based on density and depth, and a Distance-Aware Fidelity Enhancement module that improves reconstruction quality in under-fitted far-field areas through targeted supervision.
- 17
- VG4D: Vision-Language Model Goes 4D Video Recognition17
This work proposes the Vision-Language Models Goes 4D (VG4D) framework, a framework to transfer VLM knowledge from visual-text pretrained models to a 4D point cloud network, and achieves improved recognition performance.
- Improving Video Instance Segmentation via Temporal Pyramid Routing17
A Temporal Pyramid Routing strategy to conditionally align and conduct pixel-level aggregation from a feature pyramid pair of two adjacent frames is proposed, which is a light-weight and plug-and-play module and can be easily applied to existing instance segmentation methods.
- Rethinking Evaluation Metrics of Open-Vocabulary Segmentaion16
Novel evaluation metrics, namely Open mIoU, Open AP, and Open PQ, tailored for three open-vocabulary segmentation tasks are designed, which demonstrate that the relative subjectivity of similarity distance can still well evaluate the open ability of the existing open- vocabulary segmentation methods.
- 15
- Unified Dense Prediction of Video Diffusion14
A unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts and incorporates learnable task embeddings brings multiple dense prediction tasks into a single model, enhancing flexibility and further boosting performance.
- MEDIC: Zero-shot Music Editing with Disentangled Inversion Control14
MEDIC, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion, outperforms state-of-the-art inversion techniques in editing fidelity and content preservation.
- Convolution-Enhanced Evolving Attention Networks14
This work proposes a novel and generic evolving attention mechanism, which directly models the evolution of inter-token relationships through a chain of residual convolutional modules, and is the first work that explicitly models the layer-wise evolution of attention maps.
- A Generalist FaceX via Learning Unified Facial Representation14
This work presents FaceX framework, a novel facial generalist model capable of handling diverse facial tasks simultaneously, and introduces Facial Omni-Representation Decomposing (FORD) for seamless manipulation of various facial components, microscopically decomposing the core aspects of most facial editing tasks.
- RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection13
This work designs a 3D Gaussian Expansion module to mitigate inaccuracies in radar points, including position, Radar Cross-Section (RCS), and velocity, and introduces a weather-adaptive fusion module, which adaptively fuses radar and camera features based on camera signal confidence.
- From Multimodal LLM to Human-level AI: Modality, Instruction, Reasoning and Beyond13
This tutorial aims to deliver a comprehensive review of cutting-edge research in MLLMs, focusing on three key areas: MLLM architecture design, instructional learning, and multimodal reasoning of MLLMs.
- 13
- SAMTok: Representing Any Mask with Two Words12
A discrete mask tokenizer that converts any region mask into two special tokens and reconstructs the mask using these tokens with high fidelity, allowing base MLLMs to learn pixel-wise capabilities through standard next-token prediction and simple reinforcement learning, without architectural modifications and specialized loss design.
- 12
- RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything12
This work explores a new real-time segmentation setting, named all-purpose segmentation in real-time, to transfer VFMs in real-time deployment and presents Real-Time All Purpose SAM (RAP-SAM), which contains an efficient encoder and an efficient decoupled decoder to perform prompt-driven decoding.
- Open-Vocabulary SAM3D: Towards Training-Free Open-Vocabulary 3D Scene Understanding11
This paper introduces OV-SAM3D, a training-free method that contains a universal framework for understanding open-vocabulary 3D scenes without requiring prior knowledge of the scene, and surpasses existing open-vocabulary methods in unknown open-world environments.
- DreamRelation: Bridging Customization and Relation Generation11
This work introduces DreamRelation, a framework that disentangles identity and relation learning using a carefully curated dataset and proposes two key modules to tackle the two main challenges of generating accurate and natural relationships, especially when significant pose adjustments are required, and avoiding object confusion in cases of overlap.
- Exploring Self-Supervised Learning for Multi-Modal Remote Sensing Pre-Training via Asymmetric Attention Fusion11
The proposed Asymmetric Attention Fusion framework achieves an improvement of over 7% in all metrics compared to randomly initialized methods for both tasks, and when compared to early fusion and late fusion methods, AAF consistently outperforms in achieving superior improvements.
- Fast and Accurate Scene Parsing via Bi-Direction Alignment Networks11
An effective method for fast and accurate scene parsing called Bidirectional Alignment Network (BiAlignNet) is proposed by aligning two-path information into each other through a learned flow field to avoid the noise and semantic gaps.
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark10
The Visual Reasoning Tracer task is introduced, which requires models to not only localize the target object but also explicitly predict the intermediate objects that form the reasoning path, and reveals that while existing models often produce the correct final output, they struggle to ground their intermediate reasoning.
- 10
- Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene9
A 4D Large Language Model (4D-LLM) integrated with a 3D mask decoder for end-to-end generation of 4D-PSG and a 2D-to-4D visual scene transfer learning framework, where a spatial-temporal scene transcending strategy effectively transfers dimension-invariant features from abundant 2D SG annotations to 4D scenes, effectively compensating for data scarcity in 4D-PSG.
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation9
This work proposes a parallel multimodal diffusion framework, MMaDA-Parallel, that enables continuous, bidirectional interaction between text and images throughout the entire denoising trajectory, establishing a more robust paradigm for thinking-aware image synthesis.
- 9
- You Can't Ignore Either: Unifying Structure and Feature Denoising for Robust Graph Learning9
This paper develops a unified graph denoising (UGD) framework to unravel the deadlock between structure and feature denoising, and proposes to refine noisy features with reconstruction based on a graph auto-encoder.
- Iterative Robust Visual Grounding with Masked Reference based Centerpoint Supervision9
An Iterative Robust Visual Grounding (IRVG) framework with Masked Reference based Centerpoint Supervision (MRCS) is proposed with iterative multi-level vision-language fusion (IMVF) for better alignment and a multi-stage false-alarm sensitive decoder (MFSD) is presented to prevent the generation of false-alarm objects when presented with inaccurate expressions.
- Dense360: Dense Understanding from Omnidirectional Panoramas8
This work introduces an omnidirectional panoramas dataset featuring a comprehensive suite of reliability-scored annotations and introduces Dense360-Bench, the first benchmark for evaluating MLLMs on omnidirectional captioning and grounding, establishing a comprehensive framework for advancing dense visual-language understanding in panoramic settings.
- 8
- Continuous sPatial‐temporal deformable image registration and 4D frame interpolation8
This work presents a new approach to DIR implementation that addresses the issue of uncertainty in discrete volumetric motion representation, which affects the reliability of subsequent contour propagation and dose accumulation procedures.
- 8
- Point-In-Context: Understanding Point Cloud via In-Context Learning7
This paper introduces Point-In-Context (PIC), a pioneering framework for 3D point cloud understanding that leverages in-context learning with a standard transformer architecture that uniquely enables the execution of multiple tasks after a single, unified training phase, eliminating the need for fine-tuning.
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA7
This paper proposes Segmentation Augmented and Selective Averaged Sa2VA (SaSaSa2VA) to address two key bottlenecks that limit segmentation performance: sparse frame sampling and reliance on a single [SEG] token for an entire video.
- From Masks to Worlds: A Hitchhiker's Guide to World Models7
This book follows one clear road: from early masked models that unified representation learning across modalities, to unified architectures that share a single paradigm, then to interactive generative models that close the action-perception loop, and finally to memory-augmented systems that sustain consistent worlds over time.
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning7
This work proposes Reason50K, a large-scale dataset specifically curated for training and evaluating hypothetical instruction reasoning image editing, along with ReasonBrain, a novel framework designed to reason over and execute implicit hypothetical instructions across diverse scenarios.
- CyberV: Cybernetics for Test-time Scaling in Video Understanding7
This work proposes a novel framework inspired by cybernetic principles, redesigning video MLLMs as adaptive systems capable of self-monitoring, self-correction, and dynamic resource allocation during inference, and demonstrates consistent gains on general-purpose benchmarks, such as VideoMME and WorldSense.
- 6
- DynamicControl: Adaptive Condition Selection for Improved Text-to-Image Generation6
This work proposes a novel framework, DynamicControl, which supports dynamic combinations of diverse control signals, allowing adaptive selection of different numbers and types of conditions, and demonstrates its superiority over existing methods in terms of controllability, generation quality and composability under various conditional controls.
- Is Your Driving World Model an All-Around Player?5
WorldLens is introduced, a unified benchmark that measures world-model fidelity across the full spectrum, from pixel quality and 4D geometry to closed-loop driving and human perceptual alignment, through five complementary aspects and 24 standardized dimensions and forms a unified ecosystem for assessing generated worlds not merely by visual appeal, but by physical and behavioral fidelity.
- 5
- RecTok: Reconstruction Distillation along Rectified Flow5
This work proposes RecTok, which overcomes the limitations of high-dimensional visual tokenizers through two key innovations: flow semantic distillation and reconstruction--alignment distillation, to make the forward flow in flow matching semantically rich, rather than focusing on the latent space as in previous works.
- UMC: Unified Resilient Controller for Legged Robots with Joint Malfunctions4
A novel, model-free, two-stage training framework, Unified Malfunction Controller (UMC), incorporating a masking mechanism to enhance damage resilience, is proposed, incorporating a masking mechanism to enhance damage resilience.
- Auto Cherry-Picker : Learning from High-quality Generative Data Driven by Language4
This work presents Auto Cherry-Picker (ACP), a novel framework that generates high-quality cross-modality training samples at scale to augment perception and multi-modal training and finds a positive correlation between CLIS and performance gains in downstream tasks.
- PointDGMamba: Domain Generalization of Point Cloud Classification via Generalized State Space Model4
This paper presents the first work that studies the generalizability of state space models (SSMs) in DG PCC and finds that directly applying SSMs into DG PCC will encounter several challenges.
- Masked Generative Transformer Is What You Need for Image Editing3
This work presents EditMGT, an MGT-based editing framework that is the first of its kind and achieves state-of-the-art image similarity on multiple benchmarks while delivering 6x faster editing, demonstrating that MGTs offer a compelling alternative to diffusion-based editing.
- MelodyEdit: Zero-shot Music Editing with Disentangled Inversion Control3
MelodyEdit is proposed, a novel zero-shot music editing system based on innovative Disentangled Inversion Control (DIC) technique, which comprises Harmonized Attention Control and Disentangled Inversion, which outperforms state-of-the-art inversion techniques in editing fidelity and content preservation.
- 3
- 3
- CPT-Interp: Continuous sPatial and Temporal Motion Modeling for 4D Medical Image Interpolation3
This study draws inspiration from fluid mechanics to propose a novel approach for continuously modeling patient anatomic motion using implicit neural representation that ensures both spatial and temporal continuity, effectively bridging Eulerian and Lagrangian specifications together to naturally facilitate continuous frame interpolation.
- Flow2Seg: Motion-Aided Semantic Segmentation3
This paper leverages motion information densely represented by optical flow to assist the semantic segmentation task and finds that optical flow improves image-based segmentation on object boundaries especially on small thin objects.
- 4th PVUW MeViS 3rd Place Report: Sa2VA2
This report shows that with a simple modification to the test time inference method on stronger MLLMs, it can lead to stronger results on MeVIS, and adopts the recent method Sa2VA, a unified model for dense grounded understanding of both images and videos.
- PointDGRWKV: Generalizing RWKV-like Architecture to Unseen Domains for Point Cloud Classification2
This paper proposes PointDGRWKV, the first RWKV-based framework tailored for DG PCC, which introduces two key modules to enhance spatial modeling and cross-domain robustness, while maintaining RWKV's linear efficiency.
- 2
- MIMAFace: Face Animation via Motion-Identity Modulated Appearance Feature Learning2
This work meticulously examines the essential appearance features in the facial animation tasks, and introduces a Motion-Identity Modulated Appearance Learning Module (MIA) that modulates CLIP features at both motion and identity levels, and designs an Inter-clip Affinity Learning Module (ICA) to model temporal relationships across clips.
- Fine-Grained Multimodal Alignment for Image-Text Retrieval via Graph Learning1
This work proposes a Graph-based Fine-Grained multimodal Alignment (GFGA) method, which adopts a concept-based fusion module in stage I to complement typical fragment-based fusion, which lacks comprehensive interactions.
- Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking1
The interleaved Analyzing-Drafting problem-solving loop (AD-Loop), a new think paradigm that dynamically alternates between analytic and drafting operations, is introduced, highlighting AD-Loop as a principled and broadly applicable strategy for synergizing comprehension and creation.
- 1
- Bridge Feature Matching and Cross-Modal Alignment with Mutual-Filtering for Zero-Shot Anomaly Detection1
FiSeCLIP for ZSAD with training-free CLIP is introduced, combining the feature matching with the cross-modal alignment, and CLIP's inherent potential to restore its local semantic correlation is explored, adapting it for fine-grained anomaly detection tasks to enable a more accurate filtering process.
- 1
- 1
- Effective Adapter for Face Recognition in the Wild1
This paper proposes an effective adapter for augmenting existing face recognition models trained on high-quality facial datasets using two similar structures, one fixed and the other trainable, to process both the unrefined and enhanced images using two similar structures.
- Dynamic Dual Sampling Module For Fine-Grained Semantic Segmentation1
A Dynamic Dual Sampling Module (DDSM) is proposed to conduct dynamic affinity modeling and propagate semantic context to local details, which yields a more discriminative representation.
- –
- UniVR: Thinking in Visual Space for Unified Visual Reasoning–
UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations, is introduced, the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol.
- From bits to atoms: A survey of cross-layer safety in Embodied AI–
This survey organizes embodied AI safety through a four-layer reference architecture and advocates for unifying software alignment with rigorous hardware constraints to build trustworthy systems capable of operating safely within human-centric environments.
- –
- BACON: Bayesian Optimal Condensation Framework for Dataset Distillation–
The BAyesian optimal CONdensation framework (BACON) is proposed, which is the first work to introduce the Bayesian theoretical framework to the literature of DD and provides theoretical support for enhancing the performance of DD.
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.