Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks101 from public data
- A Survey on Vision Transformer4,430
This paper reviews these vision transformer models by categorizing them in different tasks and analyzing their advantages and disadvantages, and takes a brief look at the self-attention mechanism in computer vision, as it is the base component in transformer.
- Recurrent Feature Reasoning for Image Inpainting490
A Recurrent Feature Reasoning (RFR) network which is mainly constructed by a plug-and-play Recurrent feature Reasoning module and a Knowledge Consistent Attention (KCA) module, which recurrently infers the hole boundaries of the convolutional feature maps and uses them as clues for further inference.
- Perceptual Adversarial Networks for Image-to-Image Transformation396
The perceptual adversarial loss is proposed, which undergoes an adversarial training process between the image transformation network and the discriminative network and can be trained alternately to solve image-to-image transformation tasks.
- Evolutionary Generative Adversarial Networks371
A novel GAN framework called evolutionary GANs (E-GANs) is proposed for stable GAN training and improved generative performance, which overcomes the limitations of an individual adversarial training objective and always preserves the well-performing offspring, contributing to progress in, and the success of GAns.
- 369
- R1-VL: Learning to Reason with Multimodal Large Language Models via Step-Wise Group Relative Policy Optimization363
The proposed StepGRPO is designed, a new online reinforcement learning framework that enables MLLMs to self-improve reasoning ability via simple, effective and dense step-wise rewarding, and introduces R1-VL, a series of MLLMs with outstanding capabilities in step-by-step reasoning.
- Tensorized Bipartite Graph Learning for Multi-View Clustering338
A variance-based de-correlation anchor selection strategy for bipartite construction which is superior to the state-of-the-art methods and solves TBGL by an efficient algorithm which is time-economical and has good convergence.
- Patch Slimming for Efficient Vision Transformers218
A novel patch slimming approach that discards useless patches in a topdown paradigm is presented that can significantly reduce the computational costs of vision transformers without affecting their performances.
- 214
- LIIR: Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning200
This paper proposes to merge the two directions of MARL and learn each agent an intrinsic reward function which diversely stimulates the agents at each time step, and compares LIIR with a number of state-of-the-art MARL methods on battle games in StarCraft II.
- 188
- Transductive Face Sketch-Photo Synthesis170
A novel transductive face sketch-photo synthesis method that incorporates the given test samples into the learning process and optimizes the performance on these test samples and efficiently optimizes this probabilistic model by alternating optimization.
- 156
- Universal Blind Image Quality Assessment Metrics Via Natural Scene Statistics and Multiple Kernel Learning152
Two universal blind quality assessment models are presented, NSS global scheme and NSS two-step scheme, which are in remarkably high consistency with the human perception, and overwhelm representative universal blind algorithms as well as some standard full reference quality indexes.
- Manifold Regularized Dynamic Network Pruning114
A new paradigm that dynamically removes redundant filters by embedding the manifold information of all instances into the space of pruned networks (dubbed as ManiDP) is proposed, which shows better performance in terms of both accuracy and computational cost compared to the state-of-the-art methods.
- A Relay Level Set Method for Automatic Image Segmentation114
A new image segmentation method that applies an edge-based level set method in a relay fashion to detect all boundaries and automatically obtains a full segmentation without specifying the number of relays in advance is presented.
- 108
- BAG: Bi-directional Attention Entity Graph Convolutional Network for Multi-hop Reasoning Question Answering86
A Bi-directional Attention Entity Graph Convolutional Network (BAG) is proposed, leveraging relationships between nodes in an entity graph and attention information between a query and the entity graph to solve the multi-hop reasoning question answering task.
- Image-Question-Answer Synergistic Network for Visual Dialog83
A novel image-question-answer synergistic network to value the role of the answer for precise visual dialog and boosts the discriminative visual dialog model to achieve a new state-of-the-art of 57.88% normalized discounted cumulative gain.
- Bilinear Graph Networks for Visual Question Answering77
This paper revisits the bilinear attention networks in the visual question answering task from a graph perspective and develops bilInear graph networks to model the context of the joint embeddings of words and objects.
- 77
- 70
- Active Learning for Crowdsourcing Using Knowledge Transfer69
This paper proposes a new probabilistic model that transfers knowledge from abundant unlabeled data in auxiliary domains to help estimate labelers' expertise and presents a novel active learning algorithm that simultaneously selects the most informative example and queries its label from the labeler with the best expertise.
- WebUAV-3M: A Benchmark for Unveiling the Power of Million-Scale Deep UAV Tracking67
A fine-grained UAV tracking-under-scenario constraint (UTUSC) evaluation protocol and seven challenging scenario subtest sets are constructed to enable the community to develop, adapt and evaluate various types of advanced trackers.
- Hierarchical Prototype Networks for Continual Graph Representation Learning64
HPNs are presented which extract different levels of abstract knowledge in the form of prototypes to represent the continuously expanded graphs and it is proved that under mild constraints, learning new tasks will not alter the prototypes matched to previous data, thereby eliminating the forgetting problem.
- Tag Disentangled Generative Adversarial Network for Object Image Re-rendering63
A principled Tag Disentangled Generative Adversarial Networks (TD-GAN) for re-rendering new images for the object of interest from a single image of it by specifying multiple scene properties by specifyingmultiple scene properties.
- Parallelized Evolutionary Learning for Detection of Biclusters in Gene Expression Data59
This paper proposes a new biclustering algorithm based on evolutionary learning that demonstrates a significant improvement in discovering additive biclusters and is able to discover bicluster seeds within a limited computing time.
- 51
- Exposure Trajectory Recovery From Motion Blur50
Exposure trajectories are defined, which represent the motion information contained in a blurry image and explain the causes of motion blur, and a novel motion offset estimation framework is proposed to model pixel-wise displacements of the latent sharp image at multiple timepoints.
- Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural Networks48
Two dedicated forms of meta neighborhood aggregators are proposed, an exclusive meta aggregator termed as Greedy Gumbel Neighborhood Aggregator (GNA), and a diffused meta aggregators termed as Adaptable Hybrid Neighborhood AggRegator (ANA), which learn to exclusively pick one single optimal aggregator from a pool of candidates.
- 47
- Dynamic Contrastive Distillation for Image-Text Retrieval45
A novel plug-in dynamic contrastive distillation (DCD) framework to compress the large VLP models for the ITR task and proposes dynamic distillation to dynamically learn samples of different difficulties to balance better the difficulty of knowledge and students' self-learning ability.
- Heatmap Regression via Randomized Rounding45
The proposed quantization system induced by the randomized rounding operation encodes the fractional part of numerical coordinates into the ground truth heatmap using a probabilistic approach during training and decodes the predicted numerical coordinates from a set of activation points during testing.
- Exploiting Local Coherent Patterns for Unsupervised Feature Ranking45
This paper aims to develop an unsupervised feature ranking algorithm that evaluates features using discovered local coherent patterns, known as biclusters, that can yield comparable or even better performance in comparison with the well-known Fisher score, Laplacian score, and variance score using three UCI data sets.
- 38
- 36
- Active Multi-task Learning via Bandits36
This paper introduces a new active multi- task learning paradigm, which selectively samples effective instances for multi-task learning and cast the selection procedure as a bandit framework, and provides an implementation of the algorithm based on a popular multi- Task learning algorithm that is trace-norm regularization method.
- Merging Models on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging35
This study proposes a training-free projection-based continual merging method that processes models sequentially through orthogonal projections of weight matrices and adaptive scaling mechanisms, enabling efficient sequential integration of task-specific knowledge.
- SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object Detection35
A Simulated multi-modal Distillation (SimDistill) method that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space and can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment.
- Continual Learning on Graphs: Challenges, Solutions, and Opportunities34
A comprehensive review of existing continual graph learning (CGL) algorithms is provided by elucidating the different task settings and categorizing the existing methods based on their characteristics, comparing the CGL methods with traditional continual learning techniques and analyzing the applicability of the traditional continual learning techniques to CGL tasks.
- Benchmarking Reasoning Robustness in Large Language Models30
A novel benchmark is introduced, termed as Math-RoB, that exploits hallucinations triggered by missing information to expose reasoning gaps, achieved by an instruction-based approach to generate diverse datasets that closely resemble training distributions, facilitating a holistic robustness assessment and advancing the development of more robust reasoning frameworks.
- Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging29
An efficient-and-effective WEMoE (E-WEMoE) method, whose core mechanism involves eliminating non-essential elements in the critical modules of WEMoE and implementing shared routing across multiple MoE modules, thereby significantly reducing both the trainable parameters, the overall parameter count, and computational overhead of the merged model by WEMoE.
- Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam Search28
A novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking, which is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm.
- Variance Reduced Methods for Non-convex Composition Optimization27
To significantly improve the query complexity of current approaches, the stochastic composition via variance reduction (SCVR) is devised and an extension to handle the mini-batch cases is proposed, which improve thequery complexity under the optimalmini-batch size.
- A novel method for speed training acceleration of recurrent neural networks27
A particular approach for the Jordan network will be shown, however, the presented idea is applicable to other RNN structures and can be implemented in digital hardware.
- 26
- Quantum geometric machine learning for quantum circuits and control23
This work demonstrates how geometric control techniques can be used to both verify the extent to which geometrically synthesised quantum circuits lie along geodesic, and thus time-optimal, routes and synthesise those circuits.
- 21
- Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion Models20
A multi-step denoising scheduled sampling (MDSS) strategy to alleviate the exposure bias in DDPMs and demonstrates that this approach is more effective in mitigating exposure bias in DDPM, DDIM, and DPM-solver.
- SIR: Self-Supervised Image Rectification via Seeing the Same Scene From Multiple Different Lenses20
A novel self-supervised image rectification (SIR) method based on an important insight that the rectified results of distorted images of a same scene from different lenses should be the same, with comparable or even better performance than the supervised baseline method and representative state-of-the-art (SOTA) methods.
- JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models19
JustLogic is a synthetically generated deductive reasoning benchmark designed for rigorous evaluation of LLMs and reveals that state-of-the-art (SOTA) reasoning LLMs perform on par or better than the human average but significantly worse than the human ceiling, and SOTA non-reasoning models still underperform the human average.
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI19
This work introduces MMReason, a new benchmark designed to precisely and comprehensively evaluate MLLM long-chain reasoning capability with diverse, open-ended, challenging questions and designs a reference-based ternary scoring mechanism to reliably assess intermediate reasoning steps.
- Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving19
A natural reflection-based backdoor attack targeting VLM systems in autonomous driving scenarios, aiming to induce substantial response delays when specific visual triggers are present, uncovering a new class of attacks that exploit the stringent real-time requirements of autonomous driving.
- Semantic-Preserving Adversarial Text Attacks19
A Bigram and Unigram-based adaptive adaptive Semantic Preservation Optimization (BU-SPO) approach which attacks text documents not only at a unigram word level but also at a bigram level to avoid generating meaningless sentences is devised.
- Fine-Grained Zero-Shot Learning: Advances, Challenges, and Prospects17
A broad review of recent advances for fine-grained analysis in ZSL is presented and a taxonomy of existing methods and techniques with a thorough analysis of each category is provided, covering publicly available datasets, models, implementations, and some more details as a library.
- 16
- Multi-label Active Learning Based on Maximum Correntropy Criterion: Towards Robust and Discriminative Labeling16
A novel active learning approach to reduce the annotation costs greatly for multi-label classification by merging uncertainty and representativeness by using the MCC to alleviate the influence of outlier labels for discriminative labeling.
- 13
- Good Questions Help Zero-Shot Image Reasoning12
Question-Driven Visual Exploration (QVix) is introduced, a novel prompting strategy that enhances the exploratory capabilities of LVLMs in zero-shot reasoning tasks and significantly outperforms existing methods, highlighting its effectiveness in bridging the gap between complex visual data and LVL Ms' exploratory abilities.
- Graph Reasoning Networks for Visual Question Answering.12
The graph reasoning networks are developed, and the resulting model can reason the relationship and dependence between objects, which leads to realization of multi-step reasoning.
- 12
- Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning11
KG-RAR is proposed, a framework centered on process-oriented knowledge graph construction, a hierarchical retrieval strategy, and a universal post-retrieval processing and reward model (PRP-RM) that refines retrieved information and evaluates each reasoning step.
- Safety Reasoning with Guidelines9
Extensive experiments show that the pro-pose training model to perform safety reasoning for each query synthesize reasoning supervision based on pre-guidelines, training the model to reason in alignment with them, thereby effectively eliciting and utilizing latent knowledge from diverse perspectives.
- EMOv2: Pushing 5M Vision Model Frontier9
This work rethinks the lightweight infrastructure of efficient IRB and practical components in Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMBlock) for lightweight model design and deduce a modern Improved Inverted Residual Mobile Block (i2RMB).
- Learning Versatile Convolution Filters for Efficient Visual Recognition9
Experimental results on benchmark datasets and neural networks demonstrate that the versatile filters introduced are able to achieve comparable accuracy as that of original filters, but require less memory and computation cost.
- Deep Representation Calibrated Bayesian Neural Network for Semantically Explainable Face Inpainting and Editing9
This work proposes a novel deep representation calibrated Bayesian neural network (DRCBNN) for semantically explainable face inpainting and editing and exploits deep representation into Bayesian decision theory and derive a deep representation calibration evidence lower bound (ELBO).
- 9
- Rethink Sparse Signals for Pose-Guided Text-to-Image Generation8
A novel Spatial-Pose ControlNet (SP-Ctrl) is proposed, equipping sparse signals with robust controllability for pose-guided image generation and introduces keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment.
- 8
- On the Rates of Convergence From Surrogate Risk Minimizers to the Bayes Optimal Classifier7
This article introduces the notions of consistency intensity and conductivity to characterize a surrogate loss function and exploits this notion to obtain the rate of convergence from an empirical surrogate risk minimizer to the Bayes optimal classifier, enabling fair comparisons of the excess risks of different surrogaterisk minimizers.
- 6
- On Gleaning Knowledge from Multiple Domains for Active Learning6
A framework that attempts to glean knowledge from multiple domains for active learning by querying the most uncertain and representative samples from the target domain and calculating the important weights for re-weighting the source data in a single unified formulation is proposed.
- Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation5
AmbigST is introduced, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy and achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity.
- Unlikelihood Tuning on Negative Samples Amazingly Improves Zero-Shot Translation5
The authors' analyses suggest that although they work well in ideal ON settings, language IDs become fragile and lose their navigation ability when faced with off-target tokens, which commonly exist during inference but are rare in training scenarios.
- Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion4
The Federated Multimodal Knowledge Graph Completion task, aiming at training over federated MKGs for better predicting the missing links in clients without sharing sensitive knowledge, is proposed, and a framework named MMFeD3-HidE for addressing multimodal uncertain unavailability and multimodal client heterogeneity challenges of FedMKGC is proposed.
- Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs4
This work explores the potential of reusing diverse pre-tuned LoRAs without accessing their original training data, to achieve tuning-free few-shot adaptation in VFMs and distills a meta-LoRA from diverse pre-tuned LoRAs with a meta-learning objective, using surrogate data generated inversely from pre-tuned LoRAs themselves.
- 4
- SeWA: Selective Weight Average via Probabilistic Masking3
This paper proposes a simple yet efficient algorithm called Selective Weight Averaging (SeWA), which adaptively selects checkpoints during the final stages of training for averaging, and theoretically derive the SeWA's stability-based generalization bounds, which are sharper than that of SGD under both convex and non-convex assumptions.
- 3
- 3
- Self-supervised Exposure Trajectory Recovery for Dynamic Blur Estimation.3
Under mild constraints, the learned motion offsets can recover dense, (non-)linear exposure trajectories, which significantly reduce temporal disorder and ill-posed problems and further contribute to motion-aware image deblurring and warping-based video extraction from a single blurry image.
- Model Hemorrhage and the Robustness Limits of Large Language Models2
This work establishes foundational metrics for evaluating model stability during adaptation, providing practical guidelines for maintaining performance while enabling efficient LLM deployment and advancing understanding of neural network resilience under architectural transformations.
- 2
- 1
- 1
- 1
- Error Bounds for Real Function Classes Based on Discretized Vapnik-Chervonenkis Dimensions.1
This paper proposes the discretized VC dimension obtained by discretizing the range of a real function class, and points out that Sauer\'s Lemma is valid for this dimension.
- –
- –
- –
- –
- –
- –
- –
- Responsible Active Learning via Human-in-the-loop Peer Study–
This work introduces a human-in-the-loop teacher-student architecture to isolate unlabelled data from the task learner on the cloud-side by maintaining an active learner (student) on the client-side and devise a discrepancy-based active sampling criterion, Peer Study Feedback, that exploits the variability of peer students to select the most informative data to improve model stability.
- –
- –
- –
- Batch Mode Active Learning for Geographical Image Classification–
An innovative batch mode active learning by combining discriminative and representative information for hyperspectral image classification with support vector machine is proposed and a novel form of upper bound for true risk in the active learning setting is derived.
- Patch Alignment for Graph Embedding–
This paper presents locally linear embedding (LLE), which uses linear coefficients, which reconstruct a given example by its neighbors, to represent the local geometry, and then seeks a low-dimensional embedding, in which these coefficients are still suitable for reconstruction.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.