Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works101 from public data

TitleCited by
  • A Survey on Vision Transformer

    Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jian-Yuan Guo, Zhenhua Liu, Ye-Hui Tang, An Xiao, +5 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    This paper reviews these vision transformer models by categorizing them in different tasks and analyzing their advantages and disadvantages, and takes a brief look at the self-attention mechanism in computer vision, as it is the base component in transformer.

    4,430
  • Recurrent Feature Reasoning for Image Inpainting

    Jingyuan Li, N. Wang, Le-Fei Zhang, Bo Du, D. Tao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2020

    A Recurrent Feature Reasoning (RFR) network which is mainly constructed by a plug-and-play Recurrent feature Reasoning module and a Knowledge Consistent Attention (KCA) module, which recurrently infers the hole boundaries of the convolutional feature maps and uses them as clues for further inference.

    490
  • Perceptual Adversarial Networks for Image-to-Image Transformation

    Chaoyue Wang, Chang Xu, Chaohui Wang, D. Tao

    IEEE Transactions on Image Processing · 2018

    The perceptual adversarial loss is proposed, which undergoes an adversarial training process between the image transformation network and the discriminative network and can be trained alternately to solve image-to-image transformation tasks.

    396
  • Evolutionary Generative Adversarial Networks

    Chaoyue Wang, Chang Xu, Xin Yao, D. Tao

    IEEE Transactions on Evolutionary Computation · 2019

    A novel GAN framework called evolutionary GANs (E-GANs) is proposed for stable GAN training and improved generative performance, which overcomes the limitations of an individual adversarial training objective and always preserves the well-performing offspring, contributing to progress in, and the success of GAns.

    371
  • Multimodal Graph-Based Reranking for Web Image Search

    Meng Wang, Hao Li, D. Tao, K. Lu, Xindong Wu

    IEEE Transactions on Image Processing · 2012

    369
  • R1-VL: Learning to Reason with Multimodal Large Language Models via Step-Wise Group Relative Policy Optimization

    Jingyi Zhang, Jia-Xing Huang, Huan-Jin Yao, Shunyu Liu, Xi-Kun Zhang, Shijian Lu, Dacheng Tao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    The proposed StepGRPO is designed, a new online reinforcement learning framework that enables MLLMs to self-improve reasoning ability via simple, effective and dense step-wise rewarding, and introduces R1-VL, a series of MLLMs with outstanding capabilities in step-by-step reasoning.

    363
  • Tensorized Bipartite Graph Learning for Multi-View Clustering

    Wei Xia, Quan-Xue Gao, Qianqian Wang, Xin-Bo Gao, C. Ding, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    A variance-based de-correlation anchor selection strategy for bipartite construction which is superior to the state-of-the-art methods and solves TBGL by an efficient algorithm which is time-economical and has good convergence.

    338
  • Patch Slimming for Efficient Vision Transformers

    Ye-Hui Tang, Kai Han, Yunhe Wang, Chang Xu, Jian-Yuan Guo, Chao Xu, D. Tao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    A novel patch slimming approach that discards useless patches in a topdown paradigm is presented that can significantly reduce the computational costs of vision transformers without affecting their performances.

    218
  • 214
  • LIIR: Learning Individual Intrinsic Reward in Multi-Agent Reinforcement Learning

    Ya-Li Du, Lei Han, Meng Fang, Ji Liu, Tianhong Dai, D. Tao

    neural information processing systems · 2019

    This paper proposes to merge the two directions of MARL and learn each agent an intrinsic reward function which diversely stimulates the agents at each time step, and compares LIIR with a number of state-of-the-art MARL methods on battle games in StarCraft II.

    200
  • Progressive Reconstruction of Visual Structure for Image Inpainting

    Jingyuan Li, Fengxiang He, Le-Fei Zhang, Bo Du, D. Tao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2019

    188
  • Transductive Face Sketch-Photo Synthesis

    N. Wang, D. Tao, Xin-Bo Gao, Xue-Long Li, Jie Li

    IEEE Transactions on Neural Networks and Learning Systems · 2013

    A novel transductive face sketch-photo synthesis method that incorporates the given test samples into the learning process and optimizes the performance on these test samples and efficiently optimizes this probabilistic model by alternating optimization.

    170
  • Saliency propagation from simple to difficult

    Chen Gong, D. Tao, W. Liu, S. Maybank, Meng Fang, Keren Fu, Jie Yang

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2015

    156
  • Universal Blind Image Quality Assessment Metrics Via Natural Scene Statistics and Multiple Kernel Learning

    Xin-Bo Gao, Fei Gao, D. Tao, Xue-Long Li

    IEEE Transactions on Neural Networks and Learning Systems · 2013

    Two universal blind quality assessment models are presented, NSS global scheme and NSS two-step scheme, which are in remarkably high consistency with the human perception, and overwhelm representative universal blind algorithms as well as some standard full reference quality indexes.

    152
  • Manifold Regularized Dynamic Network Pruning

    Ye-Hui Tang, Yunhe Wang, Yixing Xu, Yiping Deng, Chao Xu, D. Tao, Chang Xu

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    A new paradigm that dynamically removes redundant filters by embedding the manifold information of all instances into the space of pruned networks (dubbed as ManiDP) is proposed, which shows better performance in terms of both accuracy and computational cost compared to the state-of-the-art methods.

    114
  • A Relay Level Set Method for Automatic Image Segmentation

    Xin-Bo Gao, Bin Wang, D. Tao, Xue-Long Li

    IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics) · 2010

    A new image segmentation method that applies an edge-based level set method in a relay fashion to detect all boundaries and automatically obtains a full segmentation without specifying the number of relays in advance is presented.

    114
  • Amalgamating Knowledge from Heterogeneous Graph Neural Networks

    Yongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song, D. Tao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    108
  • BAG: Bi-directional Attention Entity Graph Convolutional Network for Multi-hop Reasoning Question Answering

    Yu Cao, Meng Fang, D. Tao

    Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT) · 2019

    A Bi-directional Attention Entity Graph Convolutional Network (BAG) is proposed, leveraging relationships between nodes in an entity graph and attention information between a query and the entity graph to solve the multi-hop reasoning question answering task.

    86
  • Image-Question-Answer Synergistic Network for Visual Dialog

    Dalu Guo, Chang Xu, D. Tao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2019

    A novel image-question-answer synergistic network to value the role of the answer for precise visual dialog and boosts the discriminative visual dialog model to achieve a new state-of-the-art of 57.88% normalized discounted cumulative gain.

    83
  • Bilinear Graph Networks for Visual Question Answering

    Dalu Guo, Chang Xu, D. Tao

    IEEE Transactions on Neural Networks and Learning Systems · 2021

    This paper revisits the bilinear attention networks in the visual question answering task from a graph perspective and develops bilInear graph networks to model the context of the joint embeddings of words and objects.

    77
  • Cross-Domain Human Action Recognition

    IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics) · 2011

    77
  • Adaptive Data Structure Regularized Multiclass Discriminative Feature Selection

    Mingyu Fan, Xiaoqin Zhang, Jie Hu, N. Gu, D. Tao

    IEEE Transactions on Neural Networks and Learning Systems · 2021

    70
  • Active Learning for Crowdsourcing Using Knowledge Transfer

    Meng Fang, Jie Yin, D. Tao

    Proceedings of the AAAI Conference on Artificial Intelligence · 2014

    This paper proposes a new probabilistic model that transfers knowledge from abundant unlabeled data in auxiliary domains to help estimate labelers' expertise and presents a novel active learning algorithm that simultaneously selects the most informative example and queries its label from the labeler with the best expertise.

    69
  • WebUAV-3M: A Benchmark for Unveiling the Power of Million-Scale Deep UAV Tracking

    Chun-Hui Zhang, Guanjie Huang, Li Liu, Shan Huang, Yi-Nan Yang, Xiang Wan, Shiming Ge, Da-Cheng Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    A fine-grained UAV tracking-under-scenario constraint (UTUSC) evaluation protocol and seven challenging scenario subtest sets are constructed to enable the community to develop, adapt and evaluate various types of advanced trackers.

    67
  • Hierarchical Prototype Networks for Continual Graph Representation Learning

    Xi-Kun Zhang, Dong-Jin Song, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    HPNs are presented which extract different levels of abstract knowledge in the form of prototypes to represent the continuously expanded graphs and it is proved that under mild constraints, learning new tasks will not alter the prototypes matched to previous data, thereby eliminating the forgetting problem.

    64
  • Tag Disentangled Generative Adversarial Network for Object Image Re-rendering

    Chaoyue Wang, Chaohui Wang, Chang Xu, D. Tao

    International Joint Conference on Artificial Intelligence · 2017

    A principled Tag Disentangled Generative Adversarial Networks (TD-GAN) for re-rendering new images for the object of interest from a single image of it by specifying multiple scene properties by specifyingmultiple scene properties.

    63
  • Parallelized Evolutionary Learning for Detection of Biclusters in Gene Expression Data

    Qinghua Huang, D. Tao, Xuelong Li, Alan Wee-Chung Liew

    IEEE Transactions on Computational Biology and Bioinformatics · 2011

    This paper proposes a new biclustering algorithm based on evolutionary learning that demonstrates a significant improvement in discovering additive biclusters and is able to discover bicluster seeds within a limited computing time.

    59
  • Report on the FG 2015 Video Person Recognition Evaluation

    J. Beveridge, Hao Zhang, B. Draper, P. Flynn, Zhen-Hua Feng, P. Huber, J. Kittler, Zhiwu Huang, +13 more

    IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG) · 2015

    51
  • Exposure Trajectory Recovery From Motion Blur

    You-Jian Zhang, Chao-Yue Wang, S. Maybank, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2021

    Exposure trajectories are defined, which represent the motion information contained in a blurry image and explain the causes of motion blur, and a novel motion offset estimation framework is proposed to model pixel-wise displacements of the latent sharp image at multiple timepoints.

    50
  • Meta-Aggregator: Learning to Aggregate for 1-bit Graph Neural Networks

    Yongcheng Jing, Yiding Yang, Xinchao Wang, Mingli Song, D. Tao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    Two dedicated forms of meta neighborhood aggregators are proposed, an exclusive meta aggregator termed as Greedy Gumbel Neighborhood Aggregator (GNA), and a diffused meta aggregators termed as Adaptable Hybrid Neighborhood AggRegator (ANA), which learn to exclusively pick one single optimal aggregator from a pool of candidates.

    48
  • Learning to Track Multiple Targets

    Xiao Liu, D. Tao, Mingli Song, Lu-Ming Zhang, Jiajun Bu, Chun Chen

    IEEE Transactions on Neural Networks and Learning Systems · 2014

    47
  • Dynamic Contrastive Distillation for Image-Text Retrieval

    Jun Rao, Liang Ding, Shuhan Qi, Meng Fang, Yang Liu, Liqiong Shen, Dacheng Tao

    IEEE Transactions on Multimedia · 2023

    A novel plug-in dynamic contrastive distillation (DCD) framework to compress the large VLP models for the ITR task and proposes dynamic distillation to dynamically learn samples of different difficulties to balance better the difficulty of knowledge and students' self-learning ability.

    45
  • Heatmap Regression via Randomized Rounding

    B. Yu, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2021

    The proposed quantization system induced by the randomized rounding operation encodes the fractional part of numerical coordinates into the ground truth heatmap using a probabilistic approach during training and decodes the predicted numerical coordinates from a set of activation points during testing.

    45
  • Exploiting Local Coherent Patterns for Unsupervised Feature Ranking

    Qinghua Huang, D. Tao, Xuelong Li, Liang-Wen Jin, Gang Wei

    IEEE Transactions on Systems Man and Cybernetics Part B (Cybernetics) · 2011

    This paper aims to develop an unsupervised feature ranking algorithm that evaluates features using discovered local coherent patterns, known as biclusters, that can yield comparable or even better performance in comparison with the well-known Fisher score, Laplacian score, and variance score using three UCI data sets.

    45
  • Quantum Machine Learning: A Hands-on Tutorial for Machine Learning Practitioners and Researchers

    Yuxuan Du, Xin-Biao Wang, N. Guo, Zhan Yu, Yan Qian, Kai-Ning Zhang, Min-Hsiu Hsieh, P. Rebentrost, +1 more

    arXiv · 2025

    38
  • Bitcoin Mixing Detection Using Deep Autoencoder

    Lihao Nan, D. Tao

    IEEE International Conference on Data Science in Cyberspace · 2018

    36
  • Active Multi-task Learning via Bandits

    Meng Fang, D. Tao

    SIAM International Conference on Data Mining · 2015

    This paper introduces a new active multi- task learning paradigm, which selectively samples effective instances for multi-task learning and cast the selection procedure as a bandit framework, and provides an implementation of the algorithm based on a popular multi- Task learning algorithm that is trace-norm regularization method.

    36
  • Merging Models on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model Merging

    A. Tang, En-Neng Yang, Li Shen, Yong Luo, Han Hu, Bo Du, Da-Cheng Tao

    arXiv · 2025

    This study proposes a training-free projection-based continual merging method that processes models sequentially through orthogonal projections of weight matrices and adaptive scaling mechanisms, enabling efficient sequential integration of task-specific knowledge.

    35
  • SimDistill: Simulated Multi-Modal Distillation for BEV 3D Object Detection

    Haimei Zhao, Qiming Zhang, Shanshan Zhao, Zhe Chen, Jing Zhang, Da-Cheng Tao

    Proceedings of the AAAI Conference on Artificial Intelligence · 2024

    A Simulated multi-modal Distillation (SimDistill) method that supports intra-modal, cross-modal, and multi-modal fusion distillation simultaneously in the Bird's-eye-view space and can learn better feature representations for 3D object detection while maintaining a cost-effective camera-only deployment.

    35
  • Continual Learning on Graphs: Challenges, Solutions, and Opportunities

    Xikun Zhang, Dong-Jin Song, Dacheng Tao

    arXiv · 2024

    A comprehensive review of existing continual graph learning (CGL) algorithms is provided by elucidating the different task settings and categorizing the existing methods based on their characteristics, comparing the CGL methods with traditional continual learning techniques and analyzing the applicability of the traditional continual learning techniques to CGL tasks.

    34
  • Benchmarking Reasoning Robustness in Large Language Models

    Tong Yu, Yong-Cheng Jing, Xi-Kun Zhang, Wentao Jiang, Wenjie Wu, Yingjie Wang, Wenbin Hu, Bo Du, +1 more

    arXiv · 2025

    A novel benchmark is introduced, termed as Math-RoB, that exploits hallucinations triggered by missing information to expose reasoning gaps, achieved by an instruction-based approach to generate diverse datasets that closely resemble training distributions, facilitating a holistic robustness assessment and advancing the development of more robust reasoning frameworks.

    30
  • Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging

    Li Shen, A. Tang, En-Neng Yang, Gui-Bing Guo, Yong Luo, Le-Fei Zhang, Xiaochun Cao, Bo Du, +1 more

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2025

    An efficient-and-effective WEMoE (E-WEMoE) method, whose core mechanism involves eliminating non-essential elements in the critical modules of WEMoE and implementing shared routing across multiple MoE modules, thereby significantly reducing both the trainable parameters, the overall parameter count, and computational overhead of the merged model by WEMoE.

    29
  • Beyond Greedy Search: Tracking by Multi-Agent Reinforcement Learning-Based Beam Search

    Xiao Wang, Zhe Chen, Bo Jiang, Jin Tang, B. Luo, Da-Cheng Tao

    IEEE Transactions on Image Processing · 2022

    A novel multi-agent reinforcement learning based beam search tracking strategy, termed BeamTracking, which is mainly inspired by the image captioning task, which takes an image as input and generates diverse descriptions using beam search algorithm.

    28
  • Variance Reduced Methods for Non-convex Composition Optimization

    L. Liu, Ji Liu, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2021

    To significantly improve the query complexity of current approaches, the stochastic composition via variance reduction (SCVR) is devised and an extension to handle the mini-batch cases is proposed, which improve thequery complexity under the optimalmini-batch size.

    27
  • A novel method for speed training acceleration of recurrent neural networks

    J. Bilski, L. Rutkowski, Jacek Smoląg, D. Tao

    Information Sciences · 2020

    A particular approach for the Jordan network will be shown, however, the presented idea is applicable to other RNN structures and can be implemented in digital hardware.

    27
  • An Optimal Transport Analysis on Generalization in Deep Learning

    Jingwei Zhang, Tong-Liang Liu, D. Tao

    IEEE Transactions on Neural Networks and Learning Systems · 2021

    26
  • Quantum geometric machine learning for quantum circuits and control

    Elija Perrier, D. Tao, C. Ferrie

    New Journal of Physics · 2020

    This work demonstrates how geometric control techniques can be used to both verify the extent to which geometrically synthesised quantum circuits lie along geodesic, and thus time-optimal, routes and synthesise those circuits.

    23
  • 21
  • Multi-Step Denoising Scheduled Sampling: Towards Alleviating Exposure Bias for Diffusion Models

    Zhiyao Ren, Yi-Bing Zhan, Liang Ding, Gaoang Wang, Chao-Yue Wang, Zhongyi Fan, Da-Cheng Tao

    Proceedings of the AAAI Conference on Artificial Intelligence · 2024

    A multi-step denoising scheduled sampling (MDSS) strategy to alleviate the exposure bias in DDPMs and demonstrates that this approach is more effective in mitigating exposure bias in DDPM, DDIM, and DPM-solver.

    20
  • SIR: Self-Supervised Image Rectification via Seeing the Same Scene From Multiple Different Lenses

    Jinlong Fan, Jing Zhang, D. Tao

    IEEE Transactions on Image Processing · 2023

    A novel self-supervised image rectification (SIR) method based on an important insight that the rectified results of distorted images of a same scene from different lenses should be the same, with comparable or even better performance than the supervised baseline method and representative state-of-the-art (SOTA) methods.

    20
  • JustLogic: A Comprehensive Benchmark for Evaluating Deductive Reasoning in Large Language Models

    Michael Chen, Xikun Zhang, Dacheng Tao

    arXiv · 2025

    JustLogic is a synthetically generated deductive reasoning benchmark designed for rigorous evaluation of LLMs and reveals that state-of-the-art (SOTA) reasoning LLMs perform on par or better than the human average but significantly worse than the human ceiling, and SOTA non-reasoning models still underperform the human average.

    19
  • MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI

    Huan-Jin Yao, Jia-Xing Huang, Ya-Wen Qiu, Michael Chen, Wenzheng Liu, Wei Zhang, Wenjie Zeng, Xi-Kun Zhang, +4 more

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    This work introduces MMReason, a new benchmark designed to precisely and comprehensively evaluate MLLM long-chain reasoning capability with diverse, open-ended, challenging questions and designs a reference-based ternary scoring mechanism to reliably assess intermediate reasoning steps.

    19
  • Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving

    Ming Liu, Si-Yuan Liang, Koushik Howlader, Li-Wen Wang, Dacheng Tao, Wensheng Zhang

    arXiv · 2025

    A natural reflection-based backdoor attack targeting VLM systems in autonomous driving scenarios, aiming to induce substantial response delays when specific visual triggers are present, uncovering a new class of attacks that exploit the stringent real-time requirements of autonomous driving.

    19
  • Semantic-Preserving Adversarial Text Attacks

    Xing-Hao Yang, Yong-Shun Gong, Wei-Feng Liu, James Bailey, Tian-Qing Zhu, D. Tao, Wei Liu

    IEEE Transactions on Sustainable Computing · 2023

    A Bigram and Unigram-based adaptive adaptive Semantic Preservation Optimization (BU-SPO) approach which attacks text documents not only at a unigram word level but also at a bigram level to avoid generating meaningless sentences is devised.

    19
  • Fine-Grained Zero-Shot Learning: Advances, Challenges, and Prospects

    Jing-Cai Guo, Zhi-Jie Rao, Zhi Chen, Jingren Zhou, Da-Cheng Tao

    arXiv · 2024

    A broad review of recent advances for fine-grained analysis in ZSL is presented and a taxonomy of existing methods and techniques with a thorough analysis of each category is provided, covering publicly available datasets, models, implementations, and some more details as a library.

    17
  • 16
  • Multi-label Active Learning Based on Maximum Correntropy Criterion: Towards Robust and Discriminative Labeling

    Zeng-Mao Wang, Bo Du, Le-Fei Zhang, Liang-Pei Zhang, Meng Fang, D. Tao

    Lecture notes in computer science · 2016

    A novel active learning approach to reduce the annotation costs greatly for multi-label classification by merging uncertainty and representativeness by using the MCC to alleviate the influence of outlier labels for discriminative labeling.

    16
  • 13
  • Good Questions Help Zero-Shot Image Reasoning

    Kaiwen Yang, Tao Shen, Xinmei Tian, Xiu-Bo Geng, Chongyang Tao, Da-Cheng Tao, Tian-Yi Zhou

    arXiv · 2023

    Question-Driven Visual Exploration (QVix) is introduced, a novel prompting strategy that enhances the exploratory capabilities of LVLMs in zero-shot reasoning tasks and significantly outperforms existing methods, highlighting its effectiveness in bridging the gap between complex visual data and LVL Ms' exploratory abilities.

    12
  • Graph Reasoning Networks for Visual Question Answering.

    Dalu Guo, Chang Xu, D. Tao

    arXiv · 2019

    The graph reasoning networks are developed, and the resulting model can reason the relationship and dependence between objects, which leads to realization of multi-step reasoning.

    12
  • Discretized-Vapnik-Chervonenkis Dimension for Analyzing Complexity of Real Function Classes

    Chao Zhang, Wei Bian, D. Tao, Wei-Si Lin

    IEEE Transactions on Neural Networks and Learning Systems · 2012

    12
  • Graph-Augmented Reasoning: Evolving Step-by-Step Knowledge Graph Retrieval for LLM Reasoning

    Wenjie Wu, Yong-Cheng Jing, Yingjie Wang, Wenbin Hu, Dacheng Tao

    arXiv · 2025

    KG-RAR is proposed, a framework centered on process-oriented knowledge graph construction, a hierarchical retrieval strategy, and a universal post-retrieval processing and reward model (PRP-RM) that refines retrieved information and evaluates each reasoning step.

    11
  • Safety Reasoning with Guidelines

    Hao-Yu Wang, Zeyu Qin, Li Shen, Xueqian Wang, Minhao Cheng, Da-Cheng Tao

    arXiv · 2025

    Extensive experiments show that the pro-pose training model to perform safety reasoning for each query synthesize reasoning supervision based on pre-guidelines, training the model to reason in alignment with them, thereby effectively eliciting and utilizing latent knowledge from diverse perspectives.

    9
  • EMOv2: Pushing 5M Vision Model Frontier

    Jiang-Ning Zhang, Teng Hu, Hao-Yang He, Zhu-Cun Xue, Yabiao Wang, Chengjie Wang, Yong Li, Xiang-Tai Li, +1 more

    arXiv · 2024

    This work rethinks the lightweight infrastructure of efficient IRB and practical components in Transformer from a unified perspective, extending CNN-based IRB to attention-based models and abstracting a one-residual Meta Mobile Block (MMBlock) for lightweight model design and deduce a modern Improved Inverted Residual Mobile Block (i2RMB).

    9
  • Learning Versatile Convolution Filters for Efficient Visual Recognition

    Kai Han, Yunhe Wang, Chang Xu, Chun-Jing Xu, E. Wu, D. Tao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2021

    Experimental results on benchmark datasets and neural networks demonstrate that the versatile filters introduced are able to achieve comparable accuracy as that of original filters, but require less memory and computation cost.

    9
  • Deep Representation Calibrated Bayesian Neural Network for Semantically Explainable Face Inpainting and Editing

    Hao Xiong, Chao-Yue Wang, Xinchao Wang, D. Tao

    IEEE Access · 2020

    This work proposes a novel deep representation calibrated Bayesian neural network (DRCBNN) for semantically explainable face inpainting and editing and exploits deep representation into Bayesian decision theory and derive a deep representation calibration evidence lower bound (ELBO).

    9
  • 9
  • Rethink Sparse Signals for Pose-Guided Text-to-Image Generation

    Wenjie Xuan, Jing Zhang, Juhua Liu, Bo Du, Dacheng Tao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    A novel Spatial-Pose ControlNet (SP-Ctrl) is proposed, equipping sparse signals with robust controllability for pose-guided image generation and introduces keypoint concept learning, which encourages keypoint tokens to attend to the spatial positions of each keypoint, thus improving pose alignment.

    8
  • A Gentle Introduction to Quantum Machine Learning

    Yuxuan Du, Xin-Biao Wang, N. Guo, Zhan Yu, Yan Qian, Kai-Ning Zhang, Min-Hsiu Hsieh, P. Rebentrost, +1 more

    Springer eBooks · 2025

    8
  • On the Rates of Convergence From Surrogate Risk Minimizers to the Bayes Optimal Classifier

    Jingwei Zhang, Tong-Liang Liu, D. Tao

    IEEE Transactions on Neural Networks and Learning Systems · 2021

    This article introduces the notions of consistency intensity and conductivity to characterize a surrogate loss function and exploits this notion to obtain the rate of convergence from an empirical surrogate risk minimizer to the Bayes optimal classifier, enabling fair comparisons of the excess risks of different surrogaterisk minimizers.

    7
  • 6
  • On Gleaning Knowledge from Multiple Domains for Active Learning

    Zeng-Mao Wang, Bo Du, Le-Fei Zhang, Liang-Pei Zhang, Ruimin Hu, D. Tao

    International Joint Conference on Artificial Intelligence · 2017

    A framework that attempts to glean knowledge from multiple domains for active learning by querying the most uncertain and representative samples from the target domain and calculating the important weights for re-weighting the source data in a single unified formulation is proposed.

    6
  • Speech Sense Disambiguation: Tackling Homophone Ambiguity in End-to-End Speech Translation

    Tengfei Yu, Xue-Bo Liu, Liang Ding, Ke-Hai Chen, D. Tao, Min Zhang

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2024

    AmbigST is introduced, an innovative homophone-aware contrastive learning approach that integrates a homophone-aware masking strategy and achieves SOTA results on BLEU scores for English to German, Spanish, and French ST tasks, underlining its effectiveness in reducing speech sense ambiguity.

    5
  • Unlikelihood Tuning on Negative Samples Amazingly Improves Zero-Shot Translation

    Changtong Zan, Liang Ding, Li Shen, Yibin Lei, Yi-Bing Zhan, Wei-Feng Liu, Dacheng Tao

    arXiv · 2023

    The authors' analyses suggest that although they work well in ideal ON settings, language IDs become fragile and lose their navigation ability when faced with off-target tokens, which commonly exist during inference but are rare in training scenarios.

    5
  • Hyper-modal Imputation Diffusion Embedding with Dual-Distillation for Federated Multimodal Knowledge Graph Completion

    Ying Zhang, Yu Zhao, Xu-Hui Sui, Bao-Hang Zhou, Xiangrui Cai, Li Shen, Xiaojie Yuan, Da-Cheng Tao

    IEEE Transactions on Multimedia · 2026

    The Federated Multimodal Knowledge Graph Completion task, aiming at training over federated MKGs for better predicting the missing links in clients without sharing sensitive knowledge, is proposed, and a framework named MMFeD3-HidE for addressing multimodal uncertain unavailability and multimodal client heterogeneity challenges of FedMKGC is proposed.

    4
  • Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs

    Zixuan Hu, Yong-Xian Wei, Li Shen, Chun Yuan, Da-Cheng Tao

    arXiv · 2024

    This work explores the potential of reusing diverse pre-tuned LoRAs without accessing their original training data, to achieve tuning-free few-shot adaptation in VFMs and distills a meta-LoRA from diverse pre-tuned LoRAs with a meta-learning objective, using surrogate data generated inversely from pre-tuned LoRAs themselves.

    4
  • 4
  • SeWA: Selective Weight Average via Probabilistic Masking

    Peng Wang, Sheng-shou Hu, Zerui Tao, Guo-Xia Wang, Dian-Hai Yu, Li Shen, Quanchao Zheng, Da-Cheng Tao

    arXiv · 2025

    This paper proposes a simple yet efficient algorithm called Selective Weight Averaging (SeWA), which adaptively selects checkpoints during the final stages of training for averaging, and theoretically derive the SeWA's stability-based generalization bounds, which are sharper than that of SGD under both convex and non-convex assumptions.

    3
  • Stochastically Controlled Compositional Gradient for Composition Problems

    Liu Liu, Ji Liu, Cho-Jui Hsieh, D. Tao

    IEEE Transactions on Neural Networks and Learning Systems · 2021

    3
  • 3
  • Self-supervised Exposure Trajectory Recovery for Dynamic Blur Estimation.

    You-Jian Zhang, Chao-Yue Wang, S. Maybank, D. Tao

    arXiv · 2020

    Under mild constraints, the learned motion offsets can recover dense, (non-)linear exposure trajectories, which significantly reduce temporal disorder and ill-posed problems and further contribute to motion-aware image deblurring and warping-based video extraction from a single blurry image.

    3
  • Model Hemorrhage and the Robustness Limits of Large Language Models

    Ziyang Ma, Zu-Chao Li, Le-Fei Zhang, Gui-Song Xia, Bo Du, Liang-Pei Zhang, Dacheng Tao

    arXiv · 2025

    This work establishes foundational metrics for evaluating model stability during adaptation, providing practical guidelines for maintaining performance while enabling efficient LLM deployment and advancing understanding of neural network resilience under architectural transformations.

    2
  • InDecGAN: Learning to Generate Complex Images From Captions via Independent Object-Level Decomposition and Enhancement

    Jun Cheng, Fu-Xiang Wu, Liu Liu, Qieshi Zhang, L. Rutkowski, Da-Cheng Tao

    IEEE Transactions on Multimedia · 2023

    2
  • VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search

    Annual Meeting of the Association for Computational Linguistics (ACL) · 2025

    1
  • 1
  • Introduction

    Modern Machine Learning Techniques and Their Applications in Cartoon Animation Research · 2013

    1
  • Error Bounds for Real Function Classes Based on Discretized Vapnik-Chervonenkis Dimensions.

    Chao Zhang, D. Tao

    OPUS - Open Publications of UTS Scholars (University of Technology Sydney) · 2010

    This paper proposes the discretized VC dimension obtained by discretizing the range of a real function class, and points out that Sauer\'s Lemma is valid for this dimension.

    1
  • Reason-KE++: Aligning the Process, Not Just the Outcome, for Faithful LLM Knowledge Editing

    Yuchen Wu, Liang Ding, Li Shen, Dacheng Tao

    Findings of the Association for Computational Linguistics: ACL 2026 · 2026

    –
  • Quantum Transformer

    Springer eBooks · 2025

    –
  • Robust Knowledge Editing via Explicit Reasoning Chains for Distractor-Resilient Multi-Hop QA

    Yuchen Wu, Liang Ding, Li Shen, Dacheng Tao

    Findings of the Association for Computational Linguistics: EMNLP 2025 · 2025

    –
  • Conclusion

    Springer eBooks · 2025

    –
  • –
  • –
  • –
  • Responsible Active Learning via Human-in-the-loop Peer Study

    Yu Cao, Jingya Wang, Baosheng Yu, Dacheng Tao

    arXiv · 2022

    This work introduces a human-in-the-loop teacher-student architecture to isolate unlabelled data from the task learner on the cloud-side by maintaining an active learner (student) on the client-side and devise a discrepancy-based active sampling criterion, Peer Study Feedback, that exploits the variability of peer students to select the most informative data to improve model stability.

    –
  • –
  • Wide-angle Image Rectification: A Survey

    BIROn (Birkbeck, University of London) · 2020

    –
  • –
  • Batch Mode Active Learning for Geographical Image Classification

    Zeng-Mao Wang, Bo Du, Le-Fei Zhang, Wenbin Hu, D. Tao, Liang-Pei Zhang

    Lecture notes in computer science · 2015

    An innovative batch mode active learning by combining discriminative and representative information for hyperspectral image classification with support vector machine is proposed and a novel form of upper bound for true risk in the active learning setting is derived.

    –
  • Patch Alignment for Graph Embedding

    Yong Luo, D. Tao, Chao Xu

    Springer eBooks · 2012

    This paper presents locally linear embedding (LLE), which uses linear coefficients, which reconstruct a given example by its neighbors, to represent the local geometry, and then seeks a low-dimensional embedding, in which these coefficients are still suitable for reconstruction.

    –
  • –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.