Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Traced back 15 generations →Advisors
- Bernt SchieleFrom Academic Family Tree ↗
Students and postdocs4
- Zilin LuoFrom a thesis record ↗
- Jiahao YingFrom a thesis record ↗
- Yaoyao LiuFrom Academic Family Tree ↗
- Zhaozheng ChenFrom a thesis record ↗
Works101 from public data
- Meta-Transfer Learning for Few-Shot Learning1,293
A novel few-shot learning method called meta-transfer learning (MTL) which learns to adapt a deep NN for few shot learning tasks and introduces the hard task (HT) meta-batch scheme as an effective learning curriculum for MTL.
- 670
- Disentangled Person Image Generation574
A novel, two-stage reconstruction pipeline is proposed that learns a disentangled representation of the aforementioned image factors and generates novel person images at the same time and can manipulate the foreground, background and pose of the input image, and also sample new embedding features to generate targeted manipulations, that provide more control over the generation process.
- Feature Pyramid Transformer317
This work proposes a fully active feature interaction across both space and scales, called Feature Pyramid Transformer (FPT), which transforms any feature pyramid into another feature pyramid of the same size but with richer contexts, by using three specially designed transformers in self-level, top-down, and bottom-up interaction fashion.
- Visual Commonsense R-CNN299
A novel unsupervised feature representation learning method, Visual Commonsense Region-based Convolutional Neural Network (VC R-CNN), is presented to serve as an improved visual region encoder for high-level tasks such as captioning and VQA, and observes consistent performance boosts across them.
- Adaptive Aggregation Networks for Class-Incremental Learning296
A novel network architecture called Adaptive Aggregation Networks (AANets) is proposed in which two types of residual blocks are built at each residual level: a stable block and a plastic block, to balance stability and plasticity, dynamically.
- Counterfactual Zero-Shot and Open-Set Visual Recognition249
A novel counterfactual framework for both Zero-Shot Learning (ZSL) and Open-Set Recognition (OSR), whose common challenge is generalizing to the unseen-classes by only training on the seen-classes is presented.
- Class Re-Activation Maps for Weakly-Supervised Semantic Segmentation235
An embarrassingly simple yet surprisingly effective method: Reactivating the converged CAM with BCE by using softmax crossentropy loss (SCE), dubbed ReCAM, which not only generates high-quality masks, but also supports plug-and-play in any CAM variant with little overhead.
- 231
- Natural and Effective Obfuscation by Head Inpainting228
This work proposes a novel head inpainting obfuscation technique that generates realistic person images, while achieving superior obfuscation performance against automatic person recognizers.
- Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly Detection185
This work proposes a new MIL framework: Unbiased MIL (UMIL), to learn unbiased anomaly features that improve WSVAD and demonstrates the effectiveness of this framework on benchmarks UCF-Crime and TAD.
- Causal Attention for Unbiased Visual Recognition174
A causal attention module (CaaM) that self-annotates the confounders in unsupervised fashion is proposed that can be stacked and integrated in conventional attention CNN and self-attention Vision Transformer, and in OOD settings, deep models with CaaM outperform those without it significantly.
- RMM: Reinforced Memory Management for Class-Incremental Learning138
RMM is an optimizable and general method for memory management that can be used in any replaying-based CIL method and is an optimizable and general method for memory management that can be used in any replaying-based CIL method.
- An Ensemble of Epoch-Wise Empirical Bayes for Few-Shot Learning138
This paper proposes to meta-learn the ensemble of epoch-wise empirical Bayes models (E3BM) to achieve robust predictions and introduces four kinds of hyperprior learners by considering inductive vs. transductive, and epoch-dependent vs. epoch-independent, in the paradigm of meta-learning.
- A Hybrid Model for Identity Obfuscation by Face Replacement136
This work proposes a new hybrid approach to obfuscate identities in photos by head replacement that improves over the previous state of the art in obfuscation rate while preserving a higher similarity to the original image content.
- A Large-Scale Benchmark for Food Image Segmentation130
A new food image dataset FoodSeg103 (and its extension FoodSeg154) containing 9,490 images is built and a multi-modality pre-training approach called ReLeM that explicitly equips a segmentation model with rich and semantic food knowledge is proposed.
- 122
- Online growing neural gas for anomaly detection in changing surveillance scenes120
A neural network based model called online Growing Neural Gas (online GNG) to perform an unsupervised learning to perform anomaly detection and effectively reduces the false alarms and leak detections caused by model aging which frequently happens in changing surveillance scenes.
- A Domain Based Approach to Social Relation Recognition119
This paper provides the first dataset built on this holistic conceptualization of social life that is composed of a hierarchical label space of social domains and social relations and contributes the first models to recognize such domains and relations and find superior performance for attribute based features.
- Frame-Voyager: Learning to Query Frames for Video Large Language Models96
This paper proposes Frame-Voyager, a new data collection and labeling pipeline that learns to query informative frame combinations, based on the given textual queries in the task, and demonstrates its potential as a plug-and-play solution for Video-LLMs.
- Extracting Class Activation Maps from Non-Discriminative Features as well96
This work introduces a new computation method for CAM that explicitly captures non-discriminative features as well, thereby expanding CAM to cover whole objects and evaluates it in the challenging tasks of weakly-supervised semantic segmentation (WSSS), and plugs it in multiple state-of-the-art WSSS methods by simply replacing their original CAM with the authors'.
- Freestyle Layout-to-Image Synthesis95
This work introduces a new module called Rectified Cross-Attention (RCA) that can be conveniently plugged in the diffusion model to integrate semantic masks, and opts to leverage large-scale pre-trained text-to-image diffusion models to achieve the generation of unseen semantics.
- Class-Incremental Exemplar Compression for Class-Incremental Learning89
This paper proposes an adaptive mask generation model called class-incremental masking (CIM) to explicitly resolve two difficulties of using CAM: 1) transforming the heatmaps of CAM to 0–1 masks with an arbitrary threshold leads to a trade-off between the coverage on discriminative pixels and the quantity of exemplars, as the total memory is fixed; and 2) optimal thresholds vary for different object classes, which is particularly obvious in the dynamic environment of CIL.
- Transporting Causal Mechanisms for Unsupervised Domain Adaptation73
Transporting Causal Mechanisms (TCM) is proposed, to identify the confounder stratum and representations by using the domain-invariant disentangled causal mechanisms, which are discovered in an unsupervised fashion.
- Teacher-Student Networks with Multiple Decoders for Solving Math Word Problem72
This paper proposes a novel approach, TSN-MD, by leveraging the teacher network to integrate the knowledge of equivalent solution expressions and then to regularize the learning behavior of the student network to addressMath word problem challenges.
- 70
- Weakly-supervised Semantic Segmentation with Image-level Labels: From Traditional Models to Foundation Models59
This work conducts a comprehensive survey on traditional methods of WSSS, and investigates the applicability of visual foundation models, such as the Segment Anything Model (SAM), in the context of WSSS.
- Online Hyperparameter Optimization for Class-Incremental Learning58
This work designs an online learning method that can adaptively optimize the stability-plasticity tradeoff without knowing the setting as a priori, and consistently improves top-performing CIL methods in both TFH and TFS settings.
- Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models50
This study systematically examine the biases in foundation models and demonstrates the efficacy of the proposed Generalized Logit Adjustment (GLA) method, an optimization-based bias estimation approach for debiasing foundation models.
- Self-Regulation for Semantic Segmentation46
This paper seeks reasons for the two major failure cases in Semantic Segmentation (SS): 1) missing small objects or minor object parts, and 2) mislabeling minor parts of large objects as wrong classes and introduces several Self-Regulation (SR) losses for training SS neural networks.
- Deconfounded Visual Grounding40
This work frames the visual grounding pipeline into a causal graph, which shows the causalities among image, query, target location and underlying confounder, and proposes a confounders-agnostic approach called Referring Expression Deconfounder (RED), to remove the confounding bias.
- Exploring Diffusion Time-steps for Unsupervised Representation Learning39
A theoretical framework is built that connects the diffusion time-steps and the hidden attributes, which serves as an effective inductive bias for unsupervised learning of the modular attributes.
- Make the U in UDA Matter: Invariant Consistency Learning for Unsupervised Domain Adaptation39
This work proposes to make the U in UDA matter by giving equal status to the two domains, and learns an invariant classifier whose prediction is simultaneously consistent with the labels in the source domain and clusters in the target domain, hence the spurious correlation inconsistent in thetarget domain is removed.
- 39
- Mixed-dish Recognition with Contextual Relation Networks39
A novel approach called contextual relation networks (CR-Nets) is proposed that encodes the implicit and explicit contextual relations among multiple dishes using region-level features and label-level co-occurrence, respectively, inspired by the intuition that people are likely to choose dishes with common eating habits.
- Revisiting Local Descriptor for Improved Few-Shot Classification38
A new method, named DCAP, is presented, in which it is shown that the reliance on sophisticated classifiers is not necessary, and a simple classifier applied directly to improved feature embeddings can instead outperform most of the leading methods in the literature.
- A compact representation of human actions by sliding coordinate coding37
This article proposes to encode the relative position of visual words using a simple but very compact method called sliding coordinates coding (SCC), which is more compact than many of the spatial or spatial–temporal pooling methods in the literature.
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization35
This paper introduces Unsupervised Visual CoT (UV-CoT), a novel framework for image-level CoT reasoning via preference optimization that can improve visual comprehension, particularly in spatial reasoning tasks where textual descriptions alone fall short.
- Compositional Prompt Tuning with Motion Cues for Open-vocabulary Video Relation Detection35
This paper presents Relation Prompt (RePro) for Open-vocabulary Video Visual Relation Detection (Open-VidVRD), where conventional prompt tuning is easily biased to certain subject-object combinations and motion patterns.
- 34
- 34
- 25
- Semantic Scene Completion with Cleaner Self23
The 3D occupancy feature and the semantic relations of the “cleaner self” to supervise the counterparts of the "noisy self" to respectively address the above two incorrect predictions.
- Salient pairwise spatio-temporal interest points for real-time activity recognition23
The distribution of STIPs is organized into a salient directed graph, which reflects salient motions and can be divided into a time salient directedgraph and a space salient directed graphs, aiming at adding spatio-temporal discriminant to BoVW.
- 23
- 21
- Few-Shot Learner Parameterization by Diffusion Time-Steps20
Time-step Few-shot (TiF) learner significantly outperforms OpenCLIP and its adapters on a variety of fine-grained and customized few-shot learning tasks.
- Wakening Past Concepts without Past Data: Class-Incremental Learning from Online Placebos20
This paper trains an online placebo selection policy to quickly evaluate the quality of streaming images and use only good ones for one-time feed-forward computation of KD, and introduces an online learning algorithm to solve this MDP problem without causing much computation costs.
- Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models18
This paper proposes to automate dataset updating and provides systematic analysis regarding its effectiveness in dealing with benchmark leakage issue, difficulty control, and stability, and is the first to automate updating benchmarks for reliable and timely evaluation.
- Action Disambiguation Analysis Using Normalized Google-Like Distance Correlogram18
Normalized Google-Like Distance (NGLD) is proposed to numerically measuring this co-occurrence, due to its effectiveness in semantic correlation analysis and is proved a much richer descriptor by observably reducing action ambiguity in experiments, conducted on WEIZMANN dataset and the more challenging UCF sports.
- Invariant Training 2D-3D Joint Hard Samples for Few-Shot Point Cloud Recognition17
This work tackles the data scarcity challenge in few-shot point cloud recognition of 3D objects by using a joint prediction from a conventional 3D model and a well-trained 2D model, and proposes an invariant training strategy, called INVJOINT, which can learn more collaborative 2D and 3D representations for better ensemble.
- 16
- A novel hierarchical Bag-of-Words model for compact action representation16
This work proposes to compress residual vectors into low-dimensional residual histograms by the simple but efficient BoW quantization, which yields a hierarchical BoW (HBoW) model which is not only compact but also informative.
- 16
- 14
- Mixed Dish Recognition With Contextual Relation and Domain Alignment14
The contextual relation network is proposed that encodes the implicit and explicit contextual relations among multiple dishes from region-level features and label-level co-occurrence respectively and the domain adaption networks are introduced to align both local and global features, and eliminating domain gaps of dish features across different canteens.
- LCC: Learning to Customize and Combine Neural Networks for Few-Shot Learning14
This work aims to meta-learn how to effectively combine several base-learners and proposes to learn not only a single base-learner but an ensemble of several base-learners to obtain more robust results.
- 13
- 3D Question Answering via only 2D Vision-Language Models12
It is believed that 2D LVLMs are currently the most effective alternative (of the resource-intensive 3D LVLMs) for addressing 3D tasks, while cdViews achieves state-of-the-art performance in 3D-QA while relying solely on 2D models without fine-tuning.
- 12
- 12
- Translate-Train Embracing Translationese Artifacts11
It is found that artifacts have common patterns in different languages and can be modeled by deep learning, and subsequently proposed an approach to conduct translate-train using Translationese Embracing the effect of Artifacts (TEA), which outperforms strong baselines.
- Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames10
An automatic video compression method based on a bilevel optimization program (BOP) consisting of both model-level and frame-level optimizations called the Meta-Optimized Frames (MOF) approach is introduced, showing that MOF is a generic and efficient method that boost multiple baseline methods, and can achieve a new state-of-the-art performance.
- Learning to teach and learn for semi-supervised few-shot image classification10
The proposed LTTL combines the power of meta-learning and self-training, achieving superior performance compared with the baseline methods on two public benchmarks, and uses the meta- learning paradigm to optimize the parameters in the whole framework.
- Open-set domain adaptation by deconfounding domain gaps9
This work proposes a module of ensembling multiple transformations (EMT) to produce calibrated recognition scores, i.e., reliable normality scores, for the samples in the target domain, because of its advanced ability of correctly recognizing unknown classes.
- COSY: COunterfactual SYntax for Cross-Lingual Understanding9
This work includes the design of SYntax-aware networks as well as a COunterfactual training method to implicitly force the networks to learn not only the semantics but also the syntax, based on the observation that universal syntax is transferable across different languages.
- Human activity prediction by mapping grouplets to recurrent Self-Organizing Map9
Experimental results confirm that the method is very efficient for predicting human activity and yields better performance than state-of-the-art works.
- Inferring Ongoing Human Activities Based on Recurrent Self-Organizing Map Trajectory9
The Recurrent SelfOrganizing Map (RSOM), which was designed to process sequential data, is novelly adopted in this paper for the high-level representation of ongoing activities and the innovation lies that observed features and their spatio-temporal contexts are encoded in a trajectory of the pre-trained RSOM units.
- Unleashing Network Potentials for Semantic Scene Completion8
The proposed AMMNet introduces a cross-modal modulation enabling the interdependence of gradient flows between modalities, and a customized adversarial training scheme leveraging dynamic gradient competition, providing a promising direction for improving the effectiveness and generalization of SSC methods.
- 8
- Meta-Learning Hyperparameters for Parameter Efficient Fine-Tuning7
MetaPEFT is proposed, a method incorporating adaptive scalers that dynamically adjust module influence during fine-tuning that achieves state-of-the-art performance in cross-spectral adaptation, requiring only a small amount of trainable parameters and improving tail-class accuracy significantly.
- 7
- Meta-Aggregating Networks for Class-Incremental Learning6
This work proposes a novel network architecture called Meta-Aggregation Networks (MANets), in which two residual blocks are built at each residual level: a stable block and a plastic block, and meta-learn the aggregation weights in order to dynamically optimize and balance between the two types of blocks.
- 6
- Self-Refining Deep Symmetry Enhanced Network for Rain Removal6
Deep Symmetry Enhanced Network (DSEN) is proposed that is able to explicitly extract the rotation equivariant features from rain images and a self-refining mechanism to remove the accumulated rain streaks in a coarse-to-fine manner.
- 5
- 5
- 4
- Towards Natural Image Matting in the Wild via Real-Scenario Prior4
This work proposes SEMat, a new matting dataset based on the COCO dataset, namely COCO-Matting, which revamps the network architecture and training objectives and proves its efficacy in interactive natural image matting.
- Interventional Training for Out-Of-Distribution Natural Language Understanding4
This paper proposes a novel interventional training method called Bottom-up Automatic Intervention (BAI) that performs multi-granular intervention with identified multifactorial confounders and shows the effectiveness of BAI for tackling OOD settings.
- 4
- Generating expensive relationship features from cheap objects4
A novel Semantic Transform Generative Adversarial Network (ST-GAN) that synthesizes relationship features for rare objects, conditioned on the features from random instances of the objects, conditioned on the features from random instances of the objects is proposed.
- 3
- Attention-based Class Activation Diffusion for Weakly-Supervised Semantic Segmentation3
A new method to couple CAM and Attention matrix in a probabilistic Diffusion way is proposed, and it is shown that AD-CAM as pseudo labels can yield stronger WSSS models than the state-of-the-art variants of CAM.
- 2
- 2
- 2
- 2
- 2
- 1
- LLMs-Based Augmentation for Domain Adaptation in Long-Tailed Food Datasets1
This paper first leverage LLMs to parse food images to parse food images to generate food titles and ingredients, and project the generated texts and food images from different domains to a shared embedding space to maximize the pair similarities.
- 1
- Non-Visible Light Data Synthesis and Application: A Case Study for Synthetic Aperture Radar Imagery1
A novel prototype LoRA is introduced, as an improved version of 2LoRA, to resolve the class imbalance problem in SAR datasets, and augmentation, when integrated into the training process of SAR classification as well as segmentation models, yields notably improved performance for minor classes.
- Synthesizing Multi-Person and Rare Pose Images for Human Pose Estimation–
This paper designs a controllable pose generator named PoseFactory and introduces a multi-person image generator named MultipGenerator, which is conditioned on multiple human poses and textual descriptions of complex scenes and demonstrates its superior performance both quantitatively and qualitatively.
- SEMat: Semantic Enhanced Natural Image Interactive Matting–
This work proposes SEMat, a new matting dataset based on the COCO dataset, namely COCO-Matting, which selects real-world complex images from COCO and converts semantic segmentation masks to matting labels and revamps the network architecture and training objectives.
- –
- Generalized Visual Relation Detection With Diffusion Models–
Benefiting from the diffusion-based generative process, the Diff-VRD is able to generate visual relations beyond the pre-defined category labels of datasets, and is able to generate visual relations beyond the pre-defined category labels of datasets.
- –
- In-Context Translation: Towards Unifying Image Recognition, Processing, and Generation–
In-Context Translation (ICT), a general learning framework to unify visual recognition, low-level image processing, and conditional image generation, and edge-to-image synthesis, is proposed.
- –
- Objects, Relationships, and Context in Visual Data–
This tutorial will introduce a various of machine learning techniques for modeling visual relationships and contextual generative models, starting from fundamental theories on object detection, relationship detection, generative adversarial networks, to more advanced topics on referring expression visual grounding, pose guided person image generation, and context based image inpainting.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.