Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
View as a treeStudents and postdocs4
- Yufei WangFrom a thesis record ↗
- Eleanor TennantFrom a thesis record ↗
- Rongkai ZhangFrom a thesis record ↗
- Tao BaiFrom a thesis record ↗
Works240 from public data
- Non-local recurrent network for image restoration706
A non- local recurrent network (NLRN) is proposed as the first attempt to incorporate non-local operations into a recurrent neural network (RNN) for image restoration and achieves superior results to state-of-the-art methods with much fewer parameters.
- Recent Advances in Adversarial Training for Adversarial Robustness668
For the first time, a systematically review the recent progress on adversarial training for adversarial robustness with a novel taxonomy and highlights the challenges which are not fully tackled.
- Denoising Diffusion Models for Plug-and-Play Image Restoration511
Experimental results on three representative IR tasks demonstrate that DiffPIR achieves state-of-the-art performance on both the FFHQ and ImageNet datasets in terms of reconstruction faithfulness and perceptual quality with no more than 100 NFEs.
- 465
- SinSR: Diffusion-Based Image Super-Resolution in a Single Step378
This work first derive a deterministic sampling process from the most recent state-of-the-art (SOTA) method for accelerating diffusion-based SR, which can be distilled into a student model that performs SR within only one inference step, resulting in a remarkable up to × 10 speedup for inference.
- 279
- When Image Denoising Meets High-Level Vision Tasks: A Deep Learning Approach248
This paper proposes a convolutional neural network for image denoising which achieves the state-of-the-art performance and proposes a deep neural network solution that cascades two modules for image Denoising and various high-level tasks, respectively, and uses the joint loss for updating only the denoise network via back-propagation.
- Connecting Image Denoising and High-Level Vision Tasks via Deep Learning200
A convolutional neural network in which convolutions are conducted in various spatial resolutions via downsampling and upsampling operations in order to fuse and exploit contextual information on different scales and the benefit of exploiting image semantics simultaneously for image denoising and high-level vision tasks via deep learning is demonstrated.
- ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow Removal196
A unified diffusion framework that integrates both the image and degradation priors for highly effective shadow removal and progressively refines the estimated shadow mask as an auxiliary task of the diffusion generator, which leads to more accurate and robust shadow-free image generation.
- Structured Overcomplete Sparsifying Transform Learning with Convergence Guarantees and Applications174
The promising performance of the proposed approach in image denoising is shown, which compares quite favorably with approaches involving a single learned square transform or an overcomplete synthesis dictionary, or gaussian mixture models.
- 150
- HyperService139
The experiments show that HyperService imposes reasonable latency, in order of seconds, on the end-to-end execution of cross-chain applications, and the HyperService platform is scalable to continuously incorporate new large-scale production blockchains.
- From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image Restoration123
Experimental results demonstrate that the proposed RRC method outperforms many state-of-the-art schemes in both the objective and perceptual quality.
- A Benchmark for Sparse Coding: When Group Sparsity Meets Rank Minimization121
An adaptive dictionary is designed to bridge the gap between group-based sparse coding (GSC) and rank minimization, and WSNM is found to be the closest one to the real singular values of each patch group and is translated into a non-convex weighted norm minimization problem in GSC.
- 117
- 113
- 104
- 99
- 97
- 94
- ExposureDiffusion: Learning to Expose for Low-light Image Enhancement89
This work seamlessly integrating a diffusion model with a physics-based exposure model to achieve significantly improved performance and reduced inference time and proposes an adaptive residual layer that effectively screens out the side-effect in the iterative refinement when the intermediate results have been already well-exposed.
- DVS-Voltmeter: Stochastic Process-Based Event Simulator for Dynamic Vision Sensors79
An event simulator, dubbed DVS-Voltmeter, is proposed to enable high-performance deep networks for DVS applications and indicates that neural networks trained with DVS-Voltmeter generalize favorably on real events against state-of-the-art simulators.
- 75
- Removing Backdoor-Based Watermarks in Neural Networks with Limited Data72
A novel backdoor-based watermark removal framework using limited data, dubbed WILD, which can effectively remove the watermarks without compromising the deep model performance for the original task with the limited access to training data is proposed.
- ReLLIE: Deep Reinforcement Learning for Customized Low-Light Image Enhancement72
A novel deep reinforcement learning based method, dubbed ReLLIE, for customized low-light enhancement, which models LLIE as a markov decision process, i.e., estimating the pixel-wise image-specific curves sequentially and recurrently.
- Ground-based image analysis: A tutorial on machine-learning techniques and applications71
The advantages of using machine-learning techniques in ground-based image analysis via three primary applications: segmentation, classification, and denoising are demonstrated.
- ShadowFormer: Global Context Helps Image Shadow Removal67
A Retinex-based shadow model is proposed, from which a novel transformer-based network is derived, dubbed ShandowFormer, to exploit non-shadow regions to help shadow region restoration and achieves state-of-the-art performance by using up to 150X fewer model parameters.
- 65
- Image Recovery via Transform Learning and Low-Rank Modeling: The Power of Complementary Regularizers59
- 59
- 58
- Disentangled Feature Representation for Few-Shot Image Classification57
This work proposes a novel disentangled feature representation (DFR) framework, dubbed DFR, for few-shot learning applications, which can adaptively decouple the discriminative features that are modeled by the classification branch, from the class-irrelevant component of the variation branch.
- 57
- Learning to Solve Multiple-TSP With Time Window and Rejections via Deep Reinforcement Learning57
Experimental results demonstrate that the proposed framework outperforms strong baselines in terms of higher solution quality and shorter computation time, and the trained agents also achieve competitive performance for solving unseen larger instances.
- 53
- 50
- FRIST—flipping and rotation invariant sparsifying transform learning and applications46
This work develops a methodology for learning flipping and rotation invariant sparsifying transforms, dubbed FRIST, to better represent natural images that contain textures with various geometrical directions and provides a convergence guarantee, and demonstrates the empirical convergence behavior of the proposed FRIST learning approach.
- FRIST - Flipping and Rotation Invariant Sparsifying Transform Learning and Applications to Inverse Problems46
This work develops a methodology for learning flipping and rotation invariant sparsifying transforms, dubbed FRIST, to better represent natural images that contain textures with various geometrical directions and provides a convergence guarantee, and demonstrates the empirical convergence behavior of the proposed FRIST learning approach.
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models41
Experimental results show that FailSafe-VLM successfully helps robotic arms detect and recover from potential failures, improving the performance of three state-of-the-art VLA models by up to 22.6% on average across several tasks in ManiSkill.
- 41
- VIDOSAT: High-Dimensional Sparsifying Transform Learning for Online Video Denoising40
A novel framework for online video denoising based on high-dimensional sparsifying transform learning for spatio-temporal patches is presented and the proposed methods outperform several related and recent techniques, includingDenoising with 3D DCT, prior schemes based on dictionary learning, non-local means, background separation, and deep learning, as well as the popular VBM3D and VBM4D.
- 36
- Generating Person Images with Appearance-aware Pose Stylizer36
A novel end-to-end framework to generate realistic person images based on given person poses and appearances called Appearance-aware Pose Stylizer (APS) which generates human images by coupling the target pose with the conditioned person appearance progressively.
- 35
- 34
- SOUL-Net: A Sparse and Low-Rank Unrolling Network for Spectral CT Image Reconstruction33
A sparse and low-rank unrolling network (SOUL-Net) is proposed for spectral CT image reconstruction, that learns the parameters and thresholds in a data-driven manner and outperforms several representative state-of-the-art algorithms in terms of detail preservation and artifact reduction.
- Graph Neural Networks With Triple Attention for Few-Shot Learning33
The proposed Attentive GNN model outperforms the state-of-the-art few-shot learning methods using both GNN and non-GNN approaches and is consistent over the mini-Image net, tiered-ImageNet, CUB-200-2011, and Flowers-102 benchmarks.
- Making Your First Choice: To Address Cold Start Problem in Vision Active Learning33
This paper seeks to address the cold start problem in vision active learning by exploiting the three advantages of contrastive learning: no annotation is required; label diversity is ensured by pseudo-labels to mitigate bias; typical data is determined by contrastive features to reduce outliers.
- Deep learning‐enabled imaging flow cytometry for high‐speed Cryptosporidium and Giardia detection33
A deep learning‐enabled high‐throughput system for predicting Cryptosporidium and Giardia in drinking water that combines imaging flow cytometry and an efficient artificial neural network called MCellNet, which achieves a classification accuracy >99.6%.
- 30
- Raw Image Reconstruction with Learned Compact Metadata29
A novel framework to learn a compact representation in the latent space serving as the metadata in an end-to-end manner with the improved entropy estimation strategies, which leads to better reconstruction quality, smaller size of metadata, and faster speed.
- 29
- Targeted Universal Adversarial Examples for Remote Sensing29
Extensive experiments showed strong attackability of the two targeted adversarial variants of universal adversarial examples, and it is hoped such strong attacks can inspire and motivate research on the defenses against adversarialExamples in remote sensing.
- 28
- Towards More Efficient Security Inspection via Deep Learning: A Task-Driven X-ray Image Cropping Scheme28
A Task-Driven Cropping scheme, dubbed TDC, for improving the deep image detection algorithms towards efficient and effective luggage inspection via X-ray images, which shows that the proposed TDC algorithm can effectively boost popular detection algorithms, by achieving better detection mAPs or reducing the run time.
- Revisiting One-Stage Deep Uncalibrated Photometric Stereo via Fourier Embedding27
A one-stage deep uncalibrated photometric stereo (UPS) network, namely Fourier Uncalibrated Photometric Stereo Network (FUPS-Net), is introduced, offering better training stability, a concise end-to-end structure, and avoiding accumulated errors in disjointed networks.
- Faster nonconvex low-rank matrix learning for image low-level and high-level vision: A unified framework27
This study introduces a unified approach to tackle challenges in both low-level and high-level vision tasks for image processing, which incorporates the randomized singular value decomposition (RSVD) technique and introduces a continuous strategy for adaptive model parameter updates.
- Exploiting Non-Local Priors via Self-Convolution for Highly-Efficient Image Restoration26
It is proved that the proposed Self-Convolution based formulation can generalize the commonly-used non-local modeling methods, as well as produce results equivalent to standard methods, but with much cheaper computation.
- A joint learning method for low-light facial expression recognition25
A novel low-light FER framework (termed LL-FER) that can simultaneously enhance the images and recognition tasks of low-light facial expression images and a joint loss to train the LLENet with FER network in a cascade manner is proposed.
- 25
- 23
- Progressive Divide-and-Conquer via Subsampling Decomposition for Accelerated MRI23
This work presents a rigorous derivation of the pro-posed PDAC framework, which could be further unfolded into an end-to-end trainable network and achieves superior performance on the pub-licly available fastMRI and Stanford2D FSE datasets in both multi-coil and single-coil settings.
- 23
- DEEMO: De-identity Multimodal Emotion Recognition and Reasoning22
The De-identity Multimodal Emotion Recognition and Reasoning (DEEMO), a novel task designed to enable emotion understanding using de-identified video and audio inputs, is introduced and a Multimodal Large Language Model (MLLM) that integrates de-identified audio, video, and textual information to enhance both emotion recognition and reasoning is proposed.
- 22
- 21
- 20
- Site Selection via Learning Graph Convolutional Neural Networks: A Case Study of Singapore20
This work proposes to learn a graph convolutional network (GCN) for highly effective site selection tasks and presents a novel dataset that encompasses land use information as well as public transport networks in Singapore as a case study to benchmark site selection algorithms.
- Attentive Graph Neural Networks for Few-Shot Learning20
This work proposes a novel Attentive GNN (AGNN) to tackle few-shot learning challenges by incorporating a triple-attention mechanism, i.e., node self-att attention, neighborhood attention, and layer memory attention.
- Benchmarking White Blood Cell Classification under Domain Shift19
This paper establishes a benchmark for WBC recognition and indicates that CNN-based models achieve high accuracy when trained and tested under similar imaging conditions, however, their performance drops significantly when tested under different conditions.
- Provenance of Training without Training Data: Towards Privacy-Preserving DNN Model Ownership Verification19
This paper proposes a novel Provenance of Training (PoT) scheme, the first empirical study towards verifying DNN model ownership without accessing any original dataset while being robust against existing attacks.
- 19
- Denoising with weak signal preservation by group-sparsity transform learning19
This work has developed a novel transform learning with group sparsity (TLGS) method that jointly exploits local sparsity and internal patch self-similarity and achieves better denoising performance than existing denoised methods, in terms of signal-to-noise ratio values and visual preservation of weak signal.
- 18
- Single-Image Shadow Removal Using Deep Learning: A Comprehensive Survey18
This paper highlights the major advancements in deep learning-based single-image shadow removal methods, thoroughly review previous research across various categories, and provides insights into the historical progression of these developments.
- 18
- 18
- Two-Way Generation of High-Resolution EO and SAR Images via Dual Distortion-Adaptive GANs17
A novel image translation algorithm designed to tackle geometric distortions adaptively is proposed, and its superiority over other methods for both SAR to EO (S2E) and EO to SAR (E2S) tasks is validated, especially for urban areas in high-resolution images.
- The Power Of Triply Complementary Priors For Image Compressive Sensing17
A joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely external and internal, deep and shallow, and local and nonlocal priors is proposed, which is a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for image CS.
- 16
- 16
- Towards Adversarially Robust Continual Learning16
This work is the first to study adversarial robustness in continual learning and proposes a novel method called Task-Aware Boundary Augmentation (TABA) to boost the robustness of continual learning models.
- 16
- 16
- 15
- 15
- 15
- 15
- 15
- M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision14
The initial attempt to build a Generalized RGBT MultiSpectral foundation model (M-SpecGene), which aims to learn modalityinvariant representations from large-scale broad data in a self-supervised manner, provides new insights into multispectral fusion and integrates prior case-by-case studies into a unified paradigm.
- 14
- 14
- 14
- 14
- 13
- PIP: Physical Interaction Prediction via Mental Simulation with Span Selection13
The experiments show that PIP outperforms human, baseline, and related intuitive physics models that utilize mental simulation, and PIP's span selection module effectively identifies the frames indicating key physical interactions among objects, allowing for added interpretability.
- 13
- 13
- SCORE: Scene Context Matters in Open-Vocabulary Remote Sensing Instance Segmentation12
This work proposes SCORE (Scene Context matters in Open-vocabulary REmote sensing instance segmentation), a framework that integrates multigranularity scene context, i.e., regional context and global context, to enhance both visual and textual representations.
- Evolving Storytelling: Benchmarks and Methods for New Character Customization with Diffusion Models12
The NewEpisode benchmark is introduced, comprising refined datasets designed to evaluate generative models' adaptability in generating new stories with fresh characters using just a single example story and EpicEvo is proposed, a method that customizes a diffusion-based visual story generation model with a single story featuring the new characters seamlessly integrating them into established character dynamics.
- Beyond Learned Metadata-Based Raw Image Reconstruction12
A novel framework that learns a compact representation in the latent space, serving as metadata, in an end-to-end manner is proposed, achieving high-quality raw image reconstruction with a smaller metadata size, compared with existing SOTA methods.
- 12
- 12
- Reconciling Stochastic and Deterministic Strategies for Zero-shot Image Restoration using Diffusion Model in Dual11
This work proposes a novel zero-shot IR scheme, dubbed Reconciling Diffusion Model in Dual (RDMD), which leverages only a single pre-trained diffusion model to construct two complementary regularizers, aiming to achieve highfidelity image restoration with appealing perceptual quality.
- 11
- 11
- Benchmarking Adversarial Robustness of Image Shadow Removal with Shadow-Adaptive Attacks11
A novel approach to shadow-adaptive adversarial attack, where the optimized adversarial noise in the shadowed regions becomes visually less perceptible while permitting a greater tolerance for perturbations in non-shadow regions.
- Temporal Output Discrepancy for Loss Estimation-Based Active Learning11
A novel deep active learning approach that queries the oracle for data annotation when the unlabeled sample is believed to incorporate high loss, and shows that TOD can be utilized to select the best model of potentially the highest testing accuracy from a pool of candidate models.
- 11
- 11
- Hyper RPCA: Joint Maximum Correntropy Criterion and Laplacian Scale Mixture Modeling on-the-Fly for Moving Object Detection11
The proposed Hyper RPCA jointly applies the maximum correntropy criterion (MCC) for the modeling error, and Laplacian scale mixture (LSM) model for foreground objects to detect moving objects on the fly.
- 11
- 11
- EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model10
A dataset for Emotion Analysis in Long-sequential and De-identity videos called EALD is constructed by collecting and processing the sequences of athletes' post-match interviews and evaluating the Multimodal Large Language Models with de-identification signals (e.g., visual, speech, and NFBLs) to perform emotion analysis.
- 10
- 10
- 10
- Joint Patch-Group Based Sparse Representation for Image Inpainting10
Compared with existing sparse representation models, the proposed JPG-SR provides a powerful mechanism to integrate the local sparsity and nonlocal self-similarity of images and outperforms several state-of-the-art methods in both objective and perceptual quality.
- 9
- STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning9
This work proposes a discriminative Temporal-based Subspace Classifier (TSC) that represents each class with an orthogonal subspace basis and adopts subspace projection loss for classification, and implements inter-and intra-class orthogonal constraints into TSC.
- 9
- Compressed Event Sensing (CES) Volumes for Event Cameras9
CES volumes preserve the high temporal resolution of event streams by leveraging the sparsity property of events and the principles of compressed sensing theory, and effectively capture the frequency characteristics of events in low-dimensional representations.
- 9
- Parameter-Free Style Projection for Arbitrary Image Style Transfer9
A real-time feed-forward model to leverage Style Projection for arbitrary image style transfer, which includes a regularization term for matching the semantics between input contents and stylized outputs.
- 9
- 9
- RF4D:Neural Radar Fields for Novel View Synthesis in Outdoor Dynamic Scenes8
RF4D is presented, a radar-based neural field framework tailored for novel view synthesis in outdoor dynamic scenes that substantially outperforms existing methods in radar measurement synthesis and occupancy estimation accuracy, with particularly strong gains in dynamic outdoor environments.
- 8
- 8
- 8
- Rare bioparticle detection via deep metric learning8
A robust model based on a deep metric neural network for rare bioparticle (Cryptosporidium or Giardia) detection in drinking water is proposed and empowers imaging flow cytometry with capabilities of biomedical diagnosis, environmental monitoring, and other biosensing applications.
- R3L: Connecting Deep Reinforcement Learning To Recurrent Neural Networks For Image Denoising Via Residual Recovery8
It is demonstrated that the proposed R3L has better generalizability and robustness in image denoising when the estimated noise level varies, comparing to its counterparts using deterministic training, as well as various state-of-the-art image Denoising algorithms.
- Automating tephra fall building damage assessment using deep learning7
This is the first attempt to automate tephra fall building damage assessment solely using post-event data and is expected to perform well across other volcanic islands in the Caribbean where building types are similar, though it would benefit from additional testing.
- Video-Text Prompting for Weakly Supervised Spatio-Temporal Video Grounding7
This paper proposes Video-Text Prompting (VTP) to construct candidate feature of weakly-supervised Spatio-Temporal Video Grounding (STVG), and introduces negative contrastive samples whose candidate object is erased instead of being highlighted.
- 7
- 7
- 7
- 7
- 7
- 7
- 7
- Improving Flexible Image Tokenizers for Autoregressive Image Generation6
A flexible tokenizer with underline{Re}dundant \underline{Tok}en Padding and Hierarchical Semantic Regularization, designed to fully exploit all tokens for enhanced latent modeling, and introduces Redundant Token Padding to activate tail tokens more frequently, thereby alleviating information over-concentration in the early tokens.
- SoftShadow: Leveraging Soft Masks for Penumbra-Aware Shadow Removal6
This work introduces novel soft shadow masks specifically designed for shadow removal by leveraging the prior knowledge of pretrained SAM and integrating physical constraints, and proposes a SoftShadow framework that enables accurate predictions of penumbra and umbra areas while simultaneously facilitating end-to-end shadow removal.
- 6
- REPNP: Plug-and-Play with Deep Reinforcement Learning Prior for Robust Image Restoration6
A novel deep reinforcement learning (DRL) based PnP framework is proposed, dubbed RePNP, by leveraging a light-weight DRL-based denoiser for robust image restoration tasks and is robust to the observation model used in the PNP scheme deviating from the actual one.
- 5
- GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding5
This work proposes GroundFlow - a plug-in module for temporal reasoning on 3D point cloud sequential grounding, and introduces temporal reasoning capabilities to existing 3DVG models and achieves state-of-the-art performance in the SG3D benchmark across five datasets.
- Dual-head Genre-instance Transformer Network for Arbitrary Style Transfer5
This work proposes a Dual-head Genre-instance Transformer (DGiT) framework to simultaneously capture the genre and instance features for arbitrary style transfer and is the first work to integrate the genre features and instance features to generate a high-quality stylized image.
- Enhancing generalized spectral clustering with embedding Laplacian graph regularization5
An enhanced generalised spectral clustering framework that addresses the limitations of existing methods by incorporating the Laplacian graph and group effect into a regularisation term is presented and significantly enhances discrimination power and proves highly effective in handling noisy data.
- 5
- Removing Image Artifacts From Scratched Lens Protectors5
This work considers the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other, and demonstrates that the method outperforms the baselines qualitatively and quantitatively.
- 5
- 5
- The Power of Complementary Regularizers: Image Recovery via Transform Learning and Low-Rank Modeling5
- 5
- 5
- MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture Recognition4
A motion-aware state fusion mamba (MSF-Mamba), which enhances Mamba with local spatiotemporal modeling by fusing local contextual neighboring states, and a multiscale version named MSF-Mamba, which supports multiscale motion-aware state fusion, as well as an adaptive scale weighting module that dynamically weighs the fused states across different scales.
- StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation4
This work proposes StatePlay, a novel state-aware game world model that jointly predicts visual content and game states to promote mechanics-consistent generation and advances beyond pixel-level realism toward complete and mechanically faithful game generation.
- 4
- 4
- Leveraging Mixed Data Sources for Enhanced Road Segmentation in Synthetic Aperture Radar Images4
The results demonstrate that the HybridSAR Road Dataset and the adapted network significantly enhance the accuracy and robustness of SAR road segmentation, paving the way for future advancements in remote sensing.
- 4
- 4
- 4
- 4
- 4
- 4
- 3
- 3
- 3
- 3
- Digital Staining With Knowledge Distillation: A Unified Framework for Unpaired and Paired-but-Misaligned Data3
This work proposes a novel unsupervised deep learning framework for digital cell staining that reduces the need for extensive paired data using knowledge distillation, and applies this framework to the White Blood Cell dataset, investigating its potential for medical applications.
- Video Set Distillation: Information Diversification and Temporal Densification3
This work is the first to study Video Set Distillation, which synthesizes optimized video data by jointly addressing within-sample and inter-sample redundancies, and achieves state-of-the-art results in Video Dataset Distillation.
- 3
- 3
- 3
- 3
- 3
- 2
- On the Adversarial Vulnerabilities of Transfer Learning in Remote Sensing2
Adversarial neuron manipulation (ANM) is introduced, a novel attack strategy with two variants—ANM for single neuron (ANM-S) and ANM for multiple neurons (ANM-M)—that generate highly transferable perturbations by selectively targeting fragile neurons in pretrained models.
- 2
- 2
- 2
- Discovering Hidden Visual Concepts Beyond Linguistic Input in Infant Learning2
This work bridges cognitive science and computer vision by analyzing the internal representations of a computational model trained on an infant visual and linguistic inputs, and demonstrates that these neurons can recognize objects beyond the model’s original vocabulary.
- 2
- 2
- 2
- 2
- 1
- 1
- 1
- 1
- 1
- 1
- 1
- Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics1
This work proposes Dynamic-Aware Video Distillation (DAViD), a Reinforcement Learning (RL) approach to predict the optimal Temporal Resolution of the synthetic videos, and a teacher-in-the-loop reward function is proposed to update the RL agent policy.
- 1
- 1
- Unsupervised Deep Digital Staining for Microscopic Cell Images via Knowledge Distillation1
It is shown that the proposed unsupervised deep staining method can generate stained images with more accurate positions and shapes of the cell targets, and achieves much more promising results both qualitatively and quantitatively.
- 1
- Enhancing Low-Light Images Using Infrared Encoded Images1
A novel approach to increase the visibility of images captured under low-light environments by removing the in-camera infrared (IR) cut-off filter, which allows for the capture of more photons and results in improved signal-to-noise ratio due to the inclusion of information from the IR spectrum.
- ABCDE: An Agent-Based Cognitive Development Environment1
This work introduces ABCDE, an interactive 3D environment modeled after a typical playroom for children, the first environment aimed at mimicking a naturalistic setting for cognitive development in children; no other environment focuses on high-level concept learning through learner-teacher interactions.
- 1
- 1
- 1
- 1
- When Smart Signal Processing Meets Smart Imaging1
Some recent trends on techniques for high dynamic range (HDR) imaging, compressed sensing, computational imaging, as well as the image recovery methods with data-driven regularizers are covered.
- 1
- –
- –
- –
- –
- –
- –
- –
- Dynamic-Aware video distillation: Adaptive temporal partitioning based on video semantics for edge device–
This work proposes Adaptive Temporal Partitioning of synthetic videos to account for semantic adaptability and better temporal redundancy reduction, and is the first study to adaptively reduce temporal redundancy based on video semantics within the DD domain.
- –
- –
- –
- Filter Learning for Subgraphs: Algebras and Performance Risk Bounds–
A subgraph filter algebra based on distance-aware Laplacian constructions is developed, defining a structured and controllable class of filters for effective approximation and quantifying how well the learned operator approximates the restricted ambient mapping.
- Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models–
A dissociation between visual and linguistic attribute information is found: infant-trained models form strong visual representations for size and discriminate texture comparably to other models, but perform poorly on visual color discrimination, and in the text--vision setting they struggle to ground color and show only modest size grounding.
- –
- TPDiff: A Training-Free Triple-Path Diffusion Method for Content-Faithful Text-Driven Style Transfer–
- –
- –
- –
- –
- –
- –
- –
- –
- –
- Spectral Convergence of Complexon Shift Operators–
It is proved that when a simplicial complex sequence converges to a complexon, the eigenvalues of the corresponding CSOs converge to that of the limit complexon, which hint at learning transferability on large simplicial complexes or simplicial complex sequences, which generalize the graphon signal processing framework.
- –
- Editorial for the Special Issue on Advanced Machine Learning Techniques for Sensing and Imaging Applications–
This paper presents a meta-modelling architecture suitable for large-scale optimization of neural networks, and some examples show how this architecture can be applied to image recognition and 3D image analysis.
- –
- –
- –
- –
- MATLAB tools for EnviSAT ASAR data visualization and image enhancement–
A user-friendly ASAR data visualization and image enhancement toolbox based on the very popular MATLAB software environment is presented and a case study is presented with image enhancement for a poor AsAR data.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar; position from the scholar’s ORCID record, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.