Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
View as a treeStudents and postdocs1
- Chiat Pin TayFrom a thesis record ↗
Works205 from public data
- AANet: Attribute Attention Network for Person Re-Identifications365
The proposed AANet leverages on a baseline model that uses body parts and integrates the key attribute information in an unified learning framework and outperforms the best state-of-the-art method using ResNet-50.
- Multi-Path Region Mining for Weakly Supervised 3D Semantic Segmentation on Point Clouds165
This paper introduces a multi-path region mining module to generate pseudo point-level labels from a classification network trained with weak labels, and uses the point- level pseudo label to train a point cloud segmentation network in a fully supervised manner.
- 135
- 101
- 94
- A soft MAP framework for blind super-resolution image reconstruction87
A new soft maximum a posteriori (MAP) estimation framework to perform joint blur identification and HR image reconstruction that incorporates a soft blur prior that estimates the relevance of the best-fit parametric blur model, and induces reinforcement learning towards it.
- 87
- 54
- 53
- 53
- 53
- Interactive Change-Aware Transformer Network for Remote Sensing Image Change Captioning50
An Interactive Change-Aware Transformer Network (ICT-Net) is proposed, able to extract and incorporate the most critical changes of interest in each encoder layer to improve change description generation and achieves a state-of-the-art performance.
- Open World Object Detection: A Survey47
This survey paper offers a thorough review of the OWOD domain, covering essential aspects, including problem definitions, benchmark datasets, source codes, evaluation metrics, and a comparative study of existing methods.
- 46
- A recursive soft-decision approach to blind image deconvolution44
A new approach to blind image deconvolution based on soft-decision blur identification and hierarchical neural networks that integrates the knowledge of well-known blur models without compromising its flexibility in restoring images degraded by nonstandard blurs.
- 42
- TAPS3D: Text-Guided 3D Textured Shape Generation from Pseudo Supervision40
A novel framework, TAPS3D, is presented to train a text-guided 3D shape generator with pseudo captions based on rendered 2D images, which can generate 3D textured shapes from the given text without any additional optimization.
- 40
- 40
- 39
- Empirical Analysis Of Overfitting And Mode Drop In Gan Training39
It is shown that when stochasticity is removed from the training procedure, GANs can overfit and exhibit almost no mode drop, providing evidence against prevailing intuitions that GAns do not memorize the training set, and that mode dropping is mainly due to properties of the GAN objective rather than how it is optimized during training.
- 38
- 37
- 34
- 34
- 34
- 32
- 31
- Efficient discrete spatial techniques for blur support identification in blind image deconvolution31
Experimental results show that the proposed methods called maximum average square difference and maximum average absolute difference are effective in identifying the blur support reliably, thus providing a sound foundation for further blind image deconvolution.
- 29
- Vehicle license plate super-resolution using soft learning prior29
A new framework that adopts a soft learning approach in license plate super-resolution using optical character recognition (OCR) to perform VLP SR and results show that the proposed method is effective in handling license plate SR in both simulated and real experiments.
- A computational reinforced learning scheme to blind image deconvolution28
A novel reinforced mutation scheme is developed that combines stochastic search and pattern acquisition throughout the blur identification and is robust in alleviating the constraints and difficulties encountered by most conventional methods.
- 27
- 25
- 25
- 24
- 23
- 23
- 22
- 21
- 21
- 21
- Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and Method20
The BSCV is a collection of a proposed three-parameter corruption model for video bitstream, a large-scale dataset containing rich error patterns, multiple corruption levels, and flexible dataset branches, and a plug-and-play module in video recovery framework that serves as a benchmark.
- 20
- NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration: Methods and Results19
The NTIRE 2026 Challenge on Bitstream-Corrupted Video Restoration provides a common benchmark for evaluating restoration methods under realistic corruption settings and provides useful insights for future research on robust video restoration under practical bitstream corruption.
- Dense Supervision Propagation for Weakly Supervised Semantic Segmentation on 3D Point Clouds19
This paper proposes Dense Supervision Propagation (DSP) to train a semantic point cloud segmentation network with only a small portion of points being labeled and argues that this weakly supervised method with only 10% and 1% of labels can produce competitive results with the fully supervised counterpart.
- 19
- 18
- 17
- Toward Attribute-Controlled Fashion Image Captioning16
The results demonstrate that the proposed approach outperforms existing fashion image captioning models as well as conventional captioning methods and further validate the effectiveness of the proposed method on the MSCOCO and Flickr30K captioning datasets and achieve competitive performance.
- 16
- Bitstream-Corrupted JPEG Images are Restorable: Two-stage Compensation and Alignment Framework for Image Restoration15
A robust JPEG decoder is proposed, followed by a two-stage compensation and alignment framework to restore bitstream-corrupted JPEC images, and experimental results and ablation studies demonstrate the superiority of the proposed method.
- 15
- 15
- 14
- 14
- Content-based image retrieval using fuzzy perceptual feedback14
A fuzzy relevance feedback approach is proposed which enables the user to make a fuzzy judgement and integrates the user’s fuzzy interpretation of visual content into the notion of relevance feedback using a fuzzy radial basis function (FRBF) network.
- 13
- 13
- A Byte Sequence is Worth an Image: CNN for File Fragment Classification Using Bit Shift and n-Gram Embeddings13
This work proposes Byte2Image, a novel data augmentation technique, to introduce the neglected intra-byte information into file fragments and re-treat them as 2d gray-scale images, which allows us to capture both inter-byte and intra- byte correlations simultaneously through powerful convolutional neural networks (CNNs).
- Semantic granularity metric learning for visual search13
A new deep semantic granularity metric learning (SGML) is proposed that develops a novel idea of leveraging attribute semantic space to capture different granularity of similarity, and then integrates this information into deep metric learning.
- Efficient mobile landmark recognition based on saliency-aware scalable vocabulary tree13
This paper constructed a city-scale landmark dataset in Singapore and the experimental results show that the proposed mobile landmark recognition by incorporating saliency information outperforms the baseline SVT recognition by about 9%.
- 13
- 12
- 12
- Multi-Modality Action Recognition Based on Dual Feature Shift in Vehicle Cabin Monitoring12
A novel yet efficient multi-modality driver action recognition method based on dual feature shift based on dual feature shift, named DFS, which achieves good performance and improves the efficiency of multi-modality driver action recognition.
- SSN: Stockwell Scattering Network for SAR Image Change Detection12
Stockwell scattering network (SSN) based on Stockwell transform (ST) is proposed, which provides noise-resilient feature representation and obtains state-of-the-art performance in SAR image change detection as well as high computational efficiency.
- 12
- 12
- OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian Tracking11
This work proposes an adaptive occlusion-aware multiple pedestrian tracker, OccluTrack, that outperforms state-of-the-art methods on MOTChallenge and DanceTrack datasets and develops a pose-guided re-identification module to extract discriminative part features for partially occluded pedestrians.
- 11
- Remote detection of idling cars using infrared imaging and deep networks11
The first automatic system to detect idling cars, using infrared (IR) imaging and deep networks, is proposed, based on the differences in spatio-temporal heat signatures of idling and stopped cars and monitor the car temperature with a long-wavelength IR camera.
- 11
- 11
- 11
- 11
- PromptSR: Cascade Prompting for Lightweight Image Super-Resolution10
The experimental results demonstrate the superiority of the PromptSR method, which outperforms state-of-the-art lightweight SR methods in quantitative, qualitative, and complexity evaluations.
- 10
- 10
- Adaptive image restoration based on hierarchical neural networks10
Experimental results show that the new approach to adaptive image regularization based on a neural network, hierarchical cluster model, is superior in suppressing noise and ringing at the smooth background while effectively preserving the fine details at the texture and edge regions.
- High dimensional optical data — varifocal multiview imaging, compression and evaluation9
An efficient VFMV compression scheme based on view mountain-shape rearrangement (VMSR) and all-directional prediction structure (ADPS) that outperforms comparison schemes by quantitative, qualitative, complexity, and forgery protection evaluations is proposed.
- 9
- 9
- Multifocal multiview imaging and data compression based on angular–focal–spatial representation8
To efficiently compress MFMV data, the first, to the authors' knowledge, MFMV data compression scheme based on angular-focal-spatial representation is proposed, which exploits inter-view, inter-stack, and intra-frame predictions to eliminate data redundancy in angular, focal, and spatial dimensions, respectively.
- Collaborative learning mutual network for domain adaptation in person re-identification8
This paper proposes a new Collaborative Learning Mutual Network (CLM-Net) for domain adaptation in person re-identification (re-id) that integrates body part features learning tasks and a global saliency task to the baseline model so that information that helps to identify the pedestrian can be extracted.
- An efficient approach for scene categorization based on discriminative codebook learning in bag-of-words framework8
This paper proposes an efficient technique for learning a discriminative codebook for scene categorization by careful design of the codewords such that the resulting image histograms for each category will retain strong power, while the online categorization of the testing image is as efficient as in the baseline BoW.
- A Perceptual Subjectivity Notion in Interactive Content-Based Image Retrieval Systems8
This chapter presents a new framework called fuzzy relevance feedback in interactive content-based image retrieval (CBIR) systems that is developed using a fuzzy radial basis function network (FRBFN) based on hierarchical clustering algorithm.
- Efficient Recursive Multichannel Blind Image Restoration8
A novel multichannel recursive filtering technique that incorporates a forgetting factor to discard the old unreliable estimates, hence achieving better convergence performance and allowing the method to be adopted readily in real-life applications.
- 8
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence7
This work introduces S-Agent, a spatial tool-use agentic paradigm for understanding and reasoning over continuous multi-view images and videos, and consistently improves both open-source and closed-source VLMs in a training-free manner.
- A Structure-Aware and Motion-Adaptive Framework for 3D Human Pose Estimation with Mamba7
This work proposes a structure-aware and motion-adaptive framework to capture spatial joint topology along with diverse motion dynamics independently, named as SAMA, which enables structure-aware and motion-adaptive pose lifting.
- 7
- 7
- 7
- A Collaborative Bayesian Image Annotation Framework7
Numerical results based on images collected from the Internet demonstrated better performance resulting from the introduction of context knowledge and information fusion, resulting in a new framework, termed as Collaborative Bayesian Image Annotation (CBIA) framework.
- SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review6
Experiments show that agentic review continuously improves PRs through a generate-review-revise loop, outperforms single-turn fixed-context review in both decision accuracy and resolve rate after revision, transfers beyond review to improve issue-resolution models, and enables effective and efficient test-time scaling.
- Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges6
This paper identifies three core tasks in EgoProceAssist: egocentric procedural error detection, egocentric procedural learning, and egocentric procedural question answering, and introduces two enabling dimensions: real-time and streaming video understanding, and proactive interaction in procedural contexts.
- Towards Blind Bitstream-corrupted Video Recovery: A Visual Foundation Model-driven Framework6
This paper proposes the first blind bitstream-corrupted video recovery framework that integrates visual foundation models with recovery model, which is adapted to different types of corruption and bitstream-level prompts and introduces a novel Corruption-aware Feature Completion (CFC) module.
- MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition6
A novel multimodal fusion transformer, named Multi-Fuser, which identifies cross-modal interrelations and interactions among multimodal car cabin videos and adaptively integrates different modalities for improved representations is proposed.
- 6
- Tiered Deep Similarity Search for Fashion6
A new attribute-guided metric learning (AGML) with multitask CNNs that jointly learns fashion attributes and image embeddings while taking category and brand information into account is proposed.
- 6
- 6
- 6
- Adaptive Image Processing6
This book discusses four main approaches to image restoration: Blind Image Deconvolution, Computational Reinforced Learning Soft-Decision Method Simulation Examples, Evolutionary Computation, and Content-Based Image Retrieval.
- Image Denoising Using Stochastic Chaotic Simulated Annealing6
A noisy chaotic neural network is used, which adds noise and chaos into the Hopfield neural network to facilitate efficient searching and to avoid local minima in a novel optimization algorithm called stochastic chaotic simulated annealing.
- 6
- 5
- 5
- 5
- Beyond bag of words: Combining generative and discriminative models for natural scene categorization5
- 5
- 5
- 5
- 4
- 4
- 4
- Attribute saliency network for person re-identification4
The Attribute Saliency Network (ASNet) is proposed, a deep learning model that utilizes attribute and saliency map learning for person re-identification (re-ID) task and improves the granularity of the heatmap by generating two global person attributes and body part saliency maps.
- 4
- 4
- 4
- 4
- 4
- 4
- 4
- 4
- 4
- 4
- 3
- 3
- 3
- 3
- 3
- 3
- 3
- 3
- 3
- 3
- Context-aware mobile image annotation for media search and sharing3
This paper proposes a new mobile image annotation system that utilizes content analysis, context analysis and their integration to annotate images acquired from mobile devices and shows good potential in domain-specific mobile image annotations for image sharing.
- 3
- New regularization scheme for blind color image deconvolution3
A reinforcement regularization framework that integrates a soft parametric learning term in addressing blind color image deconvolution and is able to achieve satisfactory restored color images under different blurring conditions is proposed.
- Soft-Labeling Image Scheme Using Fuzzy Support Vector Machine3
A soft labeling framework is developed that strives to address the small sample problem in CBIR systems by studying the characteristics of labeled images and utilizing an unsupervised clustering algorithm to select unlabeled images, which are called soft-labeled images.
- Knowledge Propagation in Collaborative Tagging for Image Retrieval3
A new knowledge propagation scheme to automatically propagate keywords from a subset of annotated images to the unannotated ones is proposed, based on image content analysis and training of keyword classifiers.
- 3
- 3
- Demystifying Data Organization for Enhanced LLM Training2
This paper systematically explores the influence of data organization on LLM training by reusing pre-computed sample-level scores originally generated for data efficiency, thereby incurring minimal additional computational overhead.
- Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep2
HetCache is introduced, a training-free diffusion acceleration framework designed to exploit the inherent heterogeneity in diffusion-based masked video-to-video (MV2V) generation and editing, which reduces redundant attention operations while maintaining editing consistency and fidelity.
- 2
- 2
- CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models2
A Cross-Level HOI distillation (CL-HOI) framework is proposed, which distills instance-level HOIs from VLLMs image-level understanding without the need for manual annotations, and demonstrates its efficacy in detecting HOIs without manual labels.
- 2
- Beyond Bag-of-Words: combining generative and discriminative models for scene categorization2
The image signatures for training discriminative model are carefully designed based on the generative model, and the soft relevance value of the extracted image signatures are estimated by image signature space modeling and incorporated in Fuzzy Support Vector Machine (FSVM).
- 2
- 2
- 2
- 2
- 2
- 2
- 2
- Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition1
The experimental results on the public benchmark show that the proposed MoME framework and the HTL strategy jointly outperform representative single-modal and multimodal baselines and improves subtle multimodal understanding and offers better interpretability.
- 1
- 1
- 1
- Video sentence grounding with temporally global textual knowledge1
A Pseudo-query Intermediary Network (PIN) is proposed to achieve an improved alignment of visual and comprehensive pseudo-query features within the feature space through contrastive learning and enhances the feature alignment between visual and language for better temporal grounding.
- Learning-Based Biharmonic Augmentation for Point Cloud Classification1
Biharmonic Augmentation is a novel and efficient data augmentation technique that diversifies point cloud data by imposing smooth non-rigid deformations on existing 3D structures and AdvTune, an advanced online augmentation system that integrates adversarial training is presented.
- 1
- 1
- Lattice-Support repetitive local feature detection for visual search1
Experiments performed on benchmark datasets show that the proposed LS-RLF detection method outperforms the state-of-the-art methods by mean Average Precisions (mAP) of 4.5%, 5.5% and 3.2% on Oxford, Paris, and INRIA holidays datasets respectively.
- 1
- 1
- 1
- 1
- 1
- Adaptive resynchronization approach for scalable video over wireless channel1
It is shown from experimental results that the proposed adaptive resynchronization method can perform a graceful degradation under a variety of error conditions and shows advantages over conventional method.
- 1
- 1
- 1
- 1
- 1
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- –
- Guest Editorial Special Issue on Advances in Multimedia Computing, Communications and Applications–
Ten papers published in this special issue cover a wide range of techniques and applications in various multimedia processing tasks, including multimedia analysis and retrieval, compression and optimization, communication and networking, and multimedia system and applications.
- –
- –
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar; position from the scholar’s ORCID record, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.