Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks42 from public data
- A Survey of Embodied AI: From Simulators to Research Tasks597
An encyclopedic survey of the three main research tasks in embodied AI – visual exploration, visual navigation and embodied question answering – covering the state-of-the-art approaches, evaluation metrics and datasets is surveyed.
- 98
- 88
- 54
- 52
- A Survey and Evaluation of Adversarial Attacks in Object Detection32
This article presents a novel taxonomic framework for categorizing adversarial attacks specific to object detection architectures, synthesizes existing robustness metrics, and provides a comprehensive empirical evaluation of state-of-the-art attack methodologies on popular object detection models, including both traditional detectors and modern detectors with vision-language pretraining.
- 28
- 27
- Approaching expert-level accuracy for differentiating ACL tear types on MRI with deep learning22
The fully automated model shows potential as a highly reliable and reproducible tool that allows orthopedists to noninvasively identify the ACL status and may aid in optimizing different techniques, such as ACL remnant preservation, for ACL reconstruction.
- 12
- Industrial CT image reconstruction for faster scanning through U-Net++ with hybrid attention and loss function11
A prunable U-NET++ with hybrid attention and loss function (HAL-UNET++) deep learning network, which is a U-shaped framework-based network for industrial CT image reconstruction, achieves rapid image reconstruction.
- 10
- TRECVID 2010 Known-item Search (KIS) Task by I2R.9
This work postulates that searchers can quickly reject most of the negative videos after seeing a few keyframes of the video, and develops an intuitive and user-friendly user interface to facilitate the browsing of returned videos in automatic KIS.
- 8
- 7
- A comprehensive survey of procedural video datasets7
This survey examines the current state of procedural video datasets, in terms of their data, content and annotation characteristics, as well as processing function and evaluation.
- 5
- Spectral-spatial hyperspectral image classification using super-pixel-based spatial pyramid representation5
This work proposes a novel spectral-spatial hyperspectral image classification approach using superpixel-based spatial pyramid representation, which is a complementary yet effective feature pooling approach, and evaluation on two public hyperspectrals with superior image classification performance.
- Multipath sparse coding for scene classification in very high resolution satellite imagery5
Experimental results show that the proposed unsupervised feature learning approach for scene classification on very high resolution satellite imagery outperforms the state-of-the-art that uses the single-layer sparse coding.
- 4
- Benchmarking Visual Generative Models through Cultural Lens: A Case Study with Singapore-Centric Multi-Cultural Context4
This case study introduces and utilizes a Singapore-centric multi-cultural image dataset that reflects rich ethnic, religious, and social diversity and evaluates the performance of nine state-of-the-art generative models, including DALL-E 3 and SD3.5, on their ability to contextually represent multi-cultural diversity.
- 4
- 4
- 4
- Contextualized Visual Storytelling for Conversational Chatbot in Education3
A contextualized dense image captioning framework is investigated, which augments dense image captioning with cultural and curriculum-aligned keyword retrieval through a Retrieval-Augmented Generation (RAG) module, which enables the generation of culturally appropriate, age-level suitable, and educationally anchored captions that enhance learner engagement and pedagogical relevance.
- 2
- 2
- L-shaped corner detector for rooftop extraction from satellite/aerial imagery2
This work proposes an L-shaped corner detector for automatic rooftop extraction from high resolution satellite/aerial imagery that considers information in a spatial circle around each pixel to construct a feature map which captures the probability of L- shaped corner at every pixel.
- Feasibility of replacing true non-contrast images with virtual non-contrast images in quantitative analysis of emphysema1
The use of VNC, particularly VNC(VP) images, in chest CT has the potential to supplant TNC for the quantitative assessment of emphysema, thereby streamlining scans and reducing radiation dose.
- 1
- 1
- 1
- 1
- Regions-of-interest extraction from remote sensing imageries using visual attention modelling1
The proposed visual attention model incorporates both bottom-up spatial saliency and top-down objectness, by fusing a co-occurrence histogram saliency model with the BING objectness model, and shows that the proposed model can effectively and efficiently identify regions-of-interest in remote sensing data.
- –
- Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric–
A novel framework to evaluate the vision-language logical consistency of MLLMs on both sufficient and necessary cause-effect relations is proposed and it is suggested that, beyond accuracy, logical consistency could be employed for both accuracy and reliability.
- –
- –
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it.