Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks17 from public data
- 57
- 22
- 17
- Combined CNN Transformer Encoder for Enhanced Fine-grained Human Action Recognition12
This work investigates two frameworks that combine CNN vision backbone and Transformer Encoder to enhance fine-grained action recognition, and shows that both Transformer encoder frameworks effectively learn latent temporal semantics and cross-modality association, with improved recognition performance over CNN vision model.
- 10
- Artificial intelligence-enhanced 3D gait analysis with a single consumer-grade camera9
3DGait is introduced, an artificial intelligence-enhanced markerless 3-Dimensional gait analysis system that operates with a single consumer-grade depth camera, providing a streamlined, accessible alternative to marker-based motion capture systems.
- 7
- Identification of faces in line drawings by edge decomposition5
A method to find the faces of objects in 2D line drawings, using an approach totally different from existing ones that depends only on the topology of the drawing and not on the geometry, and is therefore applicable to line drawings with curves and straight lines, and 3D wireframes.
- Contextualized Visual Storytelling for Conversational Chatbot in Education3
A contextualized dense image captioning framework is investigated, which augments dense image captioning with cultural and curriculum-aligned keyword retrieval through a Retrieval-Augmented Generation (RAG) module, which enables the generation of culturally appropriate, age-level suitable, and educationally anchored captions that enhance learner engagement and pedagogical relevance.
- 2
- 1
- –
- –
- Towards Annotation-Free Validation of MLLMs: A Vision-Language Logical Consistency Metric–
A novel framework to evaluate the vision-language logical consistency of MLLMs on both sufficient and necessary cause-effect relations is proposed and it is suggested that, beyond accuracy, logical consistency could be employed for both accuracy and reliability.
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.