Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks48 from public data
- A novel hybrid approach for crack detection174
A novel hybrid approach for crack detection in raw images, which combines deep learning models and Bayesian probabilistic analysis for robust crack detection is proposed, which outperforms the state-of-the-art baseline approach on deep CNN classifier.
- 91
- 3D reconstruction of polyhedral objects from single parallel projections using cubic corner31
A direct method to recover the geometry of the 3D polyhedron depicted in a single parallel projection using two sets of information, the list of faces in the object, obtained automatically from the drawing, and a user-identified cubic corner, to compute for the coordinates of the vertices in the drawing.
- 28
- 22
- 19
- A new hybrid method for 3D object recovery from 2D drawings and its validation against the cubic corner method and the optimisation-based method19
The hybrid method is used to reconstruct 3D polyhedral objects from 2D line drawings by combining two known methods, the cubic corner method and the optimisation-based method, and presents comprehensive test results comparing the three methods.
- 14
- 3D reconstruction of polyhedral objects from single perspective projections using cubic corner12
A direct method to recover the geometry of the 3D planar polyhedron depicted in a single perspective projection first finds the vanishing points and then the focal length, which establishes the linear relationship between a vertex on the image plane and its corresponding point on the object.
- 11
- 10
- 8
- 8
- 8
- Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs7
Positive Attention Dynamics Enhancement is proposed, a training-free attention intervention that constructs a PAD map to identify semantically core visual regions, applies per-head Median Absolute Deviation Scaling to adaptively control the intervention strength, and leverages System-Token Compensation to maintain attention to complex user instructions and support long-term output consistency.
- 7
- Lifelog Image Retrieval Based on Semantic Relevance Mapping7
This article proposes a novel semantic relevance mapping (SRM) method that serves both as a formalism to construct a trainable model to bridge the semantic gap and an algorithm to implement the training process on real-world lifelog data.
- 7
- Masked Diffusion with Task-awareness for Procedure Planning in Instructional Videos6
A simple yet effective enhancement - a masked diffusion model that acts akin to a task-oriented attention filter, enabling the diffusion/denoising process to concentrate on a subset of action types.
- 6
- Efficient decomposition of line drawings of connected manifolds without face identification6
An algorithm for decomposing complex line drawings which depict connected 3D manifolds into multiple simpler drawings of individual manifolds is presented, which greatly improves the process for 3D reconstruction, which is the ultimate goal.
- Self-Teaching Strategy for Learning to Recognize Novel Objects in Collaborative Robots5
A self-teaching strategy for a cobot to learn to recognize novel objects efficiently and effectively is proposed, like human-to-human teaching, where the user just provides a few examples of a novel object captured by an RGB-D camera.
- Identification of faces in line drawings by edge decomposition5
A method to find the faces of objects in 2D line drawings, using an approach totally different from existing ones that depends only on the topology of the drawing and not on the geometry, and is therefore applicable to line drawings with curves and straight lines, and 3D wireframes.
- Localizing discriminative regions for fine-grained visual recognition: One could be better than many4
- 4
- 4
- 4
- Team VI-I2R Technical Report on EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 20213
This report presents the technical details of the approach to the EPIC-KITCHENS-100 Unsupervised Domain Adaptation (UDA) Challenge for Action Recognition, and achieves the 1st place in terms of top-1 action recognition accuracy, using only RGB and optical flow modalities as input.
- 3D reconstruction from drawings with straight and curved edges3
The results of the implementation show that the recovered objects correspond to the human perception of what they should be, however, work remains in producing a measure on the goodness of the result and providing handles to allow the control of the final outcome.
- Visuo-Tactile Manipulation Planning Using Reinforcement Learning with Affordance Representation2
A reinforcement learning-based motion planning framework for object manipulation which makes use of both on-the-fly multisensory feedback and a learned attention-guided deep affordance model as perceptual states is proposed.
- 2
- Towards a Programming-Free Robotic System for Assembly Tasks Using Intuitive Interactions2
A solution which enables a robot to learn new objects and new tasks from non-expert users without the need for programming, and three main modules enable any non-expert user to configure a robot for new applications in a fast and intuitive way.
- 1
- Next-Generation Metalens Vision System: Powered by AI and Applied to AI1
An end-to-end metalens vision system is presented—from hardware sensing with a custom-built RGB metalens camera, to physics-informed imaging and real-time restoration, and finally to downstream vision applications such as object detection and depth estimation.
- A Study on Differentiable Logic and LLMs for EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 20231
The innovative application of a differentiable logic loss in the training to leverage the co-occurrence relations between verb and noun, as well as the pre-trained Large Language Models to generate the logic rules for the adaptation to unseen action labels achieves the first place in terms of top-1 action recognition accuracy.
- Team VI-I2R Technical Report on EPIC-KITCHENS-100 Unsupervised Domain Adaptation Challenge for Action Recognition 20221
The technical details of the submission to the EPIC-KITCHENS-100 Unsupervised Domain Adaptation (UDA) Challenge for Action Recognition 2022 are presented and an action-aware domain adaptation framework that leverages the prior knowledge induced from the action recognition task during the adaptation is proposed.
- 1
- 1
- 1
- Editing Is a Bargaining Game: Balanced Knowledge Editing in Large Language Models–
This work proposes a balanced knowledge editing framework inspired by Nash bargaining theory, and guides the optimization process toward a Pareto stationary point, ensuring an equilibrium solution wherein any deviation from the final state would degrade the overall performance with respect to both objectives.
- Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles–
This study proposes an instruction-realization framework that leverages a large language model (LLM) to interpret instructions, generates executable scripts that schedule multiple model predictive control (MPC)-based motion planners based on real-time feedback, and converts planned trajectories into control signals.
- –
- –
- –
- –
- Team I2R-VI-FF Technical Report on EPIC-KITCHENS VISOR Hand Object Segmentation Challenge 2023–
This report presents the approach to the EPIC-KITCHENS VISOR Hand Object Segmentation Challenge, which focuses on the estimation of the relation between the hands and the objects given a single frame as input, and achieves the 1st place in terms of evaluation criteria in the VISOR HOS Challenge.
- Enhance the Efficacy of Deep CNN with Auxiliary Labels–
A new algorithm is derived to iteratively train the deep CNN model on target and auxiliary labels on the principle of minimizing cross entropy loss, which demonstrates that by introducing auxiliary attributes to training images, uncertainty can be reduced in target classification tasks, and adversarial effects avoided in multi-task formulation.
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.