Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks33 from public data
- A Review of Target Recognition Technology for Fruit Picking Robots: From Digital Image Processing to Deep Learning68
This paper systematically summarizes the research work on target recognition techniques for picking robots in recent years, analyzes the technical characteristics of different approaches, and concludes their development history.
- 66
- Learning contrastive feature distribution model for interaction recognition56
Comparing the proposed CFDM approach on CR-UESTC and SBU interaction databases, and comparing the result of CFDM with the CM and the BoW approach indicates that the recognition accuracy of three approaches is: CFDMCMBoW.
- Action Knowledge Transfer for Action Prediction with Partial Videos42
The proposed action knowledge transfer method can significantly improve the performance of action prediction, especially for the actions with small observation ratios (e.g., 10%).
- 37
- Mitigating and Evaluating Static Bias of Action Representations in the Background and the Foreground35
This paper empirically verify the existence of foreground static bias by creating test videos with conflicting signals from the static and moving portions of the video, and proposes a simple yet effective technique, StillMix, to learn robust action representations.
- Deep Dual Relation Modeling for Egocentric Interaction Recognition33
A novel interactive LSTM module is developed to explicitly model the relations between the two interacting persons based on their individual action representations, which are collaboratively learned with an interactor attention module and a global-local motion module.
- Unsupervised Learning for Optical Flow Estimation Using Pyramid Convolution LSTM23
This paper proposes an unsupervised optical flow estimation framework named PCLNet, which uses pyramid Convolution LSTM (ConvLSTM) with the constraint of adjacent frame reconstruction, which allows flexibly estimating multi-frame optical flows from any video clip.
- Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data18
SPARCL (Synthetic Perturbations for Advancing Robust Compositional Learning), which integrates image feature injection into a fast text-to-image generative model, followed by an image style transfer step, to meet the three challenges of synthesizing training images for compositional learning.
- Segmentation Network for Multi-Shape Tea Bud Leaves Based on Attention and Path Feature Aggregation9
The proposed Unet-Enhanced model achieves segmentation performance well on one bud–one leaf targets with different shapes, with a mean intersection over union (mIoU) of 91.18% and a mean pixel accuracy (mPA) of 95.10%.
- 8
- 7
- Facial micro-expression recognition based on motion magnification network and graph attention mechanism7
This paper introduces a Graph Attention Mechanism-based Motion Magnification Guided Micro-Expression Recognition Network (GAM-MM-MER), which optimizes facial key area maps and prioritizes adjacent nodes crucial for mitigating the influence of noisy neighbors, while attending to key feature information.
- 6
- Adaptive Interaction Modeling via Graph Operations Search6
This paper proposes to search the network structures with differentiable architecture search mechanism, which learns to construct adaptive structures for different videos to facilitate adaptive interaction modeling, and experimentally demonstrates that the designed basic graph operations in the search space are able to model different interactions in videos.
- Hierarchical topology based hand pose estimation from a single depth image5
A hierarchical topology based approach to estimate 3D hand poses on the hand skeleton topology model, which ensures estimated hand poses in a reasonable topology, and improves estimation accuracy.
- 3
- Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction3
Mutual Information-aware Multimodal Iterated Relational dAta GEneration (MI2RAGE), which applies Chained Cross-modal Generation (CCG) to promote diversity in the generated data and exploits a teacher network to select valuable training samples with high mutual information with the ground-truth labels is proposed.
- 2
- Epidemiological characteristics of neuroendocrine neoplasms in Beijing: a population-based retrospective study2
It is found that NENs originating from the lung had worse overall survival than extrapulmonary NENs, and male patients had worse survival than female patients, and male patients had worse survival than female patients.
- Learning to Animate Images from A Few Videos to Portray Delicate Human Actions1
FLASH (Few-shot Learning to Animate and Steer Humans), which enhances generalization of motion by training the model to reconstruct a video using the motion features and cross-frame correspondences extracted from another video with the same motion but different appearance, is proposed.
- Gauze detection and segmentation in laparoscopic liver surgery: a multi-center study1
The findings highlight the strong capability of the proposed framework to detect and segment gauze across varying levels of background complexity in laparoscopic liver surgery videos and have the potential to significantly improve gauze management during surgery.
- Weakly supervised action anticipation without object annotations1
A weakly supervised method is developed that integrates global motion and local finegrained features from current action videos to predict next action label without the need for specific scene context labels and constructs a graph convolutional network for exploiting the inherent relationships of humans and objects under present incidents.
- Omni-directional guide intelligent follow-up walker design1
The all-round intelligent follow-up walker can meet the needs of users for daily community travel, walking assistance and storage, has a high degree of integration of existing technologies and has a wide range of application prospects.
- The Comparison of the computing ability of quantum and conventional computer1
The paper shows that the current technology cannot perform the perfect quantum computer, but the future of it would be very promising, and shed light on guiding further exploration of quantum computing.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.