Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Academic lineage

View as a tree

Students and postdocs9

Show 1 more

Possible advisorsa guess from early papers, not confirmed

  • L. van Gool

    Possible advisor · last author on 5 of their early first-author papers, 2010–2012

    Suggested from co-authorship

Is this you? Claim this profile to confirm or dismiss it.

Works166 from public data

TitleCited by
  • NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions

    Junbin Xiao, Xin-Di Shang, Angela Yao, Tat-seng Chua

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2021

    NExT-QA is introduced, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions, and it is found that top-performing methods excel at shallow scene descriptions but are weak in causal and temporal action reasoning.

    1,033
  • Hough Forests for Object Detection, Tracking, and Action Recognition

    Juergen Gall, Angela Yao, Nima Razavi, L. van Gool, V. Lempitsky

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2011

    653
  • Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities

    F. Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    The first multi-view action dataset, with si-multaneous static and egocentric recordings, and a novel task of detecting mistakes is proposed, to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance.

    432
  • A Hough transform-based voting framework for action recognition

    Angela Yao, Juergen Gall, L. van Gool

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2010

    301
  • A Two-Streamed Network for Estimating Fine-Scaled Depth Maps from Single RGB Images

    Jun Li, Can Yuce, Reinhard Klein, Angela Yao

    IEEE International Conference on Computer Vision (ICCV) · 2017

    276
  • Does Human Action Recognition Benefit from Pose Estimation?

    Angela Yao, Juergen Gall, G. Fanelli, L. van Gool

    British Machine Vision Conference (BMVC) · 2011

    Comparing pose-based, appearance-based and combined pose and appearance features for action recognition in a home-monitoring scenario shows that posebased features outperform low-level appearance features, even when heavily corrupted by noise, suggesting that pose estimation is beneficial for the action recognition task.

    200
  • Dense 3D Regression for Hand Pose Estimation

    Cheng-De Wan, Thomas Probst, L. van Gool, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2018

    A simple and effective method for 3D hand pose estimation from a single depth frame based on dense pixel-wise estimation that outperforms all previous state-of-the-art approaches by a large margin and outperforms various other proposed methods.

    198
  • Temporal Aggregate Representations for Long-Range Video Understanding

    F. Sener, Dipika Singhania, Angela Yao

    Lecture notes in computer science · 2020

    This work addresses questions of temporal extent, scaling, and level of semantic abstraction with a flexible multi-granular temporal aggregation framework and shows that it is possible to achieve state of the art in both next action and dense anticipation with simple techniques such as max-pooling and attention.

    185
  • Can I Trust Your Answer? Visually Grounded Video Question Answering

    Junbin Xiao, Angela Yao, Yi-Cong Li, Tat-Seng Chua

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    NExT-GQA is constructed - an extension of NExT-QA with 10.5K temporal grounding (or location) labels tied to the original QA pairs tied to the original VLMs and aims to push towards trustworthy VLMs in VQA systems.

    183
  • Temporal Action Segmentation: An Analysis of Modern Techniques

    Guodong Ding, F. Sener, Angela Yao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2023

    This survey analyzes and summarizes the most significant contributions and trends in temporal action segmentation in videos, and systematically investigates two essential techniques of this topic, i.e., frame representation and temporal modeling.

    165
  • Crossing Nets: Combining GANs and VAEs with a Shared Latent Space for Hand Pose Estimation

    Cheng-De Wan, Thomas Probst, L. van Gool, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2017

    This work proposes modelling the statistical relationship of 3D hand poses and corresponding depth images using two deep generative models with a shared latent space to prevent over-fitting and to better exploit unlabeled depth maps.

    159
  • Video as Conditional Graph Hierarchy for Multi-Granular Question Answering

    Junbin Xiao, Angela Yao, Zhiyuan Liu, Yi-Cong Li, Wei Ji, Tat-seng Chua

    Proceedings of the AAAI Conference on Artificial Intelligence · 2022

    This work proposes to model video as a conditional graph hierarchy which weaves together visual facts of different granularity in a level-wise manner, with the guidance of corresponding textual cues to align with the multi-granular essence of linguistic concepts in language queries.

    154
  • Unsupervised Learning and Segmentation of Complex Activities from Video

    F. Sener, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2018

    This paper proposes an iterative discriminative-generative approach which alternates between discriminatively learning the appearance of sub-activities from the videos' visual features to sub-activity labels and generatively modelling the temporal structure of sub theactivities using a Generalized Mallows Model.

    128
  • Disentangling Latent Hands for Image Synthesis and Pose Estimation

    Linlin Yang, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2019

    Experiments show that the dVAE can synthesize highly realistic images of the hand specifiable by both pose and image background content and also estimate 3D hand poses from RGB images with accuracy competitive with state-of-the-art on two public benchmarks.

    127
  • Coupled Action Recognition and Pose Estimation from Multiple Views

    Angela Yao, Juergen Gall, L. van Gool

    International Journal of Computer Vision · 2012

    A framework for coupled action recognition and pose estimation is presented by formulating pose estimation as an optimization over a set of action-specific manifolds to demonstrate not only the feasibility of using extracted 3D poses for action recognition, but also improved performance in comparison to action recognition using low-level appearance features.

    118
  • Self-Supervised 3D Hand Pose Estimation Through Training by Fitting

    Cheng-De Wan, Thomas Probst, L. van Gool, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2019

    117
  • 108
  • Interactive object detection

    Angela Yao, Juergen Gall, Christian Leistner, L. van Gool

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2012

    103
  • Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement

    Kai Xu, Rongyu Chen, Gianni Franchi, Angela Yao

    arXiv · 2023

    It is demonstrated that activation pruning has a detrimental effect on OOD detection, while activation scaling enhances it, and a simple yet effective post-hoc network enhancement method is proposed, SCALE, which attains state-of-the-art Ood detection performance without compromising in-distribution (ID) accuracy.

    102
  • Hand Pose Estimation from Local Surface Normals

    Cheng-De Wan, Angela Yao, L. van Gool

    Lecture notes in computer science · 2016

    A hierarchical regression framework for estimating hand joint positions from single depth images based on local surface normals and a conditional regression forest, i.e. the Frame Conditioned Regression Forest (FCRF) which uses a new normal difference feature.

    95
  • Deep morphological networks

    G. Franchi, Amin Fehri, Angela Yao

    Pattern Recognition · 2020

    This paper demonstrates on various examples that new layers making use of the morphological non-linearities are complementary to convolution layers and can be used to integrate the non- linear operations and pooling into a joint operation.

    93
  • Aligning Latent Spaces for 3D Hand Pose Estimation

    Linlin Yang, Shile Li, Dongheui Lee, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2019

    93
  • Content-Aware Multi-Level Guidance for Interactive Instance Segmentation

    Soumajit Majumder, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2019

    91
  • Contrastive Video Question Answering via Video Graph Transformer

    Junbin Xiao, Pan Zhou, Angela Yao, Yi-Cong Li, Richang Hong, Shuicheng Yan, Tat-seng Chua

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2023

    With superior video encoding and QA solution, it is shown that CoVGT can achieve much better performances than previous arts on video reasoning tasks and can also benefit from cross-modal pretraining, yet with orders of magnitude smaller data.

    78
  • Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly Detection

    Chengwei Chen, Yuan Xie, Shaohui Lin, Angela Yao, Guannan Jiang, Wei Zhang, Yan-Yun Qu, Ruizhi Qiao, +2 more

    Proceedings of the AAAI Conference on Artificial Intelligence · 2022

    A novel bi-directional architecture with three consistency constraints to comprehensively regularize the prediction task from pixel-wise, cross-modal, and temporal-sequence levels to outperforms advanced anomaly detectors and achieves state-of-the-art results.

    78
  • Efficient Unsupervised Temporal Segmentation of Motion Data

    Björn Krüger, Anna Vögele, T. Willig, Angela Yao, Reinhard Klein, Andreas Weber

    IEEE Transactions on Multimedia · 2016

    A method for automated temporal segmentation of human motion data into distinct actions and compositing motion primitives based on self-similar structures in the motion sequence is introduced, which requires no assumptions about the motion sequences at hand and no user interaction for the segmentation or clustering.

    78
  • Zero-Shot Anticipation for Instructional Activities

    F. Sener, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2019

    A hierarchical model is presented that generalizes instructional knowledge from large-scale text-corpora and transfers the knowledge to the visual domain and predicts coherent and plausible actions multiple steps into the future, all in rich natural language.

    75
  • Variations of a Hough-Voting Action Recognition System

    D. Waltisberg, Angela Yao, Juergen Gall, L. van Gool

    Lecture notes in computer science · 2010

    Two variations of a Hough-voting framework for action recognition for group actions with human-human interactions are presented and classification results for low-resolution video and videos depicting human interactions are shown.

    73
  • Improving Deep Regression with Ordinal Entropy

    Shi-Hao Zhang, Linlin Yang, Michael Mi, Xiaoxu Zheng, Angela Yao

    arXiv · 2023

    This work provides a derivation to show that classification, with the cross-entropy loss, outperforms regression with a mean squared error loss in its ability to learn high-ent entropy feature representations.

    72
  • C2F-TCN: A Framework for Semi- and Fully-Supervised Temporal Action Segmentation

    Dipika Singhania, R. Rahaman, Angela Yao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2023

    A novel unsupervised way to learn frame-wise representation from C2F-TCN, which hinges on the clustering capabilities of the input features and the formation of multi-resolution features from the decoder's implicit structure and progressively improves in performance with more labeled data.

    69
  • Learning Probabilistic Non-Linear Latent Variable Models for Tracking Complex Activities

    Angela Yao, Juergen Gall, L. van Gool, R. Urtasun

    Neural Information Processing Systems · 2011

    An efficient stochastic gradient descent algorithm that is able to learn probabilistic non-linear latent spaces composed of multiple activities and an incremental algorithm for the online setting which can update the latent space without extensive relearning are presented.

    67
  • 2D Action Recognition Serves 3D Human Pose Estimation

    Juergen Gall, Angela Yao, L. van Gool

    Lecture notes in computer science · 2010

    This work proposes a particle-based optimization algorithm that can efficiently estimate human pose even in challenging in-house scenarios and can directly integrate the results of a 2D action recognition system as prior distribution for optimization.

    66
  • Enhancing Video Super-Resolution via Implicit Resampling-based Alignment

    Kai-yu Xu, Zi-Wei Yu, Xin Wang, Michael Bi Mi, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    Experiments on synthetic and real-world datasets show that alignment with the proposed implicit resampling enhances the performance of state-of-the-art frameworks with minimal impact on both compute and parameters.

    62
  • Towards Compact Single Image Super-Resolution via Contrastive Self-distillation

    Yanbo Wang, Shaohui Lin, Yanyun Qu, Haiyan Wu, Zhizhong Zhang, Yuan Xie, Angela Yao

    Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence · 2021

    A novel contrastive self-distillation (CSD) framework to simultaneously compress and accelerate various off-the-shelf SR models to improve the quality of SR images and PSNR/SSIM via explicit knowledge transfer is proposed.

    61
  • Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction

    Anil Armagan, Guillermo Garcia-Hernando, Seungryul Baek, Shreyas Hampali, Mahdi Rad, Zhaohui Zhang, Shipeng Xie, Mingxiu Chen, +22 more

    Lecture notes in computer science · 2020

    A public challenge to evaluate the abilities of current 3D hand pose estimators (HPEs) to interpolate and extrapolate the poses of a training set, with dramatically improved accuracy over the baseline.

    57
  • Crossing Nets: Dual Generative Models with a Shared Latent Space for Hand Pose Estimation.

    Cheng-De Wan, Thomas Probst, L. van Gool, Angela Yao

    arXiv · 2017

    The proposed discriminator network architecture is highly efficient and runs at 90 FPS on the CPU with accuracies comparable or better than state-of-art on 3 publicly available benchmarks.

    52
  • Hough Forest-Based Facial Expression Recognition from Video Sequences

    G. Fanelli, Angela Yao, P. Noel, Juergen Gall, L. van Gool

    Lecture notes in computer science · 2012

    This work presents a user-independent approach for the recognition of facial expressions from image sequences based on the eye centers' locations into tracks from which features representing shape and motion are extracted.

    46
  • SemiHand: Semi-supervised Hand Pose Estimation with Consistency

    Linlin Yang, Shicheng Chen, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    45
  • Iterative Contrast-Classify for Semi-supervised Temporal Action Segmentation

    Dipika Singhania, R. Rahaman, Angela Yao

    Proceedings of the AAAI Conference on Artificial Intelligence · 2022

    This work proposes a novel way to learn frame-wise representations from temporal convolutional networks (TCNs) by clustering input features with added time-proximity conditions and multi-resolution similarity by merging representation learning with conventional supervised learning.

    41
  • VideoQA in the Era of LLMs: An Empirical Study

    Junbin Xiao, Nanxin Huang, Hangyu Qin, Dongyang Li, Yi-Cong Li, Fengbin Zhu, Zhulin Tao, Jianxing Yu, +3 more

    International Journal of Computer Vision · 2025

    This work conducts a timely and comprehensive study of Video-LLMs’ behavior in VideoQA, aiming to elucidate their success and failure modes, and provide insights towards more human-like video understanding and question answering.

    39
  • A Generalized and Robust Framework for Timestamp Supervision in Temporal Action Segmentation

    R. Rahaman, Dipika Singhania, Alexandre H. Thiery, Angela Yao

    Lecture notes in computer science · 2022

    39
  • Multi-Scale Memory-Based Video Deblurring

    Bongil Ji, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    A memory branch is designed to memorize the blurry-sharp feature pairs in the memory bank, thus providing useful information for the blurry query input in order to achieve fine-grained deblurring.

    39
  • 39
  • Learning deep morphological networks with neural architecture search

    Yufei Hu, Nacim Belkhir, J. Angulo, Angela Yao, G. Franchi

    Pattern Recognition · 2022

    This paper proposes a method based on meta-learning to incorporate morphological operators into DNNs and demonstrates how the utility of integrating these operations in an end-to-end deep learning framework significantly increase DNN performance on various tasks, including picture classification and edge detection.

    36
  • Gesture Recognition Portfolios for Personalization

    Angela Yao, L. van Gool, Pushmeet Kohli

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2014

    35
  • Removing the Bias of Integral Pose Regression

    Ke Gu, Linlin Yang, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2021

    34
  • Make Me a BNN: A Simple Strategy for Estimating Bayesian Uncertainty from Pre-trained Models

    G. Franchi, Olivier Laurent, Maxence Legu'ery, Andrei Bursuc, Andrea Pilzer, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    The Adaptable Bayesian Neural Network (ABNN) is introduced, a simple and scalable strategy to seamlessly transform DNNs into BNNs in a post-hoc manner with minimal computational and training overheads.

    32
  • Dual Grid Net: Hand Mesh Vertex Regression from Single Depth Maps

    Cheng-De Wan, Thomas Probst, L. van Gool, Angela Yao

    Lecture notes in computer science · 2020

    A method for recovering the dense 3D surface of the hand by regressing the vertex coordinates of a mesh model from a single depth map, which achieves state-of-the-art accuracy on NYU dataset for key point localization while recovering mesh vertices and a dense correspondence map.

    30
  • DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction

    Kai Xu, T. Tse, Ji-Zong Peng, Angela Yao

    arXiv · 2024

    Compared to existing methods, DAS3R is more robust in complex motion scenarios, capable of handling videos where dynamic objects occupy a significant portion of the scene, and does not require camera pose inputs or point cloud data from SLAM-based methods.

    29
  • Bias-Compensated Integral Regression for Human Pose Estimation

    Kerui Gu, Lin-Lin Yang, Michael Bi Mi, Angela Yao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2023

    This paper uncovers an induced bias from integral regression that results from combining the softmax and the expectation operation, and proposes Bias Compensated Integral Regression (BCIR), an integral regression-based framework that compensates for the bias.

    29
  • Accelerating Video Object Segmentation with Compressed Video

    Kai-yu Xu, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    An efficient plug-and-play acceleration framework for semi-supervised video object segmentation by exploiting the temporal redundancies in videos presented by the compressed bitstream is proposed and a residual-based correction module is introduced that can fix wrongly propagated segmentation masks from noisy or erroneous motion vectors.

    29
  • Superpixel Optimization Using Higher Order Energy

    Jian-Teng Peng, Jian-Bing Shen, Angela Yao, Xue-Long Li

    IEEE Transactions on Circuits and Systems for Video Technology · 2015

    A novel superpixel extraction algorithm using a higher order energy optimization framework is proposed in this paper that generates better results with well-aligned boundaries and homogeneous effects than the existing superpixel algorithms.

    29
  • Tracking People in Broadcast Sports

    Angela Yao, Dominique Uebersax, Juergen Gall, L. van Gool

    Lecture notes in computer science · 2010

    A method for tracking people in monocular broadcast sports videos by coupling a particle filter with a vote-based confidence map of athletes, appearance features and optical flow for motion estimation that outperforms tracking with discrete target detections.

    27
  • Every Mistake Counts in Assembly

    Guodong Ding, F. Sener, Shugao Ma, Angela Yao

    arXiv · 2023

    A system that can detect ordering mistakes by utilizing a learned knowledge base that constructs a knowledge base with spatial and temporal beliefs based on observed mistakes, and demonstrates the superior performance of the belief inference algorithm in detecting ordering mistakes on the Assembly101 dataset.

    26
  • Advances in Visual Computing

    Lecture notes in computer science · 2022

    26
  • Robust Semantic Segmentation with Superpixel-Mix

    G. Franchi, Nacim Belkhir, Lan-Ha Mai, Yufei Hu, Andrei Bursuc, V. Blanz, Angela Yao

    British Machine Vision Conference (BMVC) · 2021

    This work introduces Superpixel-mix, a new superpixel-based data augmentation method with teacher-student consistency training that achieves state-of-the-art results in semi-supervised semantic segmentation on the Cityscapes dataset.

    26
  • On the Consistency of Video Large Language Models in Temporal Comprehension

    Minjoon Jung, Junbin Xiao, Byoung-Tak Zhang, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    A study on prediction consistency –a key indicator for robustness and trustworthiness of temporal grounding and proposed event temporal verification tuning that explicitly accounts for consistency.

    24
  • Temporal Action Segmentation With High-Level Complex Activity Labels

    Guodong Ding, Angela Yao

    IEEE Transactions on Multimedia · 2022

    This work is the first to propose a Constituent Action Discovery framework that only requires the video-wise high-level complex activity label as supervision for temporal action segmentation, and adopts the Hungarian matching algorithm to relate latent action prototypes to ground truth semantic classes for evaluation.

    24
  • Rethinking Visibility in Human Pose Estimation: Occluded Pose Reasoning via Transformers

    Peng-Zhan Sun, Kerui Gu, Yunsong Wang, Linlin Yang, Angela Yao

    IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) · 2024

    22
  • Discrete-Constrained Regression for Local Counting Models

    Lecture notes in computer science · 2022

    22
  • Transferring Knowledge From Text to Video: Zero-Shot Anticipation for Procedural Actions

    F. Sener, Rishabh Saraf, Angela Yao

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2022

    A hierarchical model that generalizes instructional knowledge from large-scale text corpora and transfers the knowledge to video recognizes and predicts coherent and plausible actions multiple steps into the future, all in rich natural language.

    21
  • 21
  • 19
  • Neural Network Compression via Learnable Wavelet Transforms

    M. Wolter, Shaohui Lin, Angela Yao

    Lecture notes in computer science · 2020

    This paper shows how the fast wavelet transform can be used to compress linear layers in neural networks and has significantly fewer parameters yet still perform competitively with the state-of-the-art on synthetic and real-world RNN benchmarks.

    19
  • Context-Enhanced Memory-Refined Transformer for Online Action Detection

    Zhanzhong Pang, F. Sener, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2025

    A Context-enhanced Memory-Refined Transformer (CMeRT) is proposed, which introduces a context-enhanced encoder to improve frame representations using additional near-past context and features a memory-refined decoder to leverage near-future generation to enhance performance.

    18
  • Question-Answering Dense Video Events

    Hangyu Qin, Junbin Xiao, Angela Yao

    International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) · 2025

    De DeVi is proposed, a novel training-free MLLM approach that highlights a hierarchical captioning module, a temporal event memory module, and a self-consistency checking module to respectively detect, contextualize and memorize, and ground dense-events in long videos for question answering.

    16
  • MHEntropy: Entropy Meets Multiple Hypotheses for Pose and Shape Recovery

    Rongyu Chen, Linlin Yang, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2023

    16
  • OnlineTAS: An Online Baseline for Temporal Action Segmentation

    Qing Zhong, Guodong Ding, Angela Yao

    arXiv · 2024

    An adaptive memory designed to accommodate dynamic changes in context over time is presented, alongside a feature augmentation module that enhances the frames with the memory that achieves state-of-the-art performance.

    15
  • Cross-Domain 3D Hand Pose Estimation with Dual Modalities

    Qiu-Xia Lin, Linlin Yang, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    15
  • 15
  • Learning Fine-Scaled Depth Maps from Single RGB Images.

    Jun Yu Li, Reinhard Klein, Angela Yao

    arXiv · 2016

    A multi-scale convolution neural network to learn from single RGB images fine-scaled depth maps that result in realistic 3D reconstructions and introduces spatial coordinate feature maps and a local relative depth constraint to encourage spatial coherency.

    15
  • DD-Ranking: Rethinking the Evaluation of Dataset Distillation

    Ze-Kai Li, Xin-Hao Zhong, Samir Khaki, Zhiyuan Liang, Yuhao Zhou, Ming-Jia Shi, Zi-Qiao Wang, Xuan-Lei Zhao, +22 more

    arXiv · 2025

    DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.

    14
  • 14
  • WAVE: Warping DDIM Inversion Features for Zero-Shot Text-to-Video Editing

    Yutang Feng, Sicheng Gao, Yu-Xiang Bao, Xiaodi Wang, Shumin Han, Juan Zhang, Baochang Zhang, Angela Yao

    Lecture notes in computer science · 2024

    14
  • On the Utility of 3D Hand Poses for Action Recognition

    Lecture notes in computer science · 2024

    13
  • On the Calibration of Human Pose Estimation

    Kerui Gu, Rongyu Chen, Angela Yao

    arXiv · 2023

    The proposed Calibrated ConfidenceNet (CCNet) is a light-weight post-hoc addition that improves AP by up to 1.4% on off-the-shelf pose estimation frameworks and facilitates an additional 1.0mm decrease in 3D keypoint error.

    13
  • Local and Global Point Cloud Reconstruction for 3D Hand Pose Estimation

    Zi-Wei Yu, Linlin Yang, Shicheng Chen, Angela Yao

    British Machine Vision Conference (BMVC) · 2021

    This paper presents a novel pipeline for local and global point cloud reconstruction using a 3D hand template while learning a latent representation for pose estimation, and introduces a new multi-view hand posture dataset to obtain complete 3D point clouds of the hand in the real world.

    13
  • Learning Style Compatibility for Furniture

    Divyansh Aggarwal, E. Valiyev, F. Sener, Angela Yao

    Lecture notes in computer science · 2019

    This paper investigates how Siamese networks can be used efficiently for assessing the style compatibility between images of furniture items and shows that the middle layers of pretrained CNNs can capture essential information about furniture style, which allows for efficient applications of such networks for this task.

    13
  • Scene-Text Grounding for Text-Based Video Question Answering

    Sheng Zhou, Junbin Xiao, Xun Yang, Peipei Song, Dan Guo, Angela Yao, Meng Wang, Tat-Seng Chua

    IEEE Transactions on Multimedia · 2025

    The T2S-QA model is proposed that highlights a disentangled temporal-to-spatial contrastive learning strategy for weakly-supervised scene-text grounding and grounded TextVideoQA, thus decoupling QA from scene-text recognition and promoting research towards interpretable QA.

    12
  • Deep Imbalanced Regression via Hierarchical Classification Adjustment

    Haipeng Xiong, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This work proposes a range-preserving distillation process that effectively learns a single classifier from the set of hierarchical classifiers to improve regression performance over the entire range of data.

    12
  • 12
  • 12
  • Overcoming the TradeOff between Accuracy and Plausibility in 3D Hand Shape Reconstruction

    Zi-Wei Yu, Chen Li, Linlin Yang, Xiaoxu Zheng, Michael Bi Mi, G. Lee, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    This work introduces a novel weakly-supervised hand shape estimation framework that integrates non-parametric mesh fitting with MANO model in an end-to-end fashion to yield well-aligned and high-quality 3D meshes, especially in challenging two-hand and hand-object interaction scenarios.

    11
  • UV-Based 3D Hand-Object Reconstruction with Grasp Optimization

    Zi-Wei Yu, Linlin Yang, You Xie, Ping Chen, Angela Yao

    arXiv · 2022

    A novel framework for 3D hand shape reconstruction and hand-object grasp optimization from a single RGB image in the form of a UV coordinate map is proposed and inference-time optimization is introduced to fine-tune the grasp and improve interactions between the hand and the object.

    11
  • Coherent Temporal Synthesis for Incremental Action Segmentation

    Guodong Ding, Hans Golong, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    This paper presents the first exploration of video data replay techniques for incremen-tal action segmentation, focusing on action temporal modeling with a Temporally Coherent Action (TCA) model, which represents actions using a generative model instead of storing individual frames.

    10
  • KITRO: Refining Human Mesh by 2D Clues and Kinematic-tree Rotation

    Fengyuan Yang, Kerui Gu, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    Kinematic-Tree Rotation (KITRO), a novel mesh refinement strategy that explicitly models depth and human kinematic-tree structure, is introduced, which significantly improves 3D joint estimation accuracy and achieves an ideal 2D fit simultaneously.

    9
  • The Robust Semantic Segmentation UNCV2023 Challenge Results

    Xuanlong Yu, Yi Zuo, Zitao Wang, Xiao-Wen Zhang, Jia-Xuan Zhao, Yu-Ting Yang, Licheng Jiao, Rui Peng, +22 more

    IEEE/CVF International Conference on Computer Vision Workshops (ICCVW) · 2023

    This paper outlines the winning solutions employed in addressing the MUAD uncertainty quantification challenge held at ICCV 2023, which primarily revolved around enhancing the robustness of semantic segmentation in urban scenes under varying natural adversarial conditions.

    8
  • Multi-stage Fusion for One-Click Segmentation

    Soumajit Majumder, Ansh Khurana, Abhinav Rai, Angela Yao

    Lecture notes in computer science · 2021

    This work proposes a new multi-stage guidance framework for interactive segmentation that incorporates user cues at different stages of the network to allow user interactions to impact the final segmentation output in a more direct way.

    8
  • Gated Complex Recurrent Neural Networks.

    M. Wolter, Angela Yao

    arXiv · 2018

    A novel complex gate recurrent cell is presented that exhibits excellent stability and convergence properties when used together with norm-preserving state transition matrices and demonstrates competitive performance of the complex gated RNN on the synthetic memory and adding task, as well as on the real-world task of human motion prediction.

    8
  • Ego-Grounding for Personalized Question-Answering in Egocentric Videos

    Jun Xiao, Sheng Zhang, Peng Zhu, Angela Yao

    arXiv · 2026

    MyEgo is introduced, the first egocentric VideoQA dataset designed to evaluate MLLMs'ability to understand, remember, and reason about the camera wearer, and the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos.

    7
  • Sequence Prediction Using Spectral RNNs

    M. Wolter, Juergen Gall, Angela Yao

    Lecture notes in computer science · 2020

    This work predicts time series data drawn from the chaotic Mackey-Glass differential equation and real-world power load and motion capture data and proposes to combine Fourier methods and recurrent neural network architectures.

    7
  • MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering

    Jun Xiao, Jiajun Chen, Tianxiang Sun, Xun Yang, Angela Yao

    arXiv · 2026

    MuKV, a method that features a multi-grained KV cache compression module and a semi-hierarchical retrieval approach to improve both efficiency and accuracy for long streaming VideoQA, is proposed.

    6
  • Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing

    Xiaolu Liu, Yicong Li, Qi-Yuan He, Jiayin Zhu, Wei Ji, Angela Yao, Jian-Ke Zhu

    arXiv · 2026

    The proposed Interp3D is a novel training-free framework for textured 3D morphing that harnesses generative priors and adopts a progressive alignment principle to ensure both geometric fidelity and texture coherence.

    6
  • VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation

    Xin-Yao Liao, Qi-Yuan He, Kai Xu, Xiao-Ye Qu, Yicong Li, Wei Wei, Angela Yao

    arXiv · 2025

    VA-$\pi$ is proposed, a lightweight post-training framework that directly optimizes AR models with a principled pixel-space objective, and enables rapid adaptation of existing AR generators, without neither tokenizer retraining nor external reward models.

    6
  • Deep Regression Representation Learning with Topology

    Shihao Zhang, Kenji Kawaguchi, Angela Yao

    arXiv · 2024

    PH-Reg is introduced, a regularizer specific to regression that matches the intrinsic dimension and topology of the feature space with the target space and suggests a feature space that is topologically similar to the target space will better align with the IB principle.

    6
  • TemporalUV: Capturing Loose Clothing with Temporally Coherent UV Coordinates

    You Xie, Hui-li Mao, Angela Yao, Nils Thürey

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2022

    A novel approach to generate temporally coherent UV coordinates for loose clothing by implementing a differentiable pipeline to learn UV mapping between a sequence of RGB inputs and textures via UV coordinates.

    6
  • InstructHumans: Editing Animated 3D Human Textures With Instructions

    Jiayin Zhu, Linlin Yang, Angela Yao

    IEEE Transactions on Multimedia · 2026

    This work shows that naively using SDS harms editing, as it may destroy consistency, and proposes a modified SDS for Editing (SDS-E) that selectively incorporates subterms of SDS across diffusion timesteps.

    5
  • Spatial and temporal beliefs for mistake detection in assembly tasks

    Guodong Ding, F. Sener, Shugao Ma, Angela Yao

    Computer Vision and Image Understanding · 2025

    5
  • Learning to generate training datasets for robust semantic segmentation

    Marwane Hariat, Olivier Laurent, Rémi Kazmierczak, Shihao Zhang, Andrei Bursuc, Angela Yao, G. Franchi

    IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) · 2024

    This work designs Robusta, a novel robust conditional generative adversarial network to generate realistic and plausible perturbed images that can be used to train reliable segmentation models by leveraging the synergy between label-to-image generators and image-to-label segmentation models.

    5
  • Weakly-Supervised Dense Action Anticipation

    Haotong Zhang, Fuhai Chen, Angela Yao

    British Machine Vision Conference (BMVC) · 2021

    A framework that generates pseudo-labels for future actions and their durations and adaptively refines them through a refinement module is proposed and is competitive even when compared to fully supervised state-of-the-art models.

    5
  • Two-in-One Refinement for Interactive Segmentation

    Soumajit Majumder, Abhinav Rai, Ansh Khurana, Angela Yao

    British Machine Vision Conference (BMVC) · 2020

    A simple yet intuitive two-in-one re-nement strategy placing clicks on the boundary of the object of interest and a boundary-aware loss that encourages segmentation masks to respect instance boundaries are proposed.

    5
  • Supervised Deep Kriging for Single-Image Super-Resolution

    G. Franchi, Angela Yao, A. Kolb

    Lecture notes in computer science · 2019

    This work proposes a novel single-image super-resolution approach based on the geostatistical method of kriging that combines the krigs weight generation and kriged process into a joint network that can be learned end-to-end and achieves competitive super- resolution results as other state-of-the-art methods.

    5
  • Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams

    Yun Wang, Jun-Bin Xiao, Han Lyu, Yi-Fan Wang, Jingwei Zuo, Zhan-Jie Zhang, Hongjia Huang, Da-Peng Wu, +1 more

    arXiv · 2026

    DirectMe is proposed, a framework that incrementally constructs and maintains a structured spatial memory from streaming egocentric observations that significantly improves the spatial reasoning of leading multimodal LLMs and surpasses many spatially aware and long-form streaming video models.

    4
  • Don't Pause! Every prediction matters in a streaming video

    Dibyadip Chatterjee, Zhanzhong Pang, F. Sener, Ya-Le Song, Angela Yao

    arXiv · 2026

    AsynKV, a training-free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior, serves as a strong baseline on SPOT-Bench, outperforming existing streaming models, and achieves state-of-the-art on retrospective benchmarks.

    4
  • Synthetic-to-Real Pose Estimation with Geometric Reconstruction

    Qiu-Xia Lin, Kerui Gu, Linlin Yang, Angela Yao

    neural information processing systems · 2023

    This work proposes a reconstruction-based strategy as a complement to pseudo-labelling for synthetic-to-real domain adaptation and provides a novel solution to effectively correct confident yet inaccurate keypoint locations through image reconstruction in domain adaptation.

    4
  • Opening the Vocabulary of Egocentric Actions

    neural information processing systems · 2023

    4
  • Transformed ROIs for capturing visual transformations in videos

    Abhinav Rai, F. Sener, Angela Yao

    Computer Vision and Image Understanding · 2022

    TROI, a plug-and-play module for CNNs to reason between mid-level feature representations that are otherwise separated in space and time, achieves state-of-the-art action recognition results on the large-scale datasets Something-Something-V2 and EPIC-Kitchens-100.

    4
  • 4
  • HANDS18: Methods, Techniques and Applications for Hand Observation

    I. Oikonomidis, Guillermo Garcia-Hernando, Angela Yao, Antonis A. Argyros, V. Lepetit, Tae-Kyun Kim

    Lecture notes in computer science · 2019

    The main conclusions that can be drawn are the turn of the community towards RGB data and the maturation of some methods and techniques, which in turn has led to increasing interest for real-world applications.

    4
  • Noise-Robust tiny object localization with flows

    H. Sun, Linlin Yang, Ronyu Chen, Kerui Gu, Bao-Chang Zhang, Angela Yao, Xianbin Cao

    Pattern Recognition · 2026

    This work proposes Tiny Object Localization with Flows (TOLF), a noise-robust localization framework leveraging normalizing flows for flexible error modeling and uncertainty-guided optimization, enabling robust learning under noisy supervision.

    3
  • EgoBlind: Towards Egocentric Visual Assistance for the Blind

    neural information processing systems · 2025

    3
  • Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes

    Joseph Hoche, Andrei Bursuc, David Brellmann, G. Louppe, P. Izmailov, Angela Yao, G. Franchi

    arXiv · 2025

    Semantic Gaussian Process Uncertainty (SGPU) is proposed, a Bayesian framework that quantifies semantic uncertainty by analyzing the geometric structure of answer embeddings, avoiding brittle clustering and demonstrating state-of-the-art calibration and discriminative performance.

    3
  • Analyzing and Diagnosing Pose Estimation with Attributions

    Qi-Yuan He, Linlin Yang, Kerui Gu, Qiu-Xia Lin, Angela Yao

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2023

    3
  • Learning Unorthogonalized Matrices for Rotation Estimation

    Kerui Gu, Zhihao Li, Shiyong Liu, Jianzhuang Liu, Song-Cen Xu, Youliang Yan, Michael Bi Mi, Kenji Kawaguchi, +1 more

    arXiv · 2023

    By replacing the orthogonalization incorporated representation with the proposed PRoM in various rotation-related tasks, this work achieves state-of-the-art results on large-scale benchmarks for human pose estimation.

    3
  • A Generalized & Robust Framework For Timestamp Supervision in Temporal Action Segmentation

    R. Rahaman, Dipika Singhania, Alexandre Hoang Thiery, Angela Yao

    arXiv · 2022

    This work proposes a novel Expectation-Maximization (EM) based approach that leverages the label uncertainty of unlabelled frames and is robust enough to accommodate possible annotation errors and introduces the new challenging annotation setup of Skip-tag supervision.

    3
  • 3
  • Fourier RNNs for Sequence Prediction

    M. Wolter, Angela Yao

    arXiv · 2018

    This work proposes to integrate Fourier methods into complex recurrent neural network architectures and show accuracy improvements on prediction tasks as well as computational load reductions.

    3
  • Decouple and Cache: KV Cache Construction for Streaming Video Understanding

    Zhanzhong Pang, Dibyadip Chatterjee, F. Sener, Angela Yao

    arXiv · 2026

    Decoupled Streaming Cache is proposed, a training-free cache construction mechanism that adapts pretrained offline models to streaming settings that maintains a cumulative past KV cache while constructing a separate instant cache on-demand, decoupled from past caches to preserve the informativeness of recent inputs.

    2
  • Unleashing the Power of LLMs for Medical Video Answer Localization

    Jun-Bin Xiao, Qingyun Li, Yusen Yang, Liang Qiu, Angela Yao

    Lecture notes in computer science · 2025

    2
  • A closer look at branch classifiers of multi-exit architectures

    Computer Vision and Image Understanding · 2023

    2
  • 2
  • 2
  • Object-centered Fourier Motion Estimation and Segment-Transformation Prediction.

    The European Symposium on Artificial Neural Networks · 2020

    2
  • Localized Interactive Instance Segmentation

    Soumajit Majumder, Angela Yao

    Lecture notes in computer science · 2019

    This work proposes a clicking scheme wherein user interactions are restricted to the proximity of the object, and a novel transformation of the user-provided clicks to generate a weak localization prior on the object which is consistent with image structures such as edges, textures etc.

    2
  • Fourier RNNs for Sequence Analysis and Prediction.

    M. Wolter, Angela Yao

    arXiv.org · 2018

    This work proposes to integrate Fourier methods into complex recurrent neural network architectures and show accuracy improvements on analysis and prediction tasks as well as computational load reductions.

    2
  • Vision-Based Human Motion Analysis

    Repository for Publications and Research Data (ETH Zurich) · 2012

    2
  • RelaxFlow: Text-Driven Amodal 3D Generation

    Jiayin Zhu, Guoji Fu, Xiaolu Liu, Qi-Yuan He, Yicong Li, Angela Yao

    arXiv · 2026

    This work formalizes text-driven amodal 3D generation, where text prompts steer the completion of unseen regions while strictly preserving input observation, and proposes RelaxFlow, a training-free dual-branch framework that decouples control granularity via a Multi-Prior Consensus Module and a Relaxation Mechanism.

    1
  • On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding

    Zhanzhong Pang, Dibyadip Chatterjee, F. Sener, Angela Yao

    arXiv · 2026

    GAD improves both accuracy and efficiency over generative methods, achieving state-of-the-art results on four tasks across five datasets, including an average 2.5% accuracy gain and 3x faster inference on the largest COIN benchmark.

    1
  • VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

    Xin-Yao Liao, Qi-Yuan He, Yicong Li, Jiayin Zhu, Xiao-Ye Qu, Wei Wei, Angela Yao

    arXiv · 2026

    Across class-conditional image generation with VAR, text-to-image generation with Infinity, and text-to-video generation with InfinityStar, VPG improves generation quality without retraining the base model, reducing FID on VAR by 0.36 on average and improving benchmark performance on both image and video generation.

    1
  • UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    Tanuj Sur, Shashank Tripathi, Nikos Athanasiou, Ha Linh Nguyen, Kai Xu, Michael J. Black, Angela Yao

    arXiv · 2026

    UniCon3R is introduced, a unified feed-forward framework for online human-scene 4D reconstruction from monocular videos that outperforms state-of-the-art baselines on physical plausibility and global human motion estimation while achieving real-time online inference.

    1
  • Online Test-time Adaptation for 3D Human Pose Estimation: A Practical Perspective with Estimated 2D Poses

    Qiu-Xia Lin, Kerui Gu, Linlin Yang, Angela Yao

    arXiv · 2025

    This paper addresses adapting models to streaming videos with estimated 2D poses by proposing adaptive aggregation, a two-stage optimization, and local augmentation for handling varying levels of estimated pose error.

    1
  • Improving Deep Regression with Tightness

    Shihao Zhang, Yu-Guang Yan, Angela Yao

    arXiv · 2025

    This work reveals that preserving ordinality reduces the conditional entropy of representation Z conditional on the target Y, and introduces an optimal transport-based regularizer to preserve the similarity relationships of targets in the feature space to reduce H(Z|Y) .

    1
  • AID: Attention Interpolation of Text-to-Image Diffusion

    neural information processing systems · 2024

    1
  • Computer Vision – ACCV 2024

    W. He, Z. Huang, X. Meng, X. Qi, R. Xiao, C.-G. Li

    Lecture notes in computer science · 2024

    This paper attempts to integrate an efficient clustering module into the principled framework for learning structured representation, in which the clustering module is used to provide partition information to guide the cluster-wise compression and the learned embeddings is aligned to desired geometric structures in turn to help for yielding more accurate partitions.

    1
  • 1
  • 1
  • 1
  • 1
  • Towards deep neural network compression via learnable wavelet transforms

    M. Wolter, Shaohui Lin, Angela Yao

    arXiv · 2020

    This paper shows how the fast wavelet transform can be used to compress linear layers in neural networks and has significantly fewer parameters yet still perform competitively with the state-of-the-art on synthetic and real-world RNN benchmarks.

    1
  • AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation

    Proceedings of the AAAI Conference on Artificial Intelligence · 2026

    –
  • TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction

    Yi-Yao Huang, Zhe-Dong Zheng, Zi-Wei Yu, Ya-Xiong Wang, Tze Ho Elden Tse, Angela Yao

    Proceedings - IEEE International Conference on Robotics and Automation/Proceedings · 2026

    A new Text-Instructed Generation and Refinement (TIGeR) framework, harnessing the power of intuitive text-driven priors to steer the object shape refinement and pose estimation, which shows robustness to occlusion, while maintaining compatibility with heterogeneous prior sources.

    –
  • Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models

    Guodong Ding, Angela Yao

    arXiv · 2026

    Gate-and-Merge, a zero-shot framework that enables compositional personalization without the need for co-occurrence training, is introduced, and consistent gains in performance across multiple personalization tasks in both single-concept and compositional settings are shown.

    –
  • –
  • Online Test-Time Adaptation for 3D Human Pose Estimation: A Practical Perspective With Noisy 2D Labels

    Qiu-Xia Lin, Kerui Gu, Lin-Lin Yang, Angela Yao

    IEEE Transactions on Circuits and Systems for Video Technology · 2026

    –
  • –
  • –
  • –
  • Optimizing Mesh Animation from Video via Shape Flow Guidance

    Lecture notes in computer science · 2026

    –
  • LightAVSeg: Lightweight Audio-Visual Segmentation

    Qing Zhong, Guodong Ding, Lingqiao Liu, Zai-Wen Feng, Lin-Yuan-Bo Wu, Angela Yao

    arXiv · 2026

    LightAVSeg is proposed, a lightweight framework that replaces heavy attention with a decoupled design for semantic filtering and spatial grounding, resulting in interaction costs that scale linearly with spatial resolution.

    –
  • –
  • ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

    Fengyuan Yang, Luying Huang, Jiazhi Guan, Quanwei Yang, Dong-Xu Pan, Jianglin Fu, Hao Feng, Wei He, +3 more

    arXiv · 2026

    A canonical-space injection mechanism that decouples human dynamics from environmental cues via cross-attention is introduced and a novel positional embedding strategy that establishes spatial correspondences between disparate spatial domains without any heuristic 3D alignments is proposed.

    –
  • –
  • –
  • Video Question Answering and Beyond

    Yi-Cong Li, Junbin Xiao, Angela Yao, T. Chua

    ACM International Conference on Multimedia (ACM MM) · 2025

    This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration.

    –
  • –
  • Computer Vision – ACCV 2024 Workshops

    Lecture notes in computer science · 2025

    –
  • Diagnosing Pretrained Models for Out-of-Distribution Detection

    Haipeng Xiong, Kai Xu, Angela Yao

    IEEE/CVF International Conference on Computer Vision (ICCV) · 2025

    –
  • LocalSR: Image Super-Resolution in Local Region

    Bo Ji, Angela Yao

    arXiv · 2024

    Experimental results indicate that the proposed context-based local super-resolution (CLSR) approach, with its reduced low complexity, outperforms variants that focus exclusively on the ROI.

    –
  • High-Resolution Be Aware! Improving the Self-Supervised Real-World Super-Resolution

    Yuehan Zhang, Angela Yao

    arXiv · 2024

    A controller to adjust the degradation modeling based on the quality of super-resolution results is proposed and a novel feature-alignment regularizer is introduced that directly constrains the distribution of super-resolved images.

    –
  • –
  • –
  • –
  • Workshop on Interactive and Adaptive Learning in an Open World

    Alexander Freytag, V. Ferrari, Mario Fritz, Uwe Franke, T. Boult, Juergen Gall, W. Scheirer, Angela Yao, +1 more

    Lecture notes in computer science · 2019

    This workshop at ECCV 2018 in Munich served as a discussion forum for experts in this field and in the following it is given a brief overview.

    –
  • –
  • Data Driven Synthesis of Hand Grasps from 3-D Object Models

    Soumajit Majumder, Haojiong Chen, Angela Yao

    Eurographics · 2017

    This work forms grasp synthesis as a constrained optimization problem which takes into account the anthropomorphic and kinematic limitations of a human hand as well as the local and global geometric properties of the interacting object.

    –
  • Optimizing Over a Set of Manifolds ⋆

    Juergen Gall, Angela Yao, L. van Gool

    2010

    This paper presents a comprehensive description of the algorithm that is proposed in “2D Action Recognition Serves 3D Human Pose Estimation”, which automates the very labor-intensive and therefore time-heavy and expensive process of human pose estimation.

    –

Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.