Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
View as a treeStudents and postdocs9
- Zhang ShihaoFrom a thesis record ↗
- Lin QiuxiaFrom a thesis record ↗
- Gu KeruiFrom a thesis record ↗
- Ji BoFrom a thesis record ↗
- Zhang YuehanFrom a thesis record ↗
- Pang ZhanzhongFrom a thesis record ↗
- Dipika SinghaniaFrom a thesis record ↗
- Xiong HaipengFrom a thesis record ↗
Show 1 more
- Yu ZiweiFrom a thesis record ↗
Possible advisorsa guess from early papers, not confirmed
- L. van GoolSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works166 from public data
- NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions1,033
NExT-QA is introduced, a rigorously designed video question answering (VideoQA) benchmark to advance video understanding from describing to explaining the temporal actions, and it is found that top-performing methods excel at shallow scene descriptions but are weak in causal and temporal action reasoning.
- 653
- Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities432
The first multi-view action dataset, with si-multaneous static and egocentric recordings, and a novel task of detecting mistakes is proposed, to investigate generalization to new toys, cross-view transfer, long-tailed distributions, and pose vs. appearance.
- 301
- 276
- Does Human Action Recognition Benefit from Pose Estimation?200
Comparing pose-based, appearance-based and combined pose and appearance features for action recognition in a home-monitoring scenario shows that posebased features outperform low-level appearance features, even when heavily corrupted by noise, suggesting that pose estimation is beneficial for the action recognition task.
- Dense 3D Regression for Hand Pose Estimation198
A simple and effective method for 3D hand pose estimation from a single depth frame based on dense pixel-wise estimation that outperforms all previous state-of-the-art approaches by a large margin and outperforms various other proposed methods.
- Temporal Aggregate Representations for Long-Range Video Understanding185
This work addresses questions of temporal extent, scaling, and level of semantic abstraction with a flexible multi-granular temporal aggregation framework and shows that it is possible to achieve state of the art in both next action and dense anticipation with simple techniques such as max-pooling and attention.
- Can I Trust Your Answer? Visually Grounded Video Question Answering183
NExT-GQA is constructed - an extension of NExT-QA with 10.5K temporal grounding (or location) labels tied to the original QA pairs tied to the original VLMs and aims to push towards trustworthy VLMs in VQA systems.
- Temporal Action Segmentation: An Analysis of Modern Techniques165
This survey analyzes and summarizes the most significant contributions and trends in temporal action segmentation in videos, and systematically investigates two essential techniques of this topic, i.e., frame representation and temporal modeling.
- Crossing Nets: Combining GANs and VAEs with a Shared Latent Space for Hand Pose Estimation159
This work proposes modelling the statistical relationship of 3D hand poses and corresponding depth images using two deep generative models with a shared latent space to prevent over-fitting and to better exploit unlabeled depth maps.
- Video as Conditional Graph Hierarchy for Multi-Granular Question Answering154
This work proposes to model video as a conditional graph hierarchy which weaves together visual facts of different granularity in a level-wise manner, with the guidance of corresponding textual cues to align with the multi-granular essence of linguistic concepts in language queries.
- Unsupervised Learning and Segmentation of Complex Activities from Video128
This paper proposes an iterative discriminative-generative approach which alternates between discriminatively learning the appearance of sub-activities from the videos' visual features to sub-activity labels and generatively modelling the temporal structure of sub theactivities using a Generalized Mallows Model.
- Disentangling Latent Hands for Image Synthesis and Pose Estimation127
Experiments show that the dVAE can synthesize highly realistic images of the hand specifiable by both pose and image background content and also estimate 3D hand poses from RGB images with accuracy competitive with state-of-the-art on two public benchmarks.
- Coupled Action Recognition and Pose Estimation from Multiple Views118
A framework for coupled action recognition and pose estimation is presented by formulating pose estimation as an optimization over a set of action-specific manifolds to demonstrate not only the feasibility of using extracted 3D poses for action recognition, but also improved performance in comparison to action recognition using low-level appearance features.
- 117
- 108
- 103
- Scaling for Training Time and Post-hoc Out-of-distribution Detection Enhancement102
It is demonstrated that activation pruning has a detrimental effect on OOD detection, while activation scaling enhances it, and a simple yet effective post-hoc network enhancement method is proposed, SCALE, which attains state-of-the-art Ood detection performance without compromising in-distribution (ID) accuracy.
- Hand Pose Estimation from Local Surface Normals95
A hierarchical regression framework for estimating hand joint positions from single depth images based on local surface normals and a conditional regression forest, i.e. the Frame Conditioned Regression Forest (FCRF) which uses a new normal difference feature.
- Deep morphological networks93
This paper demonstrates on various examples that new layers making use of the morphological non-linearities are complementary to convolution layers and can be used to integrate the non- linear operations and pooling into a joint operation.
- 93
- 91
- Contrastive Video Question Answering via Video Graph Transformer78
With superior video encoding and QA solution, it is shown that CoVGT can achieve much better performances than previous arts on video reasoning tasks and can also benefit from cross-modal pretraining, yet with orders of magnitude smaller data.
- Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly Detection78
A novel bi-directional architecture with three consistency constraints to comprehensively regularize the prediction task from pixel-wise, cross-modal, and temporal-sequence levels to outperforms advanced anomaly detectors and achieves state-of-the-art results.
- Efficient Unsupervised Temporal Segmentation of Motion Data78
A method for automated temporal segmentation of human motion data into distinct actions and compositing motion primitives based on self-similar structures in the motion sequence is introduced, which requires no assumptions about the motion sequences at hand and no user interaction for the segmentation or clustering.
- Zero-Shot Anticipation for Instructional Activities75
A hierarchical model is presented that generalizes instructional knowledge from large-scale text-corpora and transfers the knowledge to the visual domain and predicts coherent and plausible actions multiple steps into the future, all in rich natural language.
- Variations of a Hough-Voting Action Recognition System73
Two variations of a Hough-voting framework for action recognition for group actions with human-human interactions are presented and classification results for low-resolution video and videos depicting human interactions are shown.
- Improving Deep Regression with Ordinal Entropy72
This work provides a derivation to show that classification, with the cross-entropy loss, outperforms regression with a mean squared error loss in its ability to learn high-ent entropy feature representations.
- C2F-TCN: A Framework for Semi- and Fully-Supervised Temporal Action Segmentation69
A novel unsupervised way to learn frame-wise representation from C2F-TCN, which hinges on the clustering capabilities of the input features and the formation of multi-resolution features from the decoder's implicit structure and progressively improves in performance with more labeled data.
- Learning Probabilistic Non-Linear Latent Variable Models for Tracking Complex Activities67
An efficient stochastic gradient descent algorithm that is able to learn probabilistic non-linear latent spaces composed of multiple activities and an incremental algorithm for the online setting which can update the latent space without extensive relearning are presented.
- 2D Action Recognition Serves 3D Human Pose Estimation66
This work proposes a particle-based optimization algorithm that can efficiently estimate human pose even in challenging in-house scenarios and can directly integrate the results of a 2D action recognition system as prior distribution for optimization.
- Enhancing Video Super-Resolution via Implicit Resampling-based Alignment62
Experiments on synthetic and real-world datasets show that alignment with the proposed implicit resampling enhances the performance of state-of-the-art frameworks with minimal impact on both compute and parameters.
- Towards Compact Single Image Super-Resolution via Contrastive Self-distillation61
A novel contrastive self-distillation (CSD) framework to simultaneously compress and accelerate various off-the-shelf SR models to improve the quality of SR images and PSNR/SSIM via explicit knowledge transfer is proposed.
- Measuring Generalisation to Unseen Viewpoints, Articulations, Shapes and Objects for 3D Hand Pose Estimation Under Hand-Object Interaction57
A public challenge to evaluate the abilities of current 3D hand pose estimators (HPEs) to interpolate and extrapolate the poses of a training set, with dramatically improved accuracy over the baseline.
- Crossing Nets: Dual Generative Models with a Shared Latent Space for Hand Pose Estimation.52
The proposed discriminator network architecture is highly efficient and runs at 90 FPS on the CPU with accuracies comparable or better than state-of-art on 3 publicly available benchmarks.
- Hough Forest-Based Facial Expression Recognition from Video Sequences46
This work presents a user-independent approach for the recognition of facial expressions from image sequences based on the eye centers' locations into tracks from which features representing shape and motion are extracted.
- 45
- Iterative Contrast-Classify for Semi-supervised Temporal Action Segmentation41
This work proposes a novel way to learn frame-wise representations from temporal convolutional networks (TCNs) by clustering input features with added time-proximity conditions and multi-resolution similarity by merging representation learning with conventional supervised learning.
- VideoQA in the Era of LLMs: An Empirical Study39
This work conducts a timely and comprehensive study of Video-LLMs’ behavior in VideoQA, aiming to elucidate their success and failure modes, and provide insights towards more human-like video understanding and question answering.
- 39
- Multi-Scale Memory-Based Video Deblurring39
A memory branch is designed to memorize the blurry-sharp feature pairs in the memory bank, thus providing useful information for the blurry query input in order to achieve fine-grained deblurring.
- 39
- Learning deep morphological networks with neural architecture search36
This paper proposes a method based on meta-learning to incorporate morphological operators into DNNs and demonstrates how the utility of integrating these operations in an end-to-end deep learning framework significantly increase DNN performance on various tasks, including picture classification and edge detection.
- 35
- 34
- Make Me a BNN: A Simple Strategy for Estimating Bayesian Uncertainty from Pre-trained Models32
The Adaptable Bayesian Neural Network (ABNN) is introduced, a simple and scalable strategy to seamlessly transform DNNs into BNNs in a post-hoc manner with minimal computational and training overheads.
- Dual Grid Net: Hand Mesh Vertex Regression from Single Depth Maps30
A method for recovering the dense 3D surface of the hand by regressing the vertex coordinates of a mesh model from a single depth map, which achieves state-of-the-art accuracy on NYU dataset for key point localization while recovering mesh vertices and a dense correspondence map.
- DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction29
Compared to existing methods, DAS3R is more robust in complex motion scenarios, capable of handling videos where dynamic objects occupy a significant portion of the scene, and does not require camera pose inputs or point cloud data from SLAM-based methods.
- Bias-Compensated Integral Regression for Human Pose Estimation29
This paper uncovers an induced bias from integral regression that results from combining the softmax and the expectation operation, and proposes Bias Compensated Integral Regression (BCIR), an integral regression-based framework that compensates for the bias.
- Accelerating Video Object Segmentation with Compressed Video29
An efficient plug-and-play acceleration framework for semi-supervised video object segmentation by exploiting the temporal redundancies in videos presented by the compressed bitstream is proposed and a residual-based correction module is introduced that can fix wrongly propagated segmentation masks from noisy or erroneous motion vectors.
- Superpixel Optimization Using Higher Order Energy29
A novel superpixel extraction algorithm using a higher order energy optimization framework is proposed in this paper that generates better results with well-aligned boundaries and homogeneous effects than the existing superpixel algorithms.
- Tracking People in Broadcast Sports27
A method for tracking people in monocular broadcast sports videos by coupling a particle filter with a vote-based confidence map of athletes, appearance features and optical flow for motion estimation that outperforms tracking with discrete target detections.
- Every Mistake Counts in Assembly26
A system that can detect ordering mistakes by utilizing a learned knowledge base that constructs a knowledge base with spatial and temporal beliefs based on observed mistakes, and demonstrates the superior performance of the belief inference algorithm in detecting ordering mistakes on the Assembly101 dataset.
- 26
- Robust Semantic Segmentation with Superpixel-Mix26
This work introduces Superpixel-mix, a new superpixel-based data augmentation method with teacher-student consistency training that achieves state-of-the-art results in semi-supervised semantic segmentation on the Cityscapes dataset.
- On the Consistency of Video Large Language Models in Temporal Comprehension24
A study on prediction consistency –a key indicator for robustness and trustworthiness of temporal grounding and proposed event temporal verification tuning that explicitly accounts for consistency.
- Temporal Action Segmentation With High-Level Complex Activity Labels24
This work is the first to propose a Constituent Action Discovery framework that only requires the video-wise high-level complex activity label as supervision for temporal action segmentation, and adopts the Hungarian matching algorithm to relate latent action prototypes to ground truth semantic classes for evaluation.
- 22
- 22
- Transferring Knowledge From Text to Video: Zero-Shot Anticipation for Procedural Actions21
A hierarchical model that generalizes instructional knowledge from large-scale text corpora and transfers the knowledge to video recognizes and predicts coherent and plausible actions multiple steps into the future, all in rich natural language.
- 21
- 19
- Neural Network Compression via Learnable Wavelet Transforms19
This paper shows how the fast wavelet transform can be used to compress linear layers in neural networks and has significantly fewer parameters yet still perform competitively with the state-of-the-art on synthetic and real-world RNN benchmarks.
- Context-Enhanced Memory-Refined Transformer for Online Action Detection18
A Context-enhanced Memory-Refined Transformer (CMeRT) is proposed, which introduces a context-enhanced encoder to improve frame representations using additional near-past context and features a memory-refined decoder to leverage near-future generation to enhance performance.
- Question-Answering Dense Video Events16
De DeVi is proposed, a novel training-free MLLM approach that highlights a hierarchical captioning module, a temporal event memory module, and a self-consistency checking module to respectively detect, contextualize and memorize, and ground dense-events in long videos for question answering.
- 16
- OnlineTAS: An Online Baseline for Temporal Action Segmentation15
An adaptive memory designed to accommodate dynamic changes in context over time is presented, alongside a feature augmentation module that enhances the frames with the memory that achieves state-of-the-art performance.
- 15
- 15
- Learning Fine-Scaled Depth Maps from Single RGB Images.15
A multi-scale convolution neural network to learn from single RGB images fine-scaled depth maps that result in realistic 3D reconstructions and introduces spatial coordinate feature maps and a local relative depth constraint to encourage spatial coherency.
- DD-Ranking: Rethinking the Evaluation of Dataset Distillation14
DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods are proposed, which provide a more comprehensive and fair evaluation standard for future research advancements.
- 14
- 14
- 13
- On the Calibration of Human Pose Estimation13
The proposed Calibrated ConfidenceNet (CCNet) is a light-weight post-hoc addition that improves AP by up to 1.4% on off-the-shelf pose estimation frameworks and facilitates an additional 1.0mm decrease in 3D keypoint error.
- Local and Global Point Cloud Reconstruction for 3D Hand Pose Estimation13
This paper presents a novel pipeline for local and global point cloud reconstruction using a 3D hand template while learning a latent representation for pose estimation, and introduces a new multi-view hand posture dataset to obtain complete 3D point clouds of the hand in the real world.
- Learning Style Compatibility for Furniture13
This paper investigates how Siamese networks can be used efficiently for assessing the style compatibility between images of furniture items and shows that the middle layers of pretrained CNNs can capture essential information about furniture style, which allows for efficient applications of such networks for this task.
- Scene-Text Grounding for Text-Based Video Question Answering12
The T2S-QA model is proposed that highlights a disentangled temporal-to-spatial contrastive learning strategy for weakly-supervised scene-text grounding and grounded TextVideoQA, thus decoupling QA from scene-text recognition and promoting research towards interpretable QA.
- Deep Imbalanced Regression via Hierarchical Classification Adjustment12
This work proposes a range-preserving distillation process that effectively learns a single classifier from the set of hierarchical classifiers to improve regression performance over the entire range of data.
- 12
- 12
- Overcoming the TradeOff between Accuracy and Plausibility in 3D Hand Shape Reconstruction11
This work introduces a novel weakly-supervised hand shape estimation framework that integrates non-parametric mesh fitting with MANO model in an end-to-end fashion to yield well-aligned and high-quality 3D meshes, especially in challenging two-hand and hand-object interaction scenarios.
- UV-Based 3D Hand-Object Reconstruction with Grasp Optimization11
A novel framework for 3D hand shape reconstruction and hand-object grasp optimization from a single RGB image in the form of a UV coordinate map is proposed and inference-time optimization is introduced to fine-tune the grasp and improve interactions between the hand and the object.
- Coherent Temporal Synthesis for Incremental Action Segmentation10
This paper presents the first exploration of video data replay techniques for incremen-tal action segmentation, focusing on action temporal modeling with a Temporally Coherent Action (TCA) model, which represents actions using a generative model instead of storing individual frames.
- KITRO: Refining Human Mesh by 2D Clues and Kinematic-tree Rotation9
Kinematic-Tree Rotation (KITRO), a novel mesh refinement strategy that explicitly models depth and human kinematic-tree structure, is introduced, which significantly improves 3D joint estimation accuracy and achieves an ideal 2D fit simultaneously.
- The Robust Semantic Segmentation UNCV2023 Challenge Results8
This paper outlines the winning solutions employed in addressing the MUAD uncertainty quantification challenge held at ICCV 2023, which primarily revolved around enhancing the robustness of semantic segmentation in urban scenes under varying natural adversarial conditions.
- Multi-stage Fusion for One-Click Segmentation8
This work proposes a new multi-stage guidance framework for interactive segmentation that incorporates user cues at different stages of the network to allow user interactions to impact the final segmentation output in a more direct way.
- Gated Complex Recurrent Neural Networks.8
A novel complex gate recurrent cell is presented that exhibits excellent stability and convergence properties when used together with norm-preserving state transition matrices and demonstrates competitive performance of the complex gated RNN on the synthetic memory and adding task, as well as on the real-world task of human motion prediction.
- Ego-Grounding for Personalized Question-Answering in Egocentric Videos7
MyEgo is introduced, the first egocentric VideoQA dataset designed to evaluate MLLMs'ability to understand, remember, and reason about the camera wearer, and the crucial role of ego-grounding and long-range memory in enabling personalized QA in egocentric videos.
- Sequence Prediction Using Spectral RNNs7
This work predicts time series data drawn from the chaotic Mackey-Glass differential equation and real-world power load and motion capture data and proposes to combine Fourier methods and recurrent neural network architectures.
- MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering6
MuKV, a method that features a multi-grained KV cache compression module and a semi-hierarchical retrieval approach to improve both efficiency and accuracy for long streaming VideoQA, is proposed.
- Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing6
The proposed Interp3D is a novel training-free framework for textured 3D morphing that harnesses generative priors and adopts a progressive alignment principle to ensure both geometric fidelity and texture coherence.
- VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation6
VA-$\pi$ is proposed, a lightweight post-training framework that directly optimizes AR models with a principled pixel-space objective, and enables rapid adaptation of existing AR generators, without neither tokenizer retraining nor external reward models.
- Deep Regression Representation Learning with Topology6
PH-Reg is introduced, a regularizer specific to regression that matches the intrinsic dimension and topology of the feature space with the target space and suggests a feature space that is topologically similar to the target space will better align with the IB principle.
- TemporalUV: Capturing Loose Clothing with Temporally Coherent UV Coordinates6
A novel approach to generate temporally coherent UV coordinates for loose clothing by implementing a differentiable pipeline to learn UV mapping between a sequence of RGB inputs and textures via UV coordinates.
- InstructHumans: Editing Animated 3D Human Textures With Instructions5
This work shows that naively using SDS harms editing, as it may destroy consistency, and proposes a modified SDS for Editing (SDS-E) that selectively incorporates subterms of SDS across diffusion timesteps.
- 5
- Learning to generate training datasets for robust semantic segmentation5
This work designs Robusta, a novel robust conditional generative adversarial network to generate realistic and plausible perturbed images that can be used to train reliable segmentation models by leveraging the synergy between label-to-image generators and image-to-label segmentation models.
- Weakly-Supervised Dense Action Anticipation5
A framework that generates pseudo-labels for future actions and their durations and adaptively refines them through a refinement module is proposed and is competitive even when compared to fully supervised state-of-the-art models.
- Two-in-One Refinement for Interactive Segmentation5
A simple yet intuitive two-in-one re-nement strategy placing clicks on the boundary of the object of interest and a boundary-aware loss that encourages segmentation masks to respect instance boundaries are proposed.
- Supervised Deep Kriging for Single-Image Super-Resolution5
This work proposes a novel single-image super-resolution approach based on the geostatistical method of kriging that combines the krigs weight generation and kriged process into a joint network that can be learned end-to-end and achieves competitive super- resolution results as other state-of-the-art methods.
- Keep It in Mind: User Centric Continual Spatial Intelligence Reasoning in Egocentric Video Streams4
DirectMe is proposed, a framework that incrementally constructs and maintains a structured spatial memory from streaming egocentric observations that significantly improves the spatial reasoning of leading multimodal LLMs and surpasses many spatially aware and long-form streaming video models.
- Don't Pause! Every prediction matters in a streaming video4
AsynKV, a training-free streaming adaptation of offline models, that retains their event perception while improving their streaming behavior, serves as a strong baseline on SPOT-Bench, outperforming existing streaming models, and achieves state-of-the-art on retrospective benchmarks.
- Synthetic-to-Real Pose Estimation with Geometric Reconstruction4
This work proposes a reconstruction-based strategy as a complement to pseudo-labelling for synthetic-to-real domain adaptation and provides a novel solution to effectively correct confident yet inaccurate keypoint locations through image reconstruction in domain adaptation.
- 4
- Transformed ROIs for capturing visual transformations in videos4
TROI, a plug-and-play module for CNNs to reason between mid-level feature representations that are otherwise separated in space and time, achieves state-of-the-art action recognition results on the large-scale datasets Something-Something-V2 and EPIC-Kitchens-100.
- 4
- HANDS18: Methods, Techniques and Applications for Hand Observation4
The main conclusions that can be drawn are the turn of the community towards RGB data and the maturation of some methods and techniques, which in turn has led to increasing interest for real-world applications.
- Noise-Robust tiny object localization with flows3
This work proposes Tiny Object Localization with Flows (TOLF), a noise-robust localization framework leveraging normalizing flows for flexible error modeling and uncertainty-guided optimization, enabling robust learning under noisy supervision.
- 3
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes3
Semantic Gaussian Process Uncertainty (SGPU) is proposed, a Bayesian framework that quantifies semantic uncertainty by analyzing the geometric structure of answer embeddings, avoiding brittle clustering and demonstrating state-of-the-art calibration and discriminative performance.
- 3
- Learning Unorthogonalized Matrices for Rotation Estimation3
By replacing the orthogonalization incorporated representation with the proposed PRoM in various rotation-related tasks, this work achieves state-of-the-art results on large-scale benchmarks for human pose estimation.
- A Generalized & Robust Framework For Timestamp Supervision in Temporal Action Segmentation3
This work proposes a novel Expectation-Maximization (EM) based approach that leverages the label uncertainty of unlabelled frames and is robust enough to accommodate possible annotation errors and introduces the new challenging annotation setup of Skip-tag supervision.
- 3
- Fourier RNNs for Sequence Prediction3
This work proposes to integrate Fourier methods into complex recurrent neural network architectures and show accuracy improvements on prediction tasks as well as computational load reductions.
- Decouple and Cache: KV Cache Construction for Streaming Video Understanding2
Decoupled Streaming Cache is proposed, a training-free cache construction mechanism that adapts pretrained offline models to streaming settings that maintains a cumulative past KV cache while constructing a separate instant cache on-demand, decoupled from past caches to preserve the informativeness of recent inputs.
- 2
- 2
- 2
- 2
- 2
- Localized Interactive Instance Segmentation2
This work proposes a clicking scheme wherein user interactions are restricted to the proximity of the object, and a novel transformation of the user-provided clicks to generate a weak localization prior on the object which is consistent with image structures such as edges, textures etc.
- Fourier RNNs for Sequence Analysis and Prediction.2
This work proposes to integrate Fourier methods into complex recurrent neural network architectures and show accuracy improvements on analysis and prediction tasks as well as computational load reductions.
- 2
- RelaxFlow: Text-Driven Amodal 3D Generation1
This work formalizes text-driven amodal 3D generation, where text prompts steer the completion of unseen regions while strictly preserving input observation, and proposes RelaxFlow, a training-free dual-branch framework that decouples control granularity via a Multi-Prior Consensus Module and a Relaxation Mechanism.
- On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding1
GAD improves both accuracy and efficiency over generative methods, achieving state-of-the-art results on four tasks across five datasets, including an average 2.5% accuracy gain and 3x faster inference on the largest COIN benchmark.
- VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation1
Across class-conditional image generation with VAR, text-to-image generation with Infinity, and text-to-video generation with InfinityStar, VPG improves generation quality without retraining the base model, reducing FID on VAR by 0.36 on average and improving benchmark performance on both image and video generation.
- UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video1
UniCon3R is introduced, a unified feed-forward framework for online human-scene 4D reconstruction from monocular videos that outperforms state-of-the-art baselines on physical plausibility and global human motion estimation while achieving real-time online inference.
- Online Test-time Adaptation for 3D Human Pose Estimation: A Practical Perspective with Estimated 2D Poses1
This paper addresses adapting models to streaming videos with estimated 2D poses by proposing adaptive aggregation, a two-stage optimization, and local augmentation for handling varying levels of estimated pose error.
- Improving Deep Regression with Tightness1
This work reveals that preserving ordinality reduces the conditional entropy of representation Z conditional on the target Y, and introduces an optimal transport-based regularizer to preserve the similarity relationships of targets in the feature space to reduce H(Z|Y) .
- 1
- Computer Vision – ACCV 20241
This paper attempts to integrate an efficient clustering module into the principled framework for learning structured representation, in which the clustering module is used to provide partition information to guide the cluster-wise compression and the learned embeddings is aligned to desired geometric structures in turn to help for yielding more accurate partitions.
- 1
- 1
- 1
- 1
- Towards deep neural network compression via learnable wavelet transforms1
This paper shows how the fast wavelet transform can be used to compress linear layers in neural networks and has significantly fewer parameters yet still perform competitively with the state-of-the-art on synthetic and real-world RNN benchmarks.
- –
- TIGeR: Text-Instructed Generation and Refinement for Template-Free Hand-Object Interaction–
A new Text-Instructed Generation and Refinement (TIGeR) framework, harnessing the power of intuitive text-driven priors to steer the object shape refinement and pose estimation, which shows robustness to occlusion, while maintaining compatibility with heterogeneous prior sources.
- Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models–
Gate-and-Merge, a zero-shot framework that enables compositional personalization without the need for co-occurrence training, is introduced, and consistent gains in performance across multiple personalization tasks in both single-concept and compositional settings are shown.
- –
- –
- –
- –
- –
- –
- LightAVSeg: Lightweight Audio-Visual Segmentation–
LightAVSeg is proposed, a lightweight framework that replaces heavy attention with a decoupled design for semantic filtering and spatial grounding, resulting in interaction costs that scale linearly with spatial resolution.
- –
- ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration–
A canonical-space injection mechanism that decouples human dynamics from environmental cues via cross-attention is introduced and a novel positional embedding strategy that establishes spatial correspondences between disparate spatial domains without any heuristic 3D alignments is proposed.
- –
- –
- Video Question Answering and Beyond–
This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration.
- –
- –
- –
- LocalSR: Image Super-Resolution in Local Region–
Experimental results indicate that the proposed context-based local super-resolution (CLSR) approach, with its reduced low complexity, outperforms variants that focus exclusively on the ROI.
- High-Resolution Be Aware! Improving the Self-Supervised Real-World Super-Resolution–
A controller to adjust the degradation modeling based on the quality of super-resolution results is proposed and a novel feature-alignment regularizer is introduced that directly constrains the distribution of super-resolved images.
- –
- –
- –
- Workshop on Interactive and Adaptive Learning in an Open World–
This workshop at ECCV 2018 in Munich served as a discussion forum for experts in this field and in the following it is given a brief overview.
- –
- Data Driven Synthesis of Hand Grasps from 3-D Object Models–
This work forms grasp synthesis as a constrained optimization problem which takes into account the anthropomorphic and kinematic limitations of a human hand as well as the local and global geometric properties of the interacting object.
- Optimizing Over a Set of Manifolds ⋆–
This paper presents a comprehensive description of the algorithm that is proposed in “2D Action Recognition Serves 3D Human Pose Estimation”, which automates the very labor-intensive and therefore time-heavy and expensive process of human pose estimation.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.