Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks8 from public data
- Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives640
To push the frontier of first-person video understanding of skilled human activity, a suite of benchmark tasks and their annotations are presented, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose.
- VideoLLM-online: Online Video Large Language Model for Streaming Video280
A novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream within a continuous video stream.
- Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning200
Kimina-Prover Preview is introduced, a large language model that pioneers a novel reasoning-driven exploration paradigm for formal theorem proving, as showcased in this preview release, and the learned reasoning style shows potential to bridge the gap between formal verification and informal mathematical intuition.
- Towards A Better Metric for Text-to-Video Generation61
This paper investigates the limitations inherent in existing metrics and introduces a novel evaluation pipeline, the Text-to-Video Score (T2VScore), which integrates two pivotal criteria: Text-Video Alignment, which scrutinizes the fidelity of the video in representing the given text description, and Video Quality, which evaluates the video's overall production caliber with a mixture of experts.
- Sounding Video Generator: A Unified Framework for Text-Guided Sounding Video Generation27
This work proposes the Sounding Video Generator (SVG), a unified framework for generating realistic videos along with audio signals and presents the SVG-VQGAN, a novel hybrid contrastive learning method to model inter-modal and intra-modal consistency and improve the quantized representations.
- 18
- MM21 Pre-training for Video Understanding Challenge7
This paper proposes single-modality pretrained feature fusion technique which is composed of reasonable multi-view feature extraction method and designed multi- modality feature fusion strategy and it surpasses the state-of-the-art methods on both MSR-VTT and VATEX datasets.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.