Is this you? Claim this profile to correct it, add a bio and choose the work people see first.

Claim this profile

Works8 from public data

TitleCited by
  • Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay B. Baiyya, +22 more

    International Journal of Computer Vision · 2025

    To push the frontier of first-person video understanding of skilled human activity, a suite of benchmark tasks and their annotations are presented, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose.

    640
  • VideoLLM-online: Online Video Large Language Model for Streaming Video

    Joya Chen, Zhaoyang Lv, Shiwei Wu, Kevin Qinghong Lin, Chenan Song, Difei Gao, Jia-Wei Liu, Ziteng Gao, +2 more

    IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings - IEEE Computer Society Conference on Computer Vision and Pattern Recognition/Proceedings · 2024

    A novel Learning-In- Video-Stream (LIVE) framework, which enables temporally aligned, long-context, and real-time dialogue within a continuous video stream within a continuous video stream.

    280
  • Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning

    Haiming Wang, Mert Unsal, Xiaohan Lin, Mantas Baksys, Junqi Liu, Marco Dos Santos, Flood Sung, Marina Vinyes, +22 more

    arXiv · 2025

    Kimina-Prover Preview is introduced, a large language model that pioneers a novel reasoning-driven exploration paradigm for formal theorem proving, as showcased in this preview release, and the learned reasoning style shows potential to bridge the gap between formal verification and informal mathematical intuition.

    200
  • Towards A Better Metric for Text-to-Video Generation

    Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, +6 more

    arXiv · 2024

    This paper investigates the limitations inherent in existing metrics and introduces a novel evaluation pipeline, the Text-to-Video Score (T2VScore), which integrates two pivotal criteria: Text-Video Alignment, which scrutinizes the fidelity of the video in representing the given text description, and Video Quality, which evaluates the video's overall production caliber with a mixture of experts.

    61
  • Sounding Video Generator: A Unified Framework for Text-Guided Sounding Video Generation

    Jia-Wei Liu, Weining Wang, Sihan Chen, Xinxin Zhu, Jing Liu

    IEEE Transactions on Multimedia · 2023

    This work proposes the Sounding Video Generator (SVG), a unified framework for generating realistic videos along with audio signals and presents the SVG-VQGAN, a novel hybrid contrastive learning method to model inter-modal and intra-modal consistency and improve the quantized representations.

    27
  • Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes

    Weifeng Liu, Tianyi She, Jia-Wei Liu, Weifeng Liu, Dongyu Yao, Ziyou Liang, Run Wang

    neural information processing systems · 2024

    18
  • MM21 Pre-training for Video Understanding Challenge

    Sihan Chen, Xinxin Zhu, Dongze Hao, Wei Liu, Jia-Wei Liu, Zijia Zhao, Longteng Guo, Jing Liu

    ACM International Conference on Multimedia (ACM MM) · 2021

    This paper proposes single-modality pretrained feature fusion technique which is composed of reasonable multi-view feature extraction method and designed multi- modality feature fusion strategy and it surpasses the state-of-the-art methods on both MSR-VTT and VATEX datasets.

    7
  • AI Performance on Image-based Medical Case Scenarios: A Cross-Sectional Comparative Study

    Jia-Wei Liu, Yue‐Tong Qian, Xiao Ma, Junping Fan, Lan-Wei Guo, Hongbo Yang

    Research Square · 2025

    –

Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.

Report an error

Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.