Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Yanfeng WangSuggested from co-authorship
- Weidi XieSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works33 from public data
- 352
- Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion Models91
This work proposes a learning-based auto-regressive im-age generation model, termed as Story Gen, with a novel vision-language context module, that can generalize to unseen characters without any optimization, and generate image sequences with coherent content and consistent character.
- SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass58
This work presents SceneGen, a novel framework that takes a scene image and corresponding object masks as input, simultaneously producing multiple 3D assets with geometry and texture, and introduces a novel feature aggregation module that integrates local and global scene information from visual and geometric encoders within the feature extraction module.
- 46
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence38
This paper proposes SpatialScore, the most comprehensive and diverse multimodal spatial understanding benchmark to date, integrating VGBench with relevant data from the other 11 existing datasets, and develops SpatialAgent, a novel multi-agent system incorporating 9 specialized tools for spatial understanding.
- BabyVision: Visual Reasoning Beyond Language35
This work introduces BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for Multimodal LLMs, and explores solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit.
- Towards Universal Soccer Video Understanding33
An advanced soccer-specific visual encoder, MatchVision, is presented, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks, which highlights the superiority of the proposed data and model.
- MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further Tuning33
This paper introduces MegaFusion, a novel approach that extends existing diffusion-based text-to-image models towards efficient higher-resolution generation without additional fine-tuning or adaptation, and employs an innovative truncate and relay strategy to bridge the denoising processes across different resolutions.
- Multi-Agent System for Comprehensive Soccer Understanding23
This paper constructs SoccerWiki, the first large-scale multimodal soccer knowledge base, integrating rich domain knowledge about players, teams, referees, and venues to enable knowledge-driven reasoning and introduces SoccerAgent, a novel multi-agent system that decomposes complex soccer questions via collaborative reasoning.
- VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation19
VideoAutoArena is introduced, an arena-style benchmark inspired by LMSYS Chatbot Arena’s framework, designed to automatically assess LMMs’ video analysis abilities, and introduces a fault-driven evolution strategy, progressively increasing question complexity to push models toward handling more challenging video analysis scenarios.
- 19
- 19
- Dual-Branch Network for Portrait Image Quality Assessment16
A dual-branch network for portrait image quality assessment (PIQA) is introduced, which can effectively address how the salient person and the background of a portrait image influence its visual quality.
- VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results10
A novel benchmark comprising thousands of coarse-to-fine grained visual quality comparison tasks, spanning single images, pairs, and multi-image groups is introduced, which serves as a catalyst for future research on inter-pretable and human-aligned quality evaluation systems.
- MMRA: A Benchmark for Evaluating Multi-Granularity and Multi-Image Relational Association Capabilities in Large Visual Language Models9
These findings indicate that while LVLMs demonstrate a strong capability to perceive image details, enhancing their ability to associate information across multiple images hinges on improving the reasoning capabilities of their language model component.
- ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts9
Semantic-Temporal WAM (ST-WAM) is proposed to improve action robustness by using DINOv3 as a shared semantic representation for future prediction and history retrieval while retaining fine-grained VAE dynamics, demonstrating that semantic-temporal modeling effectively complements pixel-generative dynamics for robust manipulation.
- WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models5
WorldVQA is introduced, a benchmark designed to evaluate the atomic visual world knowledge of Multimodal Large Language Models, thereby establishing a standard for assessing the encyclopedic breadth and hallucination rates of current and next-generation frontier models.
- Towards Pixel-Level VLM Perception via Simple Points Prediction5
This work lays out that precise spatial understanding can emerge from simple point prediction, challenging the prevailing need for auxiliary components and paving the way for more unified and capable VLMs.
- PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models4
PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.
- Boost Video Frame Interpolation via Motion Adaptation4
This paper proposes a novel optimization-based VFI method that can adapt to unseen motions at test time, based on a cycle-consistency adaptation strategy that leverages the motion characteristics among video frames.
- MRGen: Segmentation Data Engine for Underrepresented MRI Modalities3
This paper investigates leveraging generative models to synthesize data, for training segmentation models for underrepresented modalities, particularly on annotation-scarce MRI, and believes that MRGen significantly improves segmentation performance on unannotated modalities by providing high-quality synthetic data.
- 3
- 2
- Improving Human Image Animation via Semantic Representation Alignment1
A novel approach named SemanticREPA is introduced that leverages these semantic representations as supervision signals through representation alignment to generate coherent and stable human structures and uses the predicted structure representations to refine identity restoration in relevant regions.
- Count Anything at Any Granularity1
This work redefine open-world counting as multi-grained counting, where visual exemplars specify target appearance and fine-grained text, with optional negative prompts, specifies the intended semantic granularity across five explicit levels.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.