Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks10 from public data
- Immersive Technologies-Driven Building Information Modeling (BIM) in the Context of Metaverse39
The macro-quantitative analysis on ImTs-driven BIM applications throughout all the stages of the building lifecycle reveals the themes, content, and characteristics of the applications across the stages, which tend to be integrated with emerging advanced technology and tools, such as Artificial Intelligence (AI), blockchain, and deep learning.
- LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA7
LaV-CoT is introduced, the first Language-aware Visual CoT framework with Multi-Aspect Reward Optimization, which adopts a two-stage training paradigm combining Supervised Fine-Tuning with Language-aware Group Relative Policy Optimization, guided by verifiable multi-aspect rewards including language consistency, structural accuracy, and semantic alignment.
- 4
- LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Multilingual Text-Centric VQA3
LaV-CoT is introduced, the first Language-aware Visual CoT framework with Multi-Aspect Reward Optimization, which outperforms open-source models of similar size by up to ~9.5% accuracy, even surpassing open-source models more than twice its size, and further exceeding several state-of-the-art proprietary models.
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis3
LogicLens is a unified framework for Visual-Textual Co-reasoning that reformulates these objectives into a joint task and establishes a significant lead over other MLLM-based methods in mF1, CSS, and the macro-average F1.
- 3
- 3
- –
- DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning–
This work proposes DocShield, the first unified framework formulating text-centric forgery analysis as a visual-logical co-reasoning problem, and introduces a Weighted Multi-Task Reward for GRPO-based optimization, aligning reasoning structure, spatial evidence, and authenticity prediction.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.