Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileAcademic lineage
Possible advisorsa guess from early papers, not confirmed
- Qing GuoSuggested from co-authorship
Is this you? Claim this profile to confirm or dismiss it.
Works14 from public data
- Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack27
This work proposes to generate AEs in the semantic image-text feature contrast space, which can project the original feature space into a semantic corpus subspace that can reduce the image feature redundancy, thereby improving adversarial transferability.
- 26
- Towards Transferable Attacks Against Vision-LLMs in Autonomous Driving with Typography22
This paper proposes a dataset-agnostic framework for automatically generating false answers that can mislead Vision-LLMs' reasoning, and presents a linguistic augmentation scheme that facilitates attacks at image-level and region-level reasoning, and extends it with attack patterns against multiple reasoning tasks simultaneously.
- Fast-dVLM: Efficient Block-Diffusion VLM via Direct Conversion from Autoregressive VLM8
A suite of multimodal diffusion adaptations, block size annealing, causal context attention, auto-truncation masking, and vision efficient concatenation, that collectively enable effective block diffusion in the VLM setting are introduced.
- Hyper3D: Efficient 3D Representation via Hybrid Triplane and Octree Feature for Enhanced 3D Shape Variational Auto-Encoders7
Hyper3D is introduced, which enhances VAE reconstruction through efficient 3D representation that integrates hybrid triplane and octree features and outperforms traditional representations by reconstructing 3D shapes with higher fidelity and finer details, making it well-suited for 3D generation pipelines.
- Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving5
This work presents Fast-dDrive, a block-diffusion VLA that performs bidirectional refinement within semantic units while enforcing strict causal ordering across them, and introduces Scaffold Speculative Decoding to achieve AR-equivalent quality at significantly higher throughput.
- Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective4
A proposed taxonomy organizes the research directions into five key problems that drive recent research development: feature enhancement, geometry awareness, model efficiency, augmentation strategies and temporal-aware models, and outlines future directions to address open challenges such as scalability, evaluation standards, and world modeling.
- PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space4
PixWorld is introduced, a single model that jointly addresses 3D reconstruction and generation that consistently outperforms prior latent-space generation methods and matches state-of-the-art reconstruction methods, demonstrating the superiority of a unified pixel-space approach.
- 1
- 1
- –
- –
- –
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.