Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks36 from public data
- 810
- BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework768
This work proposes a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods and is the first to handle realistic LiDar malfunction and can be deployed to realistic scenarios without any post-processing procedure.
- 419
- Knowledge Distillation via the Target-aware Transformer158
This work proposes a novel one-to-all spatial matching knowledge distillation approach that surpasses the state-of-the-art methods by a significant margin on various computer vision benchmarks, such as ImageNet, Pascal VOC and COCOStuff10k.
- BEVHeight: A Robust Framework for Vision-based Roadside 3D Object Detection143
This paper proposes a simple yet effective approach, dubbed BEVHeight, that regress the height to the ground to achieve a distance-agnostic formulation to ease the optimization process of camera-only perception methods.
- FusionAD: Multi-modality Fusion for Prediction and Planning Tasks of Autonomous Driving97
FusionAD is presented, to the best of the authors' knowledge, the first unified framework that fuse the information from two most critical sensors, camera and LiDAR, goes beyond perception task.
- Benchmarking the Robustness of LiDAR-Camera Fusion for 3D Object Detection91
The robust fusion dataset, benchmark, detailed documents and instructions are published and the effectiveness of the toolkit is showcased by establishing two novel robustness benchmarks on widely-adopted datasets, nuScenes and Waymo, then holistically evaluate the state-of-the-art fusion methods.
- FusionFormer: A Multi-sensory Fusion in Bird's-Eye-View and Temporal Consistent Transformer for 3D Object Detection43
This work proposes a novel end-to-end multi-modal fusion transformer-based framework, dubbed FusionFormer, that incorporates deformable attention and residual structures within the fusion encoding module that achieves state-of-the-art single model performance on a popular autonomous driving benchmark dataset, nuScenes.
- AlignMiF: Geometry-Aligned Multimodal Implicit Field for LiDAR-Camera Joint Synthesis25
This work conducts comprehensive analyses on the multimodal implicit field of LiDAR-camera joint synthesis, revealing the underlying issue lies in the misalignment of different sensors, and introduces AlignMiF, a geometrically aligned multimodal implicit field with two proposed modules: Geometry-Aware Alignment (GAA) and Shared Geometry Initialization (SGI).
- MiLA: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving20
This work proposes MiLA, a novel framework for generating high-fidelity, long-duration videos up to one minute, which utilizes a Coarse-to-Re(fine) approach to both stabilize video generation and correct distortion of dynamic objects.
- BioKGBench: A Knowledge Graph Checking Benchmark of AI Agent for Biomedical Science16
Surprisingly, it is discovered that state-of-the-art agents, both daily scenarios and biomedical ones, have either failed or inferior performance on the authors' benchmark.
- 15
- CRUISE: Cooperative Reconstruction and Editing in V2X Scenarios using Gaussian Splatting12
The experimental results demonstrate that CRUISE reconstructs real-world V2X driving scenes with high fidelity, and improves 3D detection across ego-vehicle, infrastructure, and cooperative views, as well as cooperative 3D tracking on the V2X-Seq benchmark; and CRUISE effectively generates challenging corner cases.
- 12
- Dyn-E: Local Appearance Editing of Dynamic Neural Radiance Fields11
A novel framework to edit the local appearance of dynamic NeRFs by manipulating pixels in a single frame of training video and introduces a local surface representation of the edited region, which can be inserted into and rendered along with the original NeRF and warped to arbitrary other frames through a learned invertible motion representation network.
- 10
- Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving7
This paper presents a comprehensive evaluation of five key LLM modules--tokenizer design, positional embedding, pre-training paradigms, post-training strategies, and test-time computation--within the context of motion generation for autonomous driving, and demonstrates that, when appropriately adapted, these modules can significantly improve performance for autonomous driving motion generation.
- $\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis6
This work introduces the first intelligent agent in this domain, poised to advance and potentially revolutionize the landscape of pharmaceutical research, which comprises three key end-to-end modules that perform patent question-answering, image-to-molecular-structure conversion, and core chemical structure identification.
- 6
- LiT: Unifying LiDAR "Languages" with LiDAR Translator4
The LiDAR Translator (LiT) is introduced, a framework that directly translates LiDAR data across domains, enabling both cross-domain adaptation and multi-domain joint learning.
- 3
- Real‐Time, Tissue‐Adaptive Photoacoustic Thermometry for Precision Endoscopic Thermal Therapy3
Simulated therapeutic experiments further confirm that PAET‐SAW‐guided monitoring reduces excessive thermal dose and suppresses overshoot, underscoring strong clinical potential for precision intravascular interventions.
- DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction3
This work introduces a Bird's-Eye View (BEV) based motion simulation method to model risks from three aspects: the ego-vehicle, other vehicles, and the environment, and designs a VLM-agnostic motion risk estimation framework, named DriveMRP-Agent.
- Virtual motionless photoacoustic microscopy for large-scale and high-resolution imaging based on K-Wave3
This paper uses the K-Wave simulation tool to build a single-pixel photoacoustic microscopic imaging simulation model, and uses this model to image blood vessels, demonstrating that this method can achieve a wide-field imaging with high-resolution.
- 2
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.