Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks27 from public data
- OpenESS: Event-Based Semantic Scene Understanding with Open Vocabularies50
This work synergize information from image, text, and event-data domains and introduces OpenESS to enable scalable ESS in an open-world, annotation-efficient manner and proposes a frame-to-event contrastive distillation and a text-to-event semantic consistency regularization.
- WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World29
This work introduces WorldLens, a full-spectrum benchmark evaluating how well a model builds, understands, and behaves within its generated world, and develops WorldLens-Agent, an evaluation model distilled from these annotations to enable scalable, explainable scoring.
- The RoboDrive Challenge: Drive Anytime Anywhere in Any Condition28
This challenge has set a new benchmark in the field, providing a rich repository of techniques expected to guide future research in this field, and introduced a range of innovative approaches including advanced data augmentation, multi-sensor fusion, self-supervised learning for error correction, and new algorithmic strategies to enhance sensor robustness.
- Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees23
This work proposes SARD, a convex learning algorithm built on a family of surrogate losses that are provably Bayes-consistent and $(\mathcal{R}, \mathcal{G})$-consistent, and demonstrates that these guarantees hold across classification, regression, and multi-task settings.
- The RoboDepth Challenge: Methods and Advancements Towards Robust Depth Estimation19
Out of more than two hundred participants, nine unique and top-performing solutions have appeared, with novel designs ranging from the following aspects: spatial- and frequency-domain augmentations, masked image modeling, image restoration and super-resolution, adversarial training, diffusion-based noise suppression, vision-language pre-training, learned model ensembling, and hierarchical feature enhancement.
- EventFly: Event Camera Perception from Ground to the Sky16
This work introduces EventFly, a framework for robust cross-platform adaptation in event camera perception, and introduces EXPo, a large-scale benchmark with diverse samples across vehicle, drone, and quadruped platforms to holistically assess cross-platform adaptation abilities.
- 12
- AI for Auto-Research: Roadmap & User Guide11
It is shown that greater automation can obscure rather than eliminate failure modes, making human-governed collaboration the most credible deployment paradigm, and a structured taxonomy, benchmark suite, and tool inventory, cross-stage design principles, and a practitioner-oriented playbook are provided.
- Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer9
This work proposes a decoupled surrogate: a softmax classifier head and an independent sigmoid head per expert, mirroring the two natural objects of the problem, and proves an excess-risk bound with calibration constant, the first multi-expert L2D guarantee whose constant does not grow with the expert pool when the per-expert weight is held fixed.
- Online Learning-to-Defer with Varying Experts7
An online multiclass L2D algorithm that combines queried-action bandit feedback with a dynamically varying pool of experts is introduced that achieves expected true-deferral regret under a concentrated-score condition.
- Learning-to-Defer in Non-Stationary Time Series via Switching State-Space Models7
L2D-SLDS proves sublinear regret against a changing conditional-risk oracle without exploration, when the candidate models are accurate and either the archive separates them or live feedback reveals cost differences.
- Using Visual Intelligence to Automate Maintenance Task Guidance and Monitoring on a Head-mounted Display7
An Augmented Reality Visual Intelligence (ARVI) framework, which combines visual perception with cognitive task reasoning to monitor user performance and provide contextualised guidance for maintenance tasks, is presented.
- 5
- Is Your Driving World Model an All-Around Player?5
WorldLens is introduced, a unified benchmark that measures world-model fidelity across the full spectrum, from pixel quality and 4D geometry to closed-loop driving and human perceptual alignment, through five complementary aspects and 24 standardized dimensions and forms a unified ecosystem for assessing generated worlds not merely by visual appeal, but by physical and behavioral fidelity.
- EventDrive: Event Cameras for Vision-Language Driving Intelligence2
Comprehensive evaluation across diverse tasks shows that event streams provide substantial gains in temporal precision, motion awareness, and robustness, bringing event sensing into the center of driving intelligence.
- Learning to Remove Lens Flare in Event Camera2
E-Deflare is presented, the first systematic framework for removing lens flare from event camera data, and is designed to design E-DeflareNet, which achieves state-of-the-art restoration performance.
- Automated Multi-Camera Inspection System for Aircraft–
An automated visual inspection system designed to detect defects on the upper surface of an aircraft airframe using a multi-camera PTZ set-up to capture and process images from designated regions is presented.
- SpikeCLR: Contrastive Self-Supervised Learning for Few-Shot Event-Based Vision using Spiking Neural Networks–
SpikeCLR is introduced, a contrastive self-supervised learning framework that enables SNNs to learn robust visual representations from unlabeled event data and shows that learned representations transfer across datasets, contributing to efforts for powerful event-based models in label-scarce settings.
- –
- MemoVision: A Digital Catalog for Everyday Interactions–
MemoVision is presented, a digital catalog system that captures semantic, spatial, temporal and interaction information as users move around physical environments using client devices such as smart glasses, enabling more contextualized responses compared to current multimodal large language models.
- –
- –
- Interactive Content Retrieval in Egocentric Videos Based on Vague Semantic Queries–
This work investigates the requirements for an egocentric video content retrieval framework that helps users handle vague queries and proposes a zero-shot, user-centered video content retrieval framework that leverages a VLM to provide video data and query representations that users can incrementally combine to refine queries.
- –
- –
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it.