Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks17 from public data
- Words or Vision: Do Vision-Language Models Have Blind Faith in Text?103
The need for balanced training and careful consideration of modality interactions in VLMs to enhance their robustness and reliability in handling multi-modal data inconsistencies is highlighted, including supervised fine-tuning with text augmentation with its effectiveness in reducing text bias.
- KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection92
An automated knowledge collection pipeline is proposed, using which a large-scale multimodal brand knowledge base, KnowPhish, containing 20k brands with rich information about each brand is collected, which can be used to boost the performance of existing RBPDs in a plug-and-play manner.
- Anomaly Detection under Distribution Shift71
This paper introduces a novel robust AD approach to diverse distribution shifts by minimizing the distribution gap between in-distribution and OOD normal samples in both the training and inference stages in an unsupervised way and substantially outperforms state-of-the-art AD methods and OOD generalization methods on data with various distribution shifts.
- VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents47
It is shown that current CUAs and BUAs can be deceived at rates of up to 51% and 100%, respectively, on certain platforms, and the need for robust, context-aware defenses to ensure the safe deployment of multimodal AI agents in real-world environments is highlighted.
- MANet: Multi-branch attention auxiliary learning for lung nodule detection and segmentation29
A new UNet-based backbone with multi-branch attention auxiliary learning mechanism, which contains three novel modules, namely, Projection module, Fast Cascading Context module, and Boundary Enhancement module, to further enhance the nodule feature representation.
- Automating Steering for Safe Multimodal Large Language Models20
Experiments demonstrate that AutoSteer significantly reduces the Attack Success Rate (ASR) for textual, visual, and cross-modal threats, while maintaining general abilities, position AutoSteer as a practical, interpretable, and effective framework for safer deployment of multimodal AI systems.
- 18
- Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates17
This work introduces Just-In-Time Reinforcement Learning (JitRL), a training-free framework that enables test-time policy optimization without any gradient updates, and theoretically proves that this additive update rule is the exact closed-form solution to the KL-constrained policy optimization objective.
- 12
- 11
- 9
- PhishIntel: Toward Practical Deployment of Reference-Based Phishing Detection5
PhishIntel is presented, an end-to-end phishing detection system for real-world deployment that ensures low response latency while retaining the robust detection capabilities of RBPDs for zero-day phishing threats.
- WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents4
This work proposes a defense framework in which a web agent operates in parallel with a dedicated guard agent, decoupling prompt injection detection from the agent's own reasoning, and introduces WebAgentGuard, a reasoning-driven, multimodal guard model for prompt injection detection.
- CovHuSeg: An Enhanced Approach for Kidney Pathology Segmentation3
The effectiveness of the CovHuSeg algorithm is illustrated by experimenting with multiple deep-learning models in the context of segmentation on kidney pathology images, showing that all models have increased accuracy when using the CovHuSeg algorithm.
- Are Anomaly Scores Telling the Whole Story? A Benchmark for Multilevel Anomaly Detection3
A novel setting is proposed, Multilevel AD (MAD), in which the anomaly score represents the severity of anomalies in real-world applications, and a novel benchmark, MAD-Bench, is introduced that evaluates models not only on their ability to detect anomalies, but also on how effectively their anomaly scores reflect severity.
- 3
- MELCOT: A Hybrid Learning Architecture with Marginal Preservation for Matrix-Valued Regression–
This work proposes MELCOT, a hybrid model that integrates a classical machine–learning–based Marginal Estimation block with a deep-learning–based Learnable-Cost Optimal Transport (LCOT) block, which enables MELCOT to inherit the strengths of both classical and deep learning methods.
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.