Is this you? Claim this profile to correct it, add a bio and choose the work people see first.
Claim this profileWorks12 from public data
- AudioJailbreak: Jailbreak Attacks Against End-to-End Large Audio-Language Models22
This work proposes AudioJailbreak, a novel audio jailbreak attack featuring asynchrony, universality, stealthiness, and/or over-the-air robustness, and is applicable to a more practical and broader attack scenario where the adversary cannot fully manipulate user prompts (named weak adversary).
- 15
- AutoPrompt: Automated Red-Teaming of Text-to-Image Models via LLM-Driven Adversarial Prompts6
APT (AutoPrompT) is proposed, a black-box framework that leverages large language models (LLMs) to automatically generate humanreadable adversarial suffixes for benign prompts and introduces banned-token penalties to suppress the explicit generation of bannedtokens in blacklist.
- MPAS: Breaking Sequential Constraints of Multi-Agent Communication Topologies via Individual-Epistemic Message Propagation4
To overcome underlying shortcomings of sequential structures, this work proposes a node-wise multi-agent scheme, named message passing agent system (MPAS), and extends the message propagation mechanism in graph representation learning to multi-agent scenarios and introduces individual-epistemic message propagation.
- 1
- 1
- PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors1
PersGuard is introduced, a novel backdoor-based framework designed to prevent unauthorized personalization of pre-trained T2I diffusion models, and it is demonstrated that PersGuard provides superior privacy protection compared to existing perturbation-based methods.
- –
- –
- The Emotional Baby Is Truly Deadly: Does Your Multimodal Large Reasoning Model Have Emotional Flattery Towards Humans?–
A systematic security assessment of MLRMs is conducted and a autonomous adversarial emotion-agent is proposed that orchestrates exaggerated affective prompts to hijack reasoning pathways and reveal deeper emotional cognitive misalignments in model safety.
- PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems–
PhysPatch is proposed, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems that significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs.
- –
Publication data from OpenAlex; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-11. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.