Tianwei Zhang lists you as their PhD student. Claim this profile to confirm it and keep the rest of your record right.
Claim this profileAcademic lineage
View as a treeAdvisors
- Zhang TianweiFrom a thesis record ↗
- From a thesis record ↗
Works79 from public data
- Prompt Injection attack against LLM-integrated Applications926
This study deconstructs the complexities and implications of prompt injection attacks on actual LLM-integrated applications and forms HouYi, a novel black-box prompt injection attack technique, which draws inspiration from traditional web injection attacks.
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study738
The study underscores the importance of prompt structures in jailbreaking LLMs and discusses the challenges of robust jailbreak prompt generation and prevention.
- MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots275
Jailbreaker is presented, a comprehensive framework that offers an in-depth understanding of jailbreak attacks and countermeasures, and an automatic generation method for jailbreak prompts is introduced, leveraging a fine-tuned LLM to validate the potential of automated jailbreak generation across various commercial LLM chatbots.
- MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots206
A novel method that utilizes time-based characteristics intrinsic to the generation process to deconstruct the defense mechanisms employed by popular LLM chatbot services, and an innovative method for the automatic generation of jailbreak prompts that target robustly defended LLM chatbots.
- PentestGPT: An LLM-empowered Automatic Penetration Testing Tool204
PentestGPT is an LLM-empowered automatic penetration testing tool that leverages the abundant domain knowledge inherent in LLMs and proves effective in tackling real-world penetration testing challenges.
- A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models155
It is revealed that existing white-box attacks underperform compared to universal techniques and that including special tokens in the input significantly affects the likelihood of successful attacks.
- The Threat of Offensive AI to Organizations131
The threat of offensive AI on organizations is explored through a literature review and a user study, which identifies 33 offensive AI capabilities which adversaries can use to enhance their attacks.
- Automatic Code Summarization via ChatGPT: How Far Are We?120
Evaluating ChatGPT on a widely-used Python dataset called CSN-Python and comparing it with several state-of-the-art (SOTA) code summarization models shows that in terms of BLEU and ROUGE-L,ChatGPT's code summarizing performance is significantly worse than all three SOTA models.
- Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale114
The first large-scale empirical security analysis of this emerging ecosystem, collecting 42,447 skills from two major marketplaces and systematically analyzing 31,132 using SkillScan, a multi-stage detection framework integrating static analysis with LLM-based semantic classification reveals pervasive security risks.
- A Hitchhiker’s Guide to Jailbreaking ChatGPT via Prompt Engineering103
It was discovered that GPT-3.5 and GPT-4 could still generate inappropriate content in response to malicious prompts without the need for jailbreaking, underscores the critical need for effective prompt management within LLM systems and provides valuable insights and data to spur further research in LLM testing and jailbreak prevention.
- Morest103
This paper proposes Morest, a model-based RESTful API testing technique that builds and maintains a dynamically updating RESTful-service Property Graph (RPG) to model the behaviors of REST-services and guide the call sequence generation.
- 101
- 73
- 53
- SoK: Rethinking Sensor Spoofing Attacks against Robotic Vehicles from a Systematic View51
A novel action flow model is proposed to systematically describe robotic function executions and unexplored sensor spoofing threats and two novel attack methodologies are designed to verify the feasibility of newly discovered spoofing attack vectors.
- On the (In)Security of Secure ROS251
This paper successfully identifies four security vulnerabilities in ROS2's native security module: Secure ROS2 (SROS2), and proposes a general defense solution based on the private broadcast encryption scheme to enhance the security of ROS2.
- Novel denial-of-service attacks against cloud-based multi-robot systems46
By analyzing different attack vectors in cloud-robotic platforms, this paper proposes three new DoS attacks, which manipulate the network resources, micro-architecture resources, and function parameters respectively, and alerts the robotics community to these catastrophic attacks.
- Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems45
This work introduces Document-Driven Implicit Payload Execution (DDIPE), which embeds malicious logic in code examples and configuration templates within skill documentation within skill documentation, and generates 1,070 adversarial skills from 81 seeds across 15 MITRE ATTACK categories.
- Glitch Tokens in Large Language Models: Categorization Taxonomy and Effective Detection42
This study introduces and systematically explore the phenomenon of “glitch tokens”, which are anomalous tokens produced by established tokenizers and could potentially compromise the models’ quality of response, and proposes GlitchHunter, a novel iterative clustering-based technique for efficient glitch token detection.
- An Investigation of Byzantine Threats in Multi-Robot Systems41
An in-depth investigation about the Byzantine threats in MRSs, where some robot is untrusted is presented, and a practical methodology to identify potential Byzantine risks in a given MRS workload built from the Robot Operating System (ROS) is designed.
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild34
This work constructs the first labeled dataset of malicious agent skills by behaviorally verifying 98,380 skills from two community registries, confirming 157 malicious skills with 632 vulnerabilities.
- Oedipus: LLM-enchanced Reasoning CAPTCHA Solver34
Oedipus, an innovative end-to-end framework for automated reasoning CAPTCHA solving, is introduced with a novel strategy that dissects the complex and human-easy-AI-hard tasks into a sequence of simpler and AI-easy steps.
- Digger: Detecting Copyright Content Mis-usage in Large Language Model Training34
This paper introduces a detailed framework designed to detect and assess the presence of content from potentially copyrighted books within the training datasets of LLMs, and investigates the presence of recognizable quotes from famous literary works within these datasets.
- IllusionCAPTCHA: A CAPTCHA based on Visual Illusion27
In IllusionCAPTCHA, a novel security mechanism employing the "Human-Easy but AI-Hard" paradigm, a structured, step-by-step method that generates misleading options, which particularly guide LLMs towards making incorrect choices and reduce their chances of successfully solving CAPTCHAs.
- Efficient Detection of Toxic Prompts in Large Language Models25
ToxicDetector is proposed, a lightweight greybox method designed to efficiently detect toxic prompts in LLMs that achieves high accuracy, efficiency, and scalability, making it a practical method for toxic prompt detection in LLMs.
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models23
This work constructs GenderPair, a novel pair-based benchmark designed to assess gender bias in LLMs comprehensively, and establishes pioneering criteria for gender equality benchmarks, spanning dimensions such as inclusivity, diversity, explainability, objectivity, robustness, and realisticity.
- IRCopilot: Automated Incident Response with Large Language Models22
An incremental benchmark based on real-world incident response tasks is constructed based on Large Language Models to thoroughly evaluate the performance of LLMs in this domain and proposes IRCopilot, a novel framework for automated incident response powered by LLMs.
- Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment20
This work introduces Semantic-sensitive Alignment and Generation (SSAG), a method designed to systematically manipulate output-layer logits without altering model parameters that exposes harmful responses with a 95% success rate while reducing response time by 86%.
- Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment20
FiSAO (Fine-Grained Self-Alignment Optimization), a novel self-alignment method that utilizes the model's own visual encoder as a fine-grained verifier to improve vision-language alignment without the need for additional data, is proposed.
- MeTMaP: Metamorphic Testing for Detecting False Vector Matching Problems in LLM Augmented Generation18
MeTMaP is presented, a metamorphic testing framework developed to identify false vector matching in LLM -augmented generation systems, and the results em-phasize the widespread issue of false matches in vector matching methods and the critical need for effective detection and mitigation in LLM -augmented applications.
- 16
- PhyScout: Detecting Sensor Spoofing Attacks via Spatio-temporal Consistency16
Compared to existing defense solutions, PhyScout offers rapid identification of sensor attacks (within 100ms) with low performance overhead (CPU-based), and conflict visualization, and presents new avenues for future research in robust and efficient defense mechanisms against sensor spoofing attacks.
- Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation16
Groot is introduced, the first automated framework leveraging tree-based semantic transformation for adversarial testing of text-to-image models and achieves a remarkable success rate on leading text-to-image models such as DALL-E 3 and Midjourney.
- How Your Credentials Are Leaked by LLM Agent Skills: An Empirical Study15
The first large-scale empirical study on credential leakage in agent skills is presented, identifying 520 affected skills containing 1,708 security issues, and yields a taxonomy of 10 leakage patterns.
- PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design15
PentestEval is introduced, the first comprehensive benchmark for evaluating LLMs across six decomposed penetration testing stages: Information Collection, Weakness Gathering and Filtering, Attack Decision-Making, Exploit Generation and Revision, and Revision.
- Image-Based Geolocation Using Large Vision-Language Models15
It is revealed that LVLMs can accurately determine geolocations from images, even without explicit geographic training, and an innovative framework that significantly enhances image-based geolocation accuracy is introduced, called tool{}, an innovative framework that significantly enhances image-based geolocation accuracy.
- What Makes a Good LLM Agent for Real-world Penetration Testing?14
Excalibur is presented, a penetration testing agent that couples strong tooling with difficulty-aware planning and addresses a limitation that model scaling alone does not eliminate, showing that difficulty-aware planning yields consistent end-to-end gains across models.
- Overeager Coding Agents: Measuring Out-of-Scope Actions on Benign Tasks13
OverEager-Gen is presented, a benchmark dedicated to overeager behavior on benign tasks, and OverEager-Bench contains 500 validated scenarios and ~7,500 runs across four agent products and six base models, indicating that model-layer alignment does not fully propagate through permissive permission gating.
- Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning11
A novel defense strategy, Safety Chain-of-Thought (SCoT), is proposed, which harnesses the enhanced reasoning capabilities of LLMs for proactive assessment of harmful inputs, rather than simply blocking them.
- VisionGuard: Secure and Robust Visual Perception of Autonomous Vehicles in Practice10
The key of VisionGuard is to leverage the spatiotemporal inconsistency property of PAEs to detect anomalies and it predicts the motion states from historical ones and compares them with the current driving states to identify any motion inconsistency caused by physical attacks.
- GeneRAG: Enhancing Large Language Models with Gene-Related Task by Retrieval-Augmented Generation10
GeneRAG is introduced, a frame-work that enhances LLMs’ gene-related capabilities using RAG and the Maximal Marginal Relevance (MMR) algorithm, and its potential to bridge a critical gap in LLM capabilities for more effective applications in genetics is highlighted.
- 9
- Risky-Bench: Probing Agentic Safety Risks under Real-World Deployment8
Risky-Bench organizes evaluation around domain-agnostic safety principles to derive context-aware safety rubrics that delineate safety space, and systematically evaluates safety risks across this space through realistic task execution under varying threat assumptions.
- 8
- ASTER: Automatic Speech Recognition System Accessibility Testing for Stutterers6
Aster, a technique for automatically testing the accessibility of ASR systems, is implemented as a framework and evaluated, finding that it significantly increases the word error rate, match error rates, and word information loss in the evaluated AsR systems.
- VerifyML: Obliviously Checking Model Fairness Resilient to Malicious Model Holder5
The first secure inference framework to check the fairness degree of a given Machine learning (ML) model is presented, which allows the vast majority of overhead to be performed offline, thus meeting the low latency requirements for online inference.
- ExploitFlow, cyber security exploitation routes for Game Theory and AI research in robotics5
Results indicate that EF is effective for exploring machine learning in robot cybersecurity, and several limitations in EF-driven agents are identified, including a propensity to overfit, the scarcity and production cost of datasets for generalization, and challenges in interpreting networking states across varied security settings.
- Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution4
It is found that social credibility cues are the dominant driver of short-term behavioral influence, with misleading rates up to 61%, and routine memory-saving behavior can promote short-term pollution into durable long-term memory at rates up to 91%, with cross-session behavioral influence reaching 76%.
- A Survey of LLM-Driven Penetration Testing: Taxonomy, Co-Evolution, and Open Challenges3
A systematic analysis of 81 papers between 2023 and 2026 addresses gaps in a unified taxonomy, a systematic understanding of how agent architectures and evaluation benchmarks have co-evolved, and a clear characterization of remaining capability and reliability gaps.
- DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems3
DECEIVE-AFC is proposed, an agent-based adversarial attack framework that integrates novel claim-level attack strategies and adversarial claim validity evaluation principles that systematically explores adversarial attack trajectories that disrupt search behavior, evidence retrieval, and LLM-based reasoning without relying on access to evidence sources or model internals.
- Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges3
BITE is introduced, a black-box adversarial framework that learns semantics-preserving edits to mislead an LLM judge and artificially inflate the scores it assigns and exposes a fundamental weakness in the LLM-as-a-judge paradigm and motivates robust, attack-aware evaluation.
- When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents3
Results suggest that persistent memory can turn ordinary external processing into a practical pathway for long-term agent compromise.
- A Rusty Link in the AI Supply Chain: Detecting Evil Configurations in Model Repositories3
This work presents the first comprehensive study of malicious configurations on Hugging Face, identifying three attack scenarios (file, website, and repository operations) that expose inherent risks and introduces Configscan, an LLM-based tool that analyzes configuration files in the context of their associated runtime code and critical libraries, effectively detecting suspicious elements with low false positive rates and high accuracy.
- Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer2
It is demonstrated that shared neurons form an interpretable bridge between LLMs and LVLMs, enabling low-cost transfer of inference ability into multimodal models, and across diverse mathematics and perception benchmarks, SNRF consistently enhances LVLM inference performance while preserving perceptual capabilities.
- 2
- Mission: Impossible – Image-Based Geolocation with Large Vision Language Models2
This study investigates the geolocation capabilities of state-of-the-art LVLMs, and introduces ETHAN, a framework integrating chain-of-thought (CoT) reasoning that highlights the potential trajectory of such technologies rather than their current widespread, high-accuracy applicability.
- 2
- Continuous Embedding Attacks via Clipped Inputs in Jailbreaking Large Language Models2
Clip, whose main strategy is to clip each input dimension based on the mean and standard deviation of the model vocabulary during model inference, improves the attack success rate (ASR) of continuous embedding attacks with full LLM inputs from 62% to 83% for LLaMa and from 38% to 83 % for Vicuna.
- MIRAGE: Context-Aware Prompt Injection against Mobile GUI Agents via User-Generated Content1
MIRAGE (Mobile Injection of Realistic Adversarial GUI Examples), a pipeline that turns benign mobile screenshots into prompt-injection samples by placing attacker-controlled text into ordinary user-generated content regions, without modifying the agent, the application, or the operating system is presented.
- SNARE: Adaptive Scenario Synthesis for Eliciting Overeager Behavior in Coding Agents1
SNARE (Synthesizing Non-adversarial scenarios for Adaptive Reward-guided Elicitation), a pipeline that composes benign scenarios from reusable scope and trap fragments, scores each run with a judge-free oracle flagging trap-pattern matches and unsolicited file additions or deletions, and uses Thompson sampling to steer each pair's run budget toward the scenarios that most often trigger it.
- "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems1
This work presents the first large-scale empirical study with 303 participants to measure human susceptibility to Agent-Mediated Deception, and identifies six cognitive failure modes in users and finds that their risk awareness often fails to translate to protective behavior.
- AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications1
AutoEG, a fully automated multi-agent framework for exploit generation targeting black-box web applications, is proposed, substantially outperforming state-of-the-art baselines, whose best performance reaches only 32.88%.
- Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code–
This work presents HECATE, the first tool designed to assess complexity in both the prompt and code layers of LLM-integrated applications, and establishes prompt complexity as a dimension in its own right.
- Self-Guard: Defending Large Reasoning Models via enhanced self-reflection–
Self-Guard is proposed, a lightweight safety defense framework that reinforces safety compliance at the representational level and exhibits strong generalization across diverse unseen risks and varying model scales, offering a cost-efficient solution for LRM safety alignment.
- Membership Inference Attacks Against Video Large Language Models–
It is demonstrated that Video-LLMs are vulnerable to black-box membership inference attacks, highlighting an urgent need for the community to systematically evaluate and mitigate privacy risks in VideoLLMs.
- –
- –
- –
- SAF: An AI-Agent-Ready and Browser-Accessible Static Analysis Framework for LLVM IR–
The Static Analyzer Factory is presented, an LLVM-IR analysis framework that combines a Rust kernel exposed to Python through PyO3 zero-copy bindings, a WebAssembly and Pyodide browser playground that runs the same SDK with no install, a declarative YAML language for function specifications that the built-in checkers consume at runtime, and two shipped coding-agent skills installable in Claude Code and Codex.
- Robust CAPTCHA Using Audio Illusions in the Era of Large Language Models: from Evaluation to Advances–
AI-CAPTCHA is introduced, a unified framework that offers (i) an evaluation framework, ACEval, which includes advanced LALM- and ASR-based solvers, and (ii) a novel audio CAPTCHA approach, IllusionAudio, leveraging audio illusions.
- –
- –
- –
- –
- –
- Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems–
This paper constructs two psychophysics inspired benchmarks across distinct writing systems, Chinese logographs and English alphabetic words, by splicing, recombining, and overlaying glyphs to yield visible but unreadablestimuli for models while remaining legible to humans.
- SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System Testing–
SPOLRE is a novel automated tool designed for semantic preserving object layout reconstruction in image captioning system testing that utilizes four semantic preserving transformation techniques—translation, rotation, mirroring, and scaling—to modify object layouts autonomously, eliminating the need for manual annotation.
- –
- –
Publication data from OpenAlex, with missing venues and authors filled in from Crossref; citation counts are the higher of OpenAlex and Semantic Scholar, last synced 2026-10-10. One-sentence summaries under some papers are written by Semantic Scholar’s model. Citation counts may be lower than on Google Scholar, which indexes more sources.
Report an error
Wrong papers, two people merged into one, or a profile that should not be here? Tell us and we will fix or hide it. You will be asked to sign in.