# AI Safety+Cybersecurity R&D Tracker

January 1, 2026

‍

[**Subscribe**](https://www.asenion.ai/subscribe-reg-tracker) **to our monthly AI Safety and Cybersecurity R&D Tracker updates!**

### Threats using AI models

#### Prompt Injection and Input Manipulation (Direct and Indirect)

Covers:

- OWASP LLM 01: Prompt Injection
- OWASP ML 01: Input Manipulation Attack
- MITRE ATLAS Initial Access, Privilege Escalation, and Defense Evasion

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Prompt Injection and Input Manipulation (Direct and Indirect)

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 1 | [Reference](https://www.google.com/url?q=http://Link&sa=D&source=editors&ust=1779048536614575&usg=AOvVaw1gBjI73gw1XCRQcGDCrRHh) |
| 2 | [▪ WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections (Cao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.15030&sa=D&source=editors&ust=1779048536614682&usg=AOvVaw3UI8BfdN8UkEvdBc7CPfvn) |
| 3 | [▪ Phantom Force: Injecting Adversarial Tactile Perceptions into Embodied Intelligence via EMI (Kong, Zhang, and Chau, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13492&sa=D&source=editors&ust=1779048536614780&usg=AOvVaw2RvcwLQWYHZZuRmqJzZLj3) |
| 4 | [▪ Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents (Maloyan and Namiot, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13471&sa=D&source=editors&ust=1779048536614866&usg=AOvVaw2ZBio3Ca0ln9o6VPxwwhO6) |
| 5 | [▪ Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery (Morasso et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12565&sa=D&source=editors&ust=1779048536614961&usg=AOvVaw2aS3Wwb4iR68n_9b_p1jxP) |
| 6 | [▪ Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation (Morasso et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11730&sa=D&source=editors&ust=1779048536615044&usg=AOvVaw2Cy7F4a9vCo0Rg3mYJQWCI) |
| 7 | [▪ IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection (Chia-Pei et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11868&sa=D&source=editors&ust=1779048536615120&usg=AOvVaw2ecCOXrZIqASCTjIQzhMAE) |
| 8 | [▪ Adversarial SQL Injection Generation with LLM-Based Architectures (Karakoc and Yilmaz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11188&sa=D&source=editors&ust=1779048536615203&usg=AOvVaw1M9EOmvcasgPBkRJdbK6Bt) |
| 9 | [▪ MCPShield: Content-Aware Attack Detection for LLM Agent Tool-Call Traffic (Zavrak, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11053&sa=D&source=editors&ust=1779048536615308&usg=AOvVaw3Y-QCccT3kYugTX1gjYBfA) |
| 10 | [▪ SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20129&sa=D&source=editors&ust=1779048536615392&usg=AOvVaw0S8ak8PxOJ6_cvNFKBWwnu) |
| 11 | [▪ Preventing Prompt Injection with Type-Directed Privilege Separation (Jacob et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25926&sa=D&source=editors&ust=1779048536615541&usg=AOvVaw3WhlRQUvByYmygX_8JiHCV) |
| 12 | [▪ When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications (Motlagh et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10176&sa=D&source=editors&ust=1779048536615653&usg=AOvVaw0jbpO5TU9qr3m4W5ja-h2z) |
| 13 | [▪ Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10133&sa=D&source=editors&ust=1779048536615753&usg=AOvVaw1rXt2DWPBVEfiCCrXN0Ip-) |
| 14 | [▪ Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs (Maiorano, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06669&sa=D&source=editors&ust=1779048536615845&usg=AOvVaw1mW5D9VGME4zDJPxcTVx5x) |
| 15 | [▪ One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue (Shen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05630&sa=D&source=editors&ust=1779048536615921&usg=AOvVaw2ECZ_r9OslWHAGLZXS55eg) |
| 16 | [▪ PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05974&sa=D&source=editors&ust=1779048536616001&usg=AOvVaw27PKJvA-6sl4u7lut4QI08) |
| 17 | [▪ Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs (Fei et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04446&sa=D&source=editors&ust=1779048536616081&usg=AOvVaw0NXO8gkPtGmL0QCd40vhaM) |
| 18 | [▪ Tailored Prompts, Targeted Protection: Vulnerability-Specific LLM Analysis for Smart Contracts (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03697&sa=D&source=editors&ust=1779048536616175&usg=AOvVaw1mwBiBIeOUZD_aHTPYqDAE) |
| 19 | [▪ Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis (Zhang, Zandsalimy, and Sushmita, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03441&sa=D&source=editors&ust=1779048536616248&usg=AOvVaw0vfYKZX4gKsovzeGalJv9O) |
| 20 | [▪ A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts (Young and Moody, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03179&sa=D&source=editors&ust=1779048536616346&usg=AOvVaw2CSutKzq8X0fjVdn3SvN3P) |
| 21 | [▪ LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training (Gong et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01462&sa=D&source=editors&ust=1779048536616482&usg=AOvVaw1vc8h8bxCfPPYrEa20i3NV) |
| 22 | [▪ A Sentence Relation-Based Approach to Sanitizing Malicious Instructions (Datta et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01078&sa=D&source=editors&ust=1779048536616596&usg=AOvVaw0P4mzyDTqkRSfXL9gJ3uBC) |
| 23 | [▪ Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI (Chang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.19143&sa=D&source=editors&ust=1779048536616715&usg=AOvVaw2EsFhOKQy-m3QfUx-nZEzH) |
| 24 | [▪ FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.28157&sa=D&source=editors&ust=1779048536616844&usg=AOvVaw3FcnWEJyq_Y-um8aHti1L7) |
| 25 | [▪ Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives (Khodayari et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27202&sa=D&source=editors&ust=1779048536616977&usg=AOvVaw3HIVsW228NtGhh9S-nNcs7) |
| 26 | [▪ SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts (Xin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26506&sa=D&source=editors&ust=1779048536617104&usg=AOvVaw2v0GMcijTfT4H_Eq6-iZt6) |
| 27 | [▪ "Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22040&sa=D&source=editors&ust=1779048536617241&usg=AOvVaw04TBQiDYYScOSJ7ESPQPUi) |
| 28 | [▪ SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents (Du et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25562&sa=D&source=editors&ust=1779048536617371&usg=AOvVaw3wmuyTQNxqN_7WRFaJBDMe) |
| 29 | [▪ PARASITE: Conditional System Prompt Poisoning to Hijack LLMs (Pham and Le, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16888&sa=D&source=editors&ust=1779048536617500&usg=AOvVaw30X893LaApDv0T7DmUSSX6) |
| 30 | [▪ Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review (Mitropoulos et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18740&sa=D&source=editors&ust=1779048536617616&usg=AOvVaw2zDIdq3MCUAJ5QqBarFsPr) |
| 31 | [▪ Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models (Rayhan and Jahan, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21860&sa=D&source=editors&ust=1779048536617750&usg=AOvVaw3hAuEC4zWoJb4tnRd-jOl-) |
| 32 | [▪ Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18663&sa=D&source=editors&ust=1779048536617868&usg=AOvVaw1BPsvSsr6a_f6EdLgEzKp8) |
| 33 | [▪ XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants (\\v{S}torek et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.14281&sa=D&source=editors&ust=1779048536617987&usg=AOvVaw3D7zRJH2khpd0PPGdnyOhp) |
| 34 | [▪ Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection (Munirathinam, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18248&sa=D&source=editors&ust=1779048536618102&usg=AOvVaw3VCPCddPrXT0ILNF4MJnYa) |
| 35 | [▪ CASCADE: A Cascaded Hybrid Defense Architecture for Prompt Injection Detection in MCP-Based Systems (Turgut and G\\"um\\"u\\c{s}, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17125&sa=D&source=editors&ust=1779048536618204&usg=AOvVaw35oHtxUFOgUJBCmmpdesdV) |
| 36 | [▪ LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents (Shah, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15368&sa=D&source=editors&ust=1779048536618300&usg=AOvVaw30x4RreO_g8JdBCCHFQyoo) |
| 37 | [▪ Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14604&sa=D&source=editors&ust=1779048536618393&usg=AOvVaw2asim7YxF44YJkpfWBgHLt) |
| 38 | [▪ Random Walk Learning and the Pac-Man Attack (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.05663&sa=D&source=editors&ust=1779048536618490&usg=AOvVaw1_zhaNNfuMmctW2mHJDotB) |
| 39 | [▪ Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities (Frazier and Damevski, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13955&sa=D&source=editors&ust=1779048536618599&usg=AOvVaw09EV08N_I6EIfxbP_Mg81q) |
| 40 | [▪ LLM-Guided Prompt Evolution for Password Guessing (Mazin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12601&sa=D&source=editors&ust=1779048536618718&usg=AOvVaw3KJhDlbkHOcKufHr3t4A_U) |
| 41 | [▪ DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection (Ren et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12548&sa=D&source=editors&ust=1779048536618836&usg=AOvVaw2_EHD96gdgVMpSb9vg1_Bk) |
| 42 | [▪ WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12284&sa=D&source=editors&ust=1779048536618917&usg=AOvVaw1MOxJfJS-vT_pL06QllQ4R) |
| 43 | [▪ AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.03793&sa=D&source=editors&ust=1779048536618991&usg=AOvVaw0mNOGG3gre3toW-1OEcPEG) |
| 44 | [▪ ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11790&sa=D&source=editors&ust=1779048536619060&usg=AOvVaw0vDWEgUlb5i8R9lvMoeSUY) |
| 45 | [▪ Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection (Shao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09024&sa=D&source=editors&ust=1779048536619129&usg=AOvVaw22JyopfMuUfCXOPF-m3j-d) |
| 46 | [▪ ACIArena: Toward Unified Evaluation for Agent Cascading Injection (An et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07775&sa=D&source=editors&ust=1779048536619196&usg=AOvVaw1QUfZNj0wuaB_C5y5EzQL9) |
| 47 | [▪ PIArena: A Platform for Prompt Injection Evaluation (Geng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08499&sa=D&source=editors&ust=1779048536619261&usg=AOvVaw0cuftH3Bz6jaLkVU4a7U6b) |
| 48 | [▪ Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07831&sa=D&source=editors&ust=1779048536619330&usg=AOvVaw0nQlCP5EXWfiQk_8xwfIEw) |
| 49 | [▪ Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03753&sa=D&source=editors&ust=1779048536619396&usg=AOvVaw3AeP5nTGMvxTDNXNPrVJU9) |
| 50 | [▪ AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models (Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03598&sa=D&source=editors&ust=1779048536619471&usg=AOvVaw2A82N29Q763akEB88BLdf2) |
| 51 | [▪ AgentWatcher: A Rule-based Prompt Injection Monitor (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01194&sa=D&source=editors&ust=1779048536619538&usg=AOvVaw2vXQBhH-RxwWIcDrvXrxI0) |
| 52 | [▪ Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks (Collu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20863&sa=D&source=editors&ust=1779048536619608&usg=AOvVaw3BlPbFtgO3r0Z4zzxyzgoD) |
| 53 | [▪ Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers (Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28013&sa=D&source=editors&ust=1779048536619675&usg=AOvVaw1ZXozGutuWMVD_2FarJ4wG) |
| 54 | [▪ Epistemic Bias Injection: Biasing LLMs via Selective Context Retrieval (Wu and Saxena, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00804&sa=D&source=editors&ust=1779048536619740&usg=AOvVaw0UmiLCQWPGTVZe6iNQGzmn) |
| 55 | [▪ PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25164&sa=D&source=editors&ust=1779048536619817&usg=AOvVaw2SNF96tRd_9wJUKzIKOSjI) |
| 56 | [▪ SUAD: Solid-Channel Ultrasound Injection Attack and Defense to Voice Assistants (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02116&sa=D&source=editors&ust=1779048536619908&usg=AOvVaw0Tuw4KdZgJNBLWbzNwPdeQ) |
| 57 | [▪ Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search (Shen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24203&sa=D&source=editors&ust=1779048536620005&usg=AOvVaw20C9FKWwsTUVSssn2UTqYb) |
| 58 | [▪ The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense (Lan and Kaul, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23791&sa=D&source=editors&ust=1779048536620091&usg=AOvVaw3Hr9xGmijcizjuZWwVlIz2) |
| 59 | [▪ Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22489&sa=D&source=editors&ust=1779048536620162&usg=AOvVaw0ieLyJ5_TQ82xkk8mcfnDo) |
| 60 | [▪ Are AI-assisted Development Tools Immune to Prompt Injection? (Huang, Huang, and Fard, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21642&sa=D&source=editors&ust=1779048536620227&usg=AOvVaw0ZiTyx3_yzpfHZzUuJANGg) |
| 61 | [▪ Cross-site scripting adversarial attacks based on deep reinforcement learning: Evaluation and extension study (Pasini et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.19095&sa=D&source=editors&ust=1779048536620293&usg=AOvVaw3wveOIkY9CwXyDrdy0Fd3b) |
| 62 | [▪ Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems (Alam et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18433&sa=D&source=editors&ust=1779048536620360&usg=AOvVaw32kHdV2odt6S4WUF-OGKoR) |
| 63 | [▪ Detecting Sentiment Steering Attacks on RAG-enabled Large Language Models (Andrade et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16342&sa=D&source=editors&ust=1779048536620439&usg=AOvVaw3RQldGN-JxeyN3BH_vB_sB) |
| 64 | [▪ How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition (Dziemian et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15714&sa=D&source=editors&ust=1779048536620509&usg=AOvVaw0TLpK2IK__3o6NGB6C46ef) |
| 65 | [▪ SKILLS: Structured Knowledge Injection for LLM-Driven Telecommunications Operations (Brett, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15372&sa=D&source=editors&ust=1779048536620576&usg=AOvVaw3_YrF7okstwPvoJnB8Zdcm) |
| 66 | [▪ Agent Privilege Separation in OpenClaw: A Structural Defense Against Prompt Injection (Cheng and Tsao, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13424&sa=D&source=editors&ust=1779048536620652&usg=AOvVaw2h1II1dCFJkF_HDCbsXRT-) |
| 67 | [▪ PILOT: Command-line Interface Fuzzing via Path-Guided, Iterative Large Language Model Prompting (Shiraishi, Cao, and Shinagawa, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20555&sa=D&source=editors&ust=1779048536620721&usg=AOvVaw1NxWokyIf0pNsCQoT048AF) |
| 68 | [▪ PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses (Yin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13026&sa=D&source=editors&ust=1779048536620785&usg=AOvVaw3PFkUqlkDENNLBNBiGjmbS) |
| 69 | [▪ Prompt Injection as Role Confusion (Ye, Cui, and Hadfield-Menell, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12277&sa=D&source=editors&ust=1779048536620849&usg=AOvVaw3_kZ54FTg82Ps4k1Kyz7ZF) |
| 70 | [▪ The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detection (Corll, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11875&sa=D&source=editors&ust=1779048536620918&usg=AOvVaw0676Q04L9CzT7-tbbIw_3b) |
| 71 | [▪ AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations (He et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10749&sa=D&source=editors&ust=1779048536621003&usg=AOvVaw1u7Z5rV9N3t3Oi4QNqIg0n) |
| 72 | [▪ Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions (Nagaraja et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03637&sa=D&source=editors&ust=1779048536621070&usg=AOvVaw0WOWnSASrW9scschZxhhqw) |
| 73 | [▪ VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents (Cao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.02456&sa=D&source=editors&ust=1779048536621135&usg=AOvVaw0OcPOZnqTmZwG18ToVSQNc) |
| 74 | [▪ Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection (Graves, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00164&sa=D&source=editors&ust=1779048536621201&usg=AOvVaw1kUDm5AiQa9tcDTws_PXz2) |
| 75 | [▪ AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22724&sa=D&source=editors&ust=1779048536621271&usg=AOvVaw2cvWQBXm_xFY3GxUugszsX) |
| 76 | [▪ Silent Egress: When Implicit Prompt Injection Makes LLM Agents Leak Without a Trace (Lan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22450&sa=D&source=editors&ust=1779048536621354&usg=AOvVaw2X9AReCSc1rhuSJwpW_KK8) |
| 77 | [▪ ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20708&sa=D&source=editors&ust=1779048536621427&usg=AOvVaw3zTmWc4YoXDUS92gZRMIBE) |
| 78 | [▪ AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20720&sa=D&source=editors&ust=1779048536621493&usg=AOvVaw2VUfBeHly4JOeyLG07g6g-) |
| 79 | [▪ ICSSPulse: A Modular LLM-Assisted Platform for Industrial Control System Penetration Testing (Takaronis et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20663&sa=D&source=editors&ust=1779048536621588&usg=AOvVaw1GrpIa9hLeuXUgxkCLxoLh) |
| 80 | [▪ Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks (Schmotz et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20156&sa=D&source=editors&ust=1779048536621655&usg=AOvVaw3d1jcGpZG4tZoJbBeBOAwr) |
| 81 | [▪ Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models (Wirth, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18514&sa=D&source=editors&ust=1779048536621726&usg=AOvVaw1BSEhtqnj8lc66bA7b1v3H) |
| 82 | [▪ The Vulnerability of LLM Rankers to Prompt Injection Attacks (Yin et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16752&sa=D&source=editors&ust=1779048536621792&usg=AOvVaw0hOypXxnUT1ruK9EGIq8lW) |
| 83 | [▪ Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis (Thornton, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16741&sa=D&source=editors&ust=1779048536621863&usg=AOvVaw2vEOOH51NcbVdLILeAVyIM) |
| 84 | [▪ SkillJect: Automating Stealthy Skill-Based Prompt Injection for Coding Agents with Trace-Driven Closed-Loop Refinement (Jia et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14211&sa=D&source=editors&ust=1779048536621936&usg=AOvVaw2eOLnt6K8zBTyoSOcjo6ia) |
| 85 | [▪ AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks (Jia et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13597&sa=D&source=editors&ust=1779048536622004&usg=AOvVaw3QNhnbtehK4PvO7ciLtFa6) |
| 86 | [▪ SAFuzz: Semantic-Guided Adaptive Fuzzing for LLM-Generated Code (Yang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11209&sa=D&source=editors&ust=1779048536622070&usg=AOvVaw2DL6iEN-ZI7TwWVg2I2SiY) |
| 87 | [▪ When Skills Lie: Hidden-Comment Injection in LLM Agents (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10498&sa=D&source=editors&ust=1779048536622149&usg=AOvVaw10gj_hRMYzwBQekBSbb0Wv) |
| 88 | [▪ Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI (Rajagopalan and Rao, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10481&sa=D&source=editors&ust=1779048536622217&usg=AOvVaw1mYoP-iigH6YyFAs2VKAX3) |
| 89 | [▪ The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10453&sa=D&source=editors&ust=1779048536622289&usg=AOvVaw1heK-UNXYkAsNzgaJrreuD) |
| 90 | [▪ The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism (Brodt et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09625&sa=D&source=editors&ust=1779048536622367&usg=AOvVaw3V9gUbMcU1M4ZsH1O8bLgt) |
| 91 | [▪ MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks (Syros et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09222&sa=D&source=editors&ust=1779048536622442&usg=AOvVaw18utHPq6RrM_BpxfbCu7au) |
| 92 | [▪ Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation (Hassan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08062&sa=D&source=editors&ust=1779048536622509&usg=AOvVaw1VircipzVYaUFA8HQ4v0hY) |
| 93 | [▪ CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution (Kim et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07918&sa=D&source=editors&ust=1779048536622576&usg=AOvVaw0L3_8osxH7qb3_5f9QQvnc) |
| 94 | [▪ Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection (Koide, Nakano, and Chiba, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05484&sa=D&source=editors&ust=1779048536622644&usg=AOvVaw2nFf2p9E84LBhJFl3YoLcR) |
| 95 | [▪ WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03792&sa=D&source=editors&ust=1779048536622729&usg=AOvVaw3PDTbUnfPeibTS6nFE6RZk) |
| 96 | [▪ AgentDyn: A Dynamic Open-Ended Benchmark for Evaluating Prompt Injection Attacks of Real-World Agent Security System (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03117&sa=D&source=editors&ust=1779048536622812&usg=AOvVaw2qnlEJaLHAj3eCYLzkhNxq) |
| 97 | [▪ Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11250&sa=D&source=editors&ust=1779048536622883&usg=AOvVaw2YR8L-a5W9xtHcAV4nZuMU) |
| 98 | [▪ RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01795&sa=D&source=editors&ust=1779048536622957&usg=AOvVaw2Hx7MCdR3NUkd99BXERwZK) |
| 99 | [▪ Bypassing Prompt Injection Detectors through Evasive Injections (Rahman and Alouani, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00750&sa=D&source=editors&ust=1779048536623026&usg=AOvVaw1CAZq9sEprFll6bG-CUTog) |
| 100 | [▪ zkCraft: Prompt-Guided LLM as a Zero-Shot Mutation Pattern Oracle for TCCT-Powered ZK Fuzzing (Fu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00667&sa=D&source=editors&ust=1779048536623095&usg=AOvVaw2Vu2kj0HmQaEN9gk5B9oKy) |
| 101 | [▪ Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection (Debi and Zhu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22569&sa=D&source=editors&ust=1779048536623164&usg=AOvVaw0k873EQrOEbJxJyh_oIuSW) |
| 102 | [▪ WADBERT: Dual-channel Web Attack Detection Based on BERT Models (Luo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21893&sa=D&source=editors&ust=1779048536623230&usg=AOvVaw3p2EmhSxvmfHLBmycrvC6v) |
| 103 | [▪ Analysis of LLM Vulnerability to GPU Soft Errors: An Instruction-Level Fault Injection Study (Chai et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19912&sa=D&source=editors&ust=1779048536623298&usg=AOvVaw3blNSJkeo-LIdWEVcWnL9p) |
| 104 | [▪ MalURLBench: A Benchmark Evaluating Agents' Vulnerabilities When Processing Web URLs (Kong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.18113&sa=D&source=editors&ust=1779048536623367&usg=AOvVaw2oLcghR4u0qKA4MWarxuiw) |
| 105 | [▪ Prompt Injection Evaluations: Refusal Boundary Instability and Artifact-Dependent Compliance in GPT-4-Series Models (Heverin, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17911&sa=D&source=editors&ust=1779048536623480&usg=AOvVaw00GfLxFlRs1Lt8GwLKB04C) |
| 106 | [▪ Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents (Maloyan and Namiot, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17549&sa=D&source=editors&ust=1779048536623570&usg=AOvVaw3qRv5tHhKlV7F3Rg70ewf_) |
| 107 | [▪ Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems (Maloyan and Namiot, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17548&sa=D&source=editors&ust=1779048536623645&usg=AOvVaw0GjzP0Wsoy8NBu-PlcRc2Y) |
| 108 | [▪ Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art (Trinh et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17379&sa=D&source=editors&ust=1779048536623719&usg=AOvVaw3F-hm1vmC9eELGb6PWnT39) |
| 109 | [▪ On the Insecurity of Keystroke-Based AI Authorship Detection: Timing-Forgery Attacks Against Motor-Signal Verification (Condrey, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17280&sa=D&source=editors&ust=1779048536623796&usg=AOvVaw1Kxa-EYfTnihAvoDXa8MeE) |
| 110 | [▪ RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models (Chugh, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15331&sa=D&source=editors&ust=1779048536623863&usg=AOvVaw03qkDwmoxclu8texoK1_CO) |
| 111 | [▪ SPECTRE: Conditional System Prompt Poisoning to Hijack LLMs (Pham and Le, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16888&sa=D&source=editors&ust=1779048536623937&usg=AOvVaw0EymXBuO2tuQ87Cs5KHkTE) |
| 112 | [▪ LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation (Lazo, Jelodar, and Razavi-Far, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14528&sa=D&source=editors&ust=1779048536624007&usg=AOvVaw0WogOyj8CJYyT3DylUkFH4) |
| 113 | [▪ Sockpuppetting: Jailbreaking LLMs Without Optimization Through Output Prefix Injection (Dotsinski and Eustratiadis, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13359&sa=D&source=editors&ust=1779048536624075&usg=AOvVaw1WO1mz1Bnjv5fDKAfnGx-3) |
| 114 | [▪ PINA: Prompt Injection Attack against Navigation Agents (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13612&sa=D&source=editors&ust=1779048536624140&usg=AOvVaw3xgJQAcluX7FJ62YvBHU0y) |
| 115 | [▪ Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading (Rizvani, Apruzzese, and Laskov, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13082&sa=D&source=editors&ust=1779048536624209&usg=AOvVaw2y9QM09DNlSxAqgRLvQjIn) |
| 116 | [▪ Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs (Sekar et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12359&sa=D&source=editors&ust=1779048536624277&usg=AOvVaw25uG7r7BLBtpSuSoTW-tqq) |
| 117 | [▪ SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation (Masoud, Arazzi, and Nocera, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11199&sa=D&source=editors&ust=1779048536624343&usg=AOvVaw3Ov87YAr6ufbXWLJ1pUM9q) |
| 118 | [▪ Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG (Guo and Wei, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10923&sa=D&source=editors&ust=1779048536624426&usg=AOvVaw3ScM2zizS5-DoFH73asM5T) |
| 119 | [▪ Reasoning Hijacking: Subverting LLM Classification via Decision-Criteria Injection (Liu, Tang, and Tun, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10294&sa=D&source=editors&ust=1779048536624498&usg=AOvVaw2HU_lVPpNoMj64YTmv1_Kn) |
| 120 | [▪ ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10173&sa=D&source=editors&ust=1779048536624564&usg=AOvVaw1nYPl_ZUNs_fTZ3Sf6YTSh) |
| 121 | [▪ The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multi-Step Malware (Nassi, Schneier, and Brodt, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09625&sa=D&source=editors&ust=1779048536624632&usg=AOvVaw3qo-_M9Duks_FIldtaNYAB) |
| 122 | [▪ Baiting AI: Deceptive Adversary Against AI-Protected Industrial Infrastructures (Pasikhani et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08481&sa=D&source=editors&ust=1779048536624698&usg=AOvVaw2LEXdviqeoGexTnj4pt6pB) |
| 123 | [▪ Small Symbols, Big Risks: Exploring Emoticon Semantic Confusion in Large Language Models (Jiang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07885&sa=D&source=editors&ust=1779048536624767&usg=AOvVaw1rOA8NnDi6CFBEXn5W7v7u) |
| 124 | [▪ Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning (Nguyen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06641&sa=D&source=editors&ust=1779048536624833&usg=AOvVaw0OQx-KCQe-KUdS3Fsa9Gb6) |
| 125 | [▪ SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations (Ali et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07835&sa=D&source=editors&ust=1779048536624908&usg=AOvVaw2CS0Rch_pHpcdQPwcv-Rek) |
| 126 | [▪ How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test (Tessa et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07084&sa=D&source=editors&ust=1779048536624982&usg=AOvVaw2oaRKcJgb1MaSjmbAppECk) |
| 127 | [▪ Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems (Chang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07072&sa=D&source=editors&ust=1779048536625051&usg=AOvVaw0ShJ0utUvAlpqL5UWwuDuL) |
| 128 | [▪ Defense Against Indirect Prompt Injection via Tool Result Parsing (Yu, Cheng, and Liu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04795&sa=D&source=editors&ust=1779048536625118&usg=AOvVaw2wLkFtPakMuXaQhI3rGkOk) |
| 129 | [▪ Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning (Chang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04666&sa=D&source=editors&ust=1779048536625186&usg=AOvVaw2CdaOwB3l4gSll3z_DEuyT) |
| 130 | [▪ Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models (Gong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.01386&sa=D&source=editors&ust=1779048536625255&usg=AOvVaw1t0gctWt6KBFeWEGO2aaNV) |
| 131 | [▪ Prompt Injection attack against LLM-integrated Applications (Liu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2306.05499&sa=D&source=editors&ust=1779048536625319&usg=AOvVaw2HsL0Tn5p8lQOHnO1bKknX) |
| 132 | [▪ Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks (Syed, Almutairi, and Moaty, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23557&sa=D&source=editors&ust=1779048536625386&usg=AOvVaw1ujVobmukzHjWXwwy9bbyH) |
| 133 | [▪ ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected (Gharami et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20405&sa=D&source=editors&ust=1779048536625514&usg=AOvVaw1lKI4EMv79Z1_ExKwC3pFM) |
| 134 | [▪ AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens (Li, Wu, and Liu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17375&sa=D&source=editors&ust=1779048536625582&usg=AOvVaw3in_Je9Zb2sepRvhFJbZTa) |
| 135 | [▪ From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows (Ferrag et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.23260&sa=D&source=editors&ust=1779048536625651&usg=AOvVaw2jLqXm24wA9GntGc5LY-RG) |
| 136 | [▪ Detecting Prompt Injection Attacks Against Application Using Classifiers (Shaheer et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12583&sa=D&source=editors&ust=1779048536625720&usg=AOvVaw0qStFrff4t00I8YxZM45P4) |
| 137 | [▪ CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21228&sa=D&source=editors&ust=1779048536625787&usg=AOvVaw2V8dgVQgRWkhgQ-n0zKyqy) |
| 138 | [▪ When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection (Sahoo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10449&sa=D&source=editors&ust=1779048536625855&usg=AOvVaw3ai-qkHly7eOWbTVTYipOR) |
| 139 | [▪ Llama-based source code vulnerability detection: Prompt engineering vs Fine tuning (Ouchebara and Dupont, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09006&sa=D&source=editors&ust=1779048536625926&usg=AOvVaw0N6PHX2ZpoJYnJHsxBgSLu) |
| 140 | [▪ ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data (Wang, Jia, and Gong, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09321&sa=D&source=editors&ust=1779048536625995&usg=AOvVaw2k4GSsOiaXCIx3U_TvHxtr) |
| 141 | [▪ Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs (Zhong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08417&sa=D&source=editors&ust=1779048536626063&usg=AOvVaw13skwQ58It53CwBC__I-2l) |
| 142 | [▪ Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation (Song et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06304&sa=D&source=editors&ust=1779048536626131&usg=AOvVaw38JmtCyh_BuDHnJ6meJML4) |
| 143 | [▪ In-Context Representation Hijacking (Yona et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03771&sa=D&source=editors&ust=1779048536626192&usg=AOvVaw1rlYsJW3ePxLLiYe3l-Ve5) |
| 144 | [▪ HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03420&sa=D&source=editors&ust=1779048536626261&usg=AOvVaw0SwVKm0I3XI0z_XovrONau) |
| 145 | [▪ Rethinking Security in Semantic Communication: Latent Manipulation as a New Threat (Xi and Zhu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03361&sa=D&source=editors&ust=1779048536626325&usg=AOvVaw0ypp-cb_4EZWrPJn2ZJTAI) |
| 146 | [▪ Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs (Mathew et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.03768&sa=D&source=editors&ust=1779048536626392&usg=AOvVaw156xaBBxwi00KIoGy292Xn) |
| 147 | [▪ Securing Large Language Models (LLMs) from Prompt Injection Attacks (Suri and McCrae, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01326&sa=D&source=editors&ust=1779048536626465&usg=AOvVaw2k4ULynZHzWZ7wFuGUNamV) |
| 148 | [▪ Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis (Kang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00966&sa=D&source=editors&ust=1779048536626530&usg=AOvVaw2zLDeuGIgnwoCfeVJ_NMPy) |
| 149 | [▪ BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20597&sa=D&source=editors&ust=1779048536626596&usg=AOvVaw3YoA2XjgiRRwv3fWmbt6Iq) |
| 150 | [▪ On the Feasibility of Hijacking MLLMs' Decision Chain via One Perturbation (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20002&sa=D&source=editors&ust=1779048536626663&usg=AOvVaw0k1TkKnJEmdrCVIY1wFYOB) |
| 151 | [▪ Effective Command-line Interface Fuzzing with Path-Aware Large Language Model Orchestration (Shiraishi, Cao, and Shinagawa, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20555&sa=D&source=editors&ust=1779048536626730&usg=AOvVaw2Qdho-vdiEwHFaQE8v8OLr) |
| 152 | [▪ Prompt Fencing: A Cryptographic Approach to Establishing Security Boundaries in Large Language Model Prompts (Peh, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19727&sa=D&source=editors&ust=1779048536626799&usg=AOvVaw2SJpT2kGm15rWkhyTdpKwT) |
| 153 | [▪ RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation (Tafreshian, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18790&sa=D&source=editors&ust=1779048536626865&usg=AOvVaw1V0OOe0NhvMvL87DaCjGPQ) |
| 154 | [▪ PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization (Jawad and Brunel, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16209&sa=D&source=editors&ust=1779048536626944&usg=AOvVaw1V6W2CIsmxGk44Kg3UgLrd) |
| 155 | [▪ Securing AI Agents Against Prompt Injection Attacks (Ramakrishnan and Balaji, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15759&sa=D&source=editors&ust=1779048536627011&usg=AOvVaw1d16lPQkFG0nDWVFnhErxf) |
| 156 | [▪ On-Premise SLMs vs. Commercial LLMs: Prompt Engineering and Incident Classification in SOCs and CSIRTs (Almeida et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14908&sa=D&source=editors&ust=1779048536627080&usg=AOvVaw1OmnomzjX_q9cbGVcMTdxb) |
| 157 | [▪ Cybersecurity AI: Hacking the AI Hackers via Prompt Injection (Mayoral-Vilches and Rynning, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21669&sa=D&source=editors&ust=1779048536627146&usg=AOvVaw11Sfm_UsG2nYFKpS3F9nc9) |
| 158 | [▪ Whose Narrative is it Anyway? A KV Cache Manipulation Attack (Ganesh, Iyer, and Ananthan, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12752&sa=D&source=editors&ust=1779048536627210&usg=AOvVaw3q-nVRlm0mJ-vUs6yoynCg) |
| 159 | [▪ GRAPHTEXTACK: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNs (Ma, Trivedi, and Koutra, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12423&sa=D&source=editors&ust=1779048536627285&usg=AOvVaw1oS7iZWPBU0ehJnMiE4Ddu) |
| 160 | [▪ Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification (Jayathilaka, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12295&sa=D&source=editors&ust=1779048536627352&usg=AOvVaw3yXHziGTSQ8lNh63OyQUHB) |
| 161 | [▪ Prompt Engineering vs. Fine-Tuning for LLM-Based Vulnerability Detection in Solana and Algorand Smart Contracts (Boi and Esposito, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11250&sa=D&source=editors&ust=1779048536627427&usg=AOvVaw3_eGPvTdrwxhgDkkRa7ZKa) |
| 162 | [▪ PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization (Geng et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10720&sa=D&source=editors&ust=1779048536627494&usg=AOvVaw2PN5Im1OTKkHiGKsbTHkys) |
| 163 | [▪ BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10714&sa=D&source=editors&ust=1779048536627561&usg=AOvVaw0ZKbvT1DLnWXjx0XjGwyNr) |
| 164 | [▪ DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.11358&sa=D&source=editors&ust=1779048536627627&usg=AOvVaw26qew8jTjPsn-7HBE9oNEf) |
| 165 | [▪ MPMA: Preference Manipulation Attack Against Model Context Protocol (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11154&sa=D&source=editors&ust=1779048536627693&usg=AOvVaw0zg3yvt08mHXWo3iaSXuOu) |
| 166 | [▪ Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy (Gudi\\~no-Rosero et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.04281&sa=D&source=editors&ust=1779048536627758&usg=AOvVaw0M3-UeoWh-b7lmQLS0aL_7) |
| 167 | [▪ Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs (Fastowski et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05919&sa=D&source=editors&ust=1779048536627823&usg=AOvVaw262r2_5bv-t4sWSdzOEAkx) |
| 168 | [▪ Black-Box Guardrail Reverse-engineering Attack (Yao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04215&sa=D&source=editors&ust=1779048536627886&usg=AOvVaw3F5Kr-BaewzpvUSx8BWmgl) |
| 169 | [▪ Hybrid Fuzzing with LLM-Guided Input Mutation and Semantic Feedback (Lin, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03995&sa=D&source=editors&ust=1779048536627968&usg=AOvVaw0wTAz804lZ3HXxkgNOKIFt) |
| 170 | [▪ Death by a Thousand Prompts: Open Model Vulnerability Analysis (Chang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03247&sa=D&source=editors&ust=1779048536628052&usg=AOvVaw03MH6YQA7dAS12MXw7PiSM) |
| 171 | [▪ "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers (Zhou et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01287&sa=D&source=editors&ust=1779048536628125&usg=AOvVaw1m0CZ9hoR1fX_41P7_r77p) |
| 172 | [▪ Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models (Ganiuly and Smaiyl, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01634&sa=D&source=editors&ust=1779048536628193&usg=AOvVaw17SaVYU92A6iMlUXNEM4c8) |
| 173 | [▪ Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization (Verma, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05831&sa=D&source=editors&ust=1779048536628260&usg=AOvVaw16lv8BGiAE67KeVc6JNaGX) |
| 174 | [▪ Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels (Du et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27140&sa=D&source=editors&ust=1779048536628327&usg=AOvVaw3v9vunji_iToxaXHccMT4Q) |
| 175 | [▪ Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token (Zychlinski and Kainan, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26847&sa=D&source=editors&ust=1779048536628395&usg=AOvVaw3uk87iAQMOMqwfvpmnK106) |
| 176 | [▪ QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents (Xie et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23675&sa=D&source=editors&ust=1779048536628467&usg=AOvVaw1JrY2XWVyXWQ51DXpSS7xN) |
| 177 | [▪ CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22963&sa=D&source=editors&ust=1779048536628533&usg=AOvVaw2hqfcL5StgntMR7p04Ogfi) |
| 178 | [▪ Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22944&sa=D&source=editors&ust=1779048536628599&usg=AOvVaw20FOSm5C4vHmIX2824gI47) |
| 179 | [▪ DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12104&sa=D&source=editors&ust=1779048536628665&usg=AOvVaw2Vsq4afEPiB3ACmo_cr6bE) |
| 180 | [▪ PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents (Murray, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.17884&sa=D&source=editors&ust=1779048536628732&usg=AOvVaw3REkMi0bRhcFNGMCAWbDYL) |
| 181 | [▪ GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments? (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20333&sa=D&source=editors&ust=1779048536628829&usg=AOvVaw1mQGBqwkKlDsI7xURGcqio) |
| 182 | [▪ CourtGuard: A Local, Multiagent Prompt Injection Classifier (Wu and Maslowski, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19844&sa=D&source=editors&ust=1779048536628897&usg=AOvVaw31Isz0RuoEjqjYUmtdreWk) |
| 183 | [▪ When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16823&sa=D&source=editors&ust=1779048536628965&usg=AOvVaw2RY9r2UijQObLQ6TNfpp2D) |
| 184 | [▪ RoBCtrl: Attacking GNN-Based Social Bot Detectors via Reinforced Manipulation of Bots Control Interaction (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16035&sa=D&source=editors&ust=1779048536629031&usg=AOvVaw0Y19-Ci26d9Xy1AdIIzm7U) |
| 185 | [▪ Black-box Optimization of LLM Outputs by Asking for Directions (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16794&sa=D&source=editors&ust=1779048536629096&usg=AOvVaw1yWFP8jUKLnKL-6GRwVAA9) |
| 186 | [▪ Prompt injections as a tool for preserving identity in GAI image descriptions (Glazko and Mankoff, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16128&sa=D&source=editors&ust=1779048536629162&usg=AOvVaw2hAfttxxtOwqIdz-5enJwo) |
| 187 | [▪ Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15738&sa=D&source=editors&ust=1779048536629229&usg=AOvVaw3VQZd_7hzjzlWD4skpF_7u) |
| 188 | [▪ Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14381&sa=D&source=editors&ust=1779048536629294&usg=AOvVaw0_Gx_p-ZE12RslZ_dyPeFG) |
| 189 | [▪ RHINO: Guided Reasoning for Mapping Network Logs to Adversarial Tactics and Techniques with Large Language Models (Meng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14233&sa=D&source=editors&ust=1779048536629363&usg=AOvVaw1ktaRwT2xBLAfHg3Uee9C8) |
| 190 | [▪ PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features (Zou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14005&sa=D&source=editors&ust=1779048536629433&usg=AOvVaw2rZmw0dJ53BZ9Tjant1BMS) |
| 191 | [▪ Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection (Guo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.14622&sa=D&source=editors&ust=1779048536629501&usg=AOvVaw3LiVLjAJHOn6z2m6F88pQ5) |
| 192 | [▪ In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers (Cohen, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13543&sa=D&source=editors&ust=1779048536629568&usg=AOvVaw3NcNMeyjEcjCfZ023ajFA0) |
| 193 | [▪ PromptLocate: Localizing Prompt Injection Attacks (Jia et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12252&sa=D&source=editors&ust=1779048536629632&usg=AOvVaw1PsQD-cLZHm8O937h7sryK) |
| 194 | [▪ LineBreaker: Finding Token-Inconsistency Bugs with Large Language Models (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.01668&sa=D&source=editors&ust=1779048536629698&usg=AOvVaw2mgcomaSx6PT2Ur_uPcVGp) |
| 195 | [▪ CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense (Zhuochen, Wai, and Vrizlynn, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11137&sa=D&source=editors&ust=1779048536629767&usg=AOvVaw0JsqzYFVCliAuQUfgxNPtB) |
| 196 | [▪ VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands (Liu and Tang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09699&sa=D&source=editors&ust=1779048536629834&usg=AOvVaw05PwVs4K3JjjPXVu9tJXQM) |
| 197 | [▪ CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization (Das et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08829&sa=D&source=editors&ust=1779048536629899&usg=AOvVaw2ixfqxsyspE3EeOV0nSN5z) |
| 198 | [▪ AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00088&sa=D&source=editors&ust=1779048536629970&usg=AOvVaw0bJ6W1FmOoRVlLkKzfzy9G) |
| 199 | [▪ Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks? (Bhagwatkar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05244&sa=D&source=editors&ust=1779048536630039&usg=AOvVaw0Rl6v6BXd1DlpOlavrIKjD) |
| 200 | [▪ System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection (Li, Guo, and Cai, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.06493&sa=D&source=editors&ust=1779048536630106&usg=AOvVaw19ENH6UGgYeWjgEYhE82go) |
| 201 | [▪ RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection (Wen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04885&sa=D&source=editors&ust=1779048536630189&usg=AOvVaw0nwqTXqNeQDbpF_286Ucl0) |
| 202 | [▪ Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers (KumarRavindran, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04528&sa=D&source=editors&ust=1779048536630262&usg=AOvVaw2khZIKZDaKTA5pmi7CxKBj) |
| 203 | [▪ VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy (Cui et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04261&sa=D&source=editors&ust=1779048536630368&usg=AOvVaw1DhkDwew9hBBDBZDph2oHE) |
| 204 | [▪ AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04257&sa=D&source=editors&ust=1779048536630477&usg=AOvVaw01ec7Ck9y7wibJxucuvLvw) |
| 205 | [▪ Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.02372&sa=D&source=editors&ust=1779048536630554&usg=AOvVaw0HKitxeFYJADxkb5DIvake) |
| 206 | [▪ Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations (Teja et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02319&sa=D&source=editors&ust=1779048536630637&usg=AOvVaw3zfM78GimsxIKGXj_03ig6) |
| 207 | [▪ Bypassing Prompt Guards in Production with Controlled-Release Prompting (Fairoze et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01529&sa=D&source=editors&ust=1779048536630740&usg=AOvVaw0giKdEtgyd25nW0TjovkVz) |
| 208 | [▪ In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b (Durner, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01259&sa=D&source=editors&ust=1779048536630840&usg=AOvVaw1cTrYYXUT5TPr1zzdV710M) |
| 209 | [▪ WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01354&sa=D&source=editors&ust=1779048536630940&usg=AOvVaw01OjcNjFGUG1jQrVzwM164) |
| 210 | [▪ Fingerprinting LLMs via Prompt Injection (Hu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25448&sa=D&source=editors&ust=1779048536631038&usg=AOvVaw3yKdzztgWdGHagjgGPL60D) |
| 211 | [▪ SecInfer: Preventing Prompt Injection via Inference-time Scaling (Liu et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24967&sa=D&source=editors&ust=1779048536631137&usg=AOvVaw1OQW7RCnRgsvOzYWwQg3zx) |
| 212 | [▪ Prompt to Pwn: Automated Exploit Generation for Smart Contracts (Xiao et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01371&sa=D&source=editors&ust=1779048536631237&usg=AOvVaw1zDFc2EPZsx6xEfH7QwCQj) |
| 213 | [▪ AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection (Wang et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01249&sa=D&source=editors&ust=1779048536631311&usg=AOvVaw375cUSZd3sAwUSBiC4UiWJ) |
| 214 | [▪ LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks (Panebianco et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00602&sa=D&source=editors&ust=1779048536631380&usg=AOvVaw1KjLfpSfVdvDthsoU2n5nM) |
| 215 | [▪ Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding (Pathade, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22304&sa=D&source=editors&ust=1779048536631454&usg=AOvVaw1ELeV0zxV3jHpb_ziEjDPB) |
| 216 | [▪ Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities (Samancioglu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21133&sa=D&source=editors&ust=1779048536631545&usg=AOvVaw2a8G3Z8jyYQBFYkPYuFfdS) |
| 217 | [▪ Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models (Downer et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20704&sa=D&source=editors&ust=1779048536631618&usg=AOvVaw1De5JW02tTDTLE4uzW9wH_) |
| 218 | [▪ MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? (Wahed et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19598&sa=D&source=editors&ust=1779048536631685&usg=AOvVaw2rKQQh1L-U48rLECp9xzcI) |
| 219 | [▪ Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree (Johnson, Pham, and Le, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14799&sa=D&source=editors&ust=1779048536631752&usg=AOvVaw1HTGuv-pwKd-cGHcattBNi) |
| 220 | [▪ Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design (Charles, Curry, and Charles, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14207&sa=D&source=editors&ust=1779048536631819&usg=AOvVaw0c5SNp07t-hvI_FWJRija7) |
| 221 | [▪ Can Indirect Prompt Injection Attacks Be Detected and Removed? (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16580&sa=D&source=editors&ust=1779048536631883&usg=AOvVaw1KQ1873FgMYLMRa1Vf7rUw) |
| 222 | [▪ TopicAttack: An Indirect Prompt Injection Attack via Topic Transition (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13686&sa=D&source=editors&ust=1779048536631957&usg=AOvVaw2OPqC4dPyfllORdSxqFxCx) |
| 223 | [▪ Prompt Injection 2.0: Hybrid AI Threats (McHugh, \\v{S}ekrst, and Cefalu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13169&sa=D&source=editors&ust=1779048536632020&usg=AOvVaw3LJ1lP5cahpP2RZARw-qxg) |
| 224 | [▪ MAD-Spear: A Conformity-Driven Prompt Injection Attack on Multi-Agent Debate Systems (Cui and Du, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13038&sa=D&source=editors&ust=1779048536632087&usg=AOvVaw1-elyGYNFtep0ga6NdFOOJ) |
| 225 | [▪ Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems (Hackett et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.11168&sa=D&source=editors&ust=1779048536632154&usg=AOvVaw1IBdLbXpX-kqoH-OdDvpY1) |
| 226 | [▪ Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems (Atta et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10457&sa=D&source=editors&ust=1779048536632219&usg=AOvVaw35OicZAYahjujj8gx-Lnqk) |
| 227 | [▪ May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks (Pandya et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07417&sa=D&source=editors&ust=1779048536632287&usg=AOvVaw2G8z0tH7R-oZcOTxUwbS4y) |
| 228 | [▪ How Not to Detect Prompt Injections with an LLM (Choudhary et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05630&sa=D&source=editors&ust=1779048536632368&usg=AOvVaw1u9iQJ8g_NkCSFi1_zKot_) |
| 229 | [▪ FrameShift: Learning to Resize Fuzzer Inputs Without Breaking Them (Green, Goues, and Brown, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05421&sa=D&source=editors&ust=1779048536632444&usg=AOvVaw22vV6kiCaZ5vPZKoTCxWg9) |
| 230 | [▪ Attention Tracker: Detecting Prompt Injection Attacks in LLMs (Hung et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.00348&sa=D&source=editors&ust=1779048536632516&usg=AOvVaw1aSQi8geqPOR5TX0d_CGI3) |
| 231 | [▪ SINCon: Mitigate LLM-Generated Malicious Message Injection Attack for Rumor Detection (Zhang et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.07135&sa=D&source=editors&ust=1779048536632592&usg=AOvVaw18pXnuZs5FuBWJ-NdYLhTi) |
| 232 | [▪ RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage (Zhong et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08966&sa=D&source=editors&ust=1779048536632681&usg=AOvVaw0xYDHu2oKF0fsCoLw9yxXU) |
| 233 | [▪ EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Liao et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11295&sa=D&source=editors&ust=1779048536632803&usg=AOvVaw1CaWP8bWElZ_f06k-z-5kk) |
| 234 | [▪ Self-interpreting Adversarial Images (Zhang et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.08970&sa=D&source=editors&ust=1779048536632931&usg=AOvVaw1wXbCaOY80z0-9XLH6b2Un) |
| 235 | [▪ SecAlign: Defending Against Prompt Injection with Preference Optimization (Chen et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.05451&sa=D&source=editors&ust=1779048536633052&usg=AOvVaw00my3zScchkXZzMWlWnwTG) |
| 236 | [▪ Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM (Guo et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15614&sa=D&source=editors&ust=1779048536633186&usg=AOvVaw326iMgkX5HwOo4q2ObQ4GH) |
| 237 | [▪ Defending LVLMs Against Vision Attacks through Partial-Perception Supervision (Zhou et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.12722&sa=D&source=editors&ust=1779048536633262&usg=AOvVaw2B1JhpKR4RZPOjQ3dEGJO9) |
| 238 | [▪ PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks (Feng et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.02629&sa=D&source=editors&ust=1779048536633329&usg=AOvVaw3nlBhxxPWHzOZ5820I0mQl) |
| 239 | [▪ Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization (Cooper, Narnoli, and Surdeanu, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.12212&sa=D&source=editors&ust=1779048536633400&usg=AOvVaw13uwe1az7XdECI8xrskVNl) |
| 240 | [▪ PGD-Imp: Rethinking and Unleashing Potential of Classic PGD with Dual Strategies for Imperceptible Adversarial Attacks (Li et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.11168&sa=D&source=editors&ust=1779048536633477&usg=AOvVaw0OG0K77lwVrtjADHy-g-eL) |
| 241 | [▪ Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks? (Chen et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.03411&sa=D&source=editors&ust=1779048536633541&usg=AOvVaw1CPBH-ts4haIQP1NZGPUrz) |
| 242 | [▪ Failures to Find Transferable Image Jailbreaks Between Vision-Language Models (Shaeffer et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.15211&sa=D&source=editors&ust=1779048536633618&usg=AOvVaw2lxQqV7rkbfEV_kGgeES70) |
| 243 | [▪ HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models (Gao et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.13896&sa=D&source=editors&ust=1779048536633685&usg=AOvVaw2TaEbRh0N6PSiGfyYg8Kgx) |
| 244 | [▪ Comprehensive Assessment of Jailbreak Attacks Against LLMs (Chu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.05668&sa=D&source=editors&ust=1779048536633747&usg=AOvVaw2QMJUsQD-7jV_6d8cWidH3) |
| 245 | [▪ From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection (Wang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.10198&sa=D&source=editors&ust=1779048536633813&usg=AOvVaw3-l_a_1oMFNiKLklBGXZ1l) |
| 246 | [▪ FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks (Wang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.21492&sa=D&source=editors&ust=1779048536633878&usg=AOvVaw23S3SYwdHkNdMr1Ghsa6bc) |
| 247 | [▪ Formalizing and Benchmarking Prompt Injection Attacks and Defenses (Liu et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2310.12815&sa=D&source=editors&ust=1779048536633948&usg=AOvVaw28JMOC_AO3FYimKqJE9SVv) |
| 248 | [▪ Universal and Context-Independent Triggers for Precise Control of LLM Outputs (Liang, Li, and Yu, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14738&sa=D&source=editors&ust=1779048536634027&usg=AOvVaw1mslSTnSRfuD8LSTZFM9JH) |
| 249 | [▪ SecONN: An Optical Neural Network Framework with Concurrent Detection of Thermal Fault Injection Attacks (Nishida et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14741&sa=D&source=editors&ust=1779048536634100&usg=AOvVaw2f81ZsSiKkjFGRJzRwSR9K) |
| 250 | [▪ Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks (Pasquini et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.20911&sa=D&source=editors&ust=1779048536634167&usg=AOvVaw1LgTuZoQhaM6wxWGtzd_pS) |
| 251 | [▪ Optimization-based Prompt Injection Attack to LLM-as-a-Judge (Shi et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.17710&sa=D&source=editors&ust=1779048536634231&usg=AOvVaw0EX4vkzwuu35oKYUSvy-t5) |
| 252 | [▪ Goal-guided Generative Prompt Injection Attack on Large Language Models (Zhang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.07234&sa=D&source=editors&ust=1779048536634297&usg=AOvVaw3n_Ohsr-nzB871rzQFFks4) |
| 253 | [▪ Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures (Benjamin et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.23308&sa=D&source=editors&ust=1779048536634362&usg=AOvVaw3RX6fdnqTl6NSmHBwCm75e) |
| 254 | [▪ InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models (Li, Liu, and Xiao, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22770&sa=D&source=editors&ust=1779048536634439&usg=AOvVaw37Fa4o0l0q6sQC9qq3rx-r) |
| 255 | [▪ FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks (Wang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.21492&sa=D&source=editors&ust=1779048536634508&usg=AOvVaw2NIuZtCvuF9hTAJaAPRm8d) |
| 256 | [▪ Embedding-based classifiers can detect prompt injection attacks (Ayub and Majumdar, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22284&sa=D&source=editors&ust=1779048536634571&usg=AOvVaw29NX1xqYSUm78Ai4_rA14C) |
| 257 | [▪ Imprompter: Tricking LLM Agents into Improper Tool Use (Fu et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.14923&sa=D&source=editors&ust=1779048536634636&usg=AOvVaw0Xtq47zy2hS7932olr45si) |
| 258 | [▪ System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective (Wu, Cecchetti, and Xiao, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.19091&sa=D&source=editors&ust=1779048536634706&usg=AOvVaw0v2rHSBOK8MjcPSM92I_Tg) |
| 259 | [▪ EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Liao et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11295&sa=D&source=editors&ust=1779048536634773&usg=AOvVaw0l6vC5fMCy5M1gUR8e2pKf) |
| 260 | [▪ Efficient Detection of Toxic Prompts in Large Language Models (Liu et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.11727&sa=D&source=editors&ust=1779048536634871&usg=AOvVaw1pDcs1bki_IOlakTSvY_4K) |
| 261 | [▪ Goal-guided Generative Prompt Injection Attack on Large Language Models (Zhang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.07234&sa=D&source=editors&ust=1779048536634960&usg=AOvVaw16IYxh6ZcKCoTHIv_RrgRm) |
| 262 | [▪ Soft Prompts Go Hard: Steering Visual Language Models with Hidden Meta-Instructions (Zhang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.08970&sa=D&source=editors&ust=1779048536635092&usg=AOvVaw2LkRs6ZAVMw1FU2KWA94po) |
| 263 | [▪ Optimization-based Prompt Injection Attack to LLM-as-a-Judge (Shi et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.17710&sa=D&source=editors&ust=1779048536635206&usg=AOvVaw2XZA7imULqWXrjibFKsu-r) |
| 264 | [▪ Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks (De Stefano, Schönherr, Pellegrino, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.05025%23&sa=D&source=editors&ust=1779048536635477&usg=AOvVaw1NnSYVxBSoFdBVx4QsyePZ) |
| 265 | [▪ InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents (Zhang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.02691&sa=D&source=editors&ust=1779048536635610&usg=AOvVaw0DYQ_mAwppStQe-xQnQh1C) |
| 266 | [▪ LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks (Happe, Kaplan, and Cito, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2310.11409&sa=D&source=editors&ust=1779048536635847&usg=AOvVaw3GpphrZNolXcFFsRwvN_cZ) |
| 267 | [▪ Securing the Diagnosis of Medical Imaging: An In-depth Analysis of AI-Resistant Attacks (Biswas et al, Aug 2024)](https://docs.google.com/spreadsheets/d/e/2PACX-1vQ77JhxxsGIu1MKbNvQCrCmtCVNcpiD_ROrBCx19oFjQ8pgjKcgO2YeP2Kw_AAHE5bpWW7CoPXwr3Vh/pubhtml/efaidnbmnnnibpcajpcglclefindmkaj/https://arxiv.org/pdf/2408.00348v1) |
| 268 | [▪ On Feasibility of Intent Obfuscating Attacks (Li and Shafto, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02674&sa=D&source=editors&ust=1779048536636092&usg=AOvVaw0dClnSRH7wXM-0Sz49e2GI) |

|     |
| --- |
| Prompt Injection and Input Manipulation (Direct and Indirect) |

**>**

**<**

‍

#### System and Meta Prompt Extraction

Covers:

- MITRE ATLAS Discovery and Exfiltration

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : System and Meta Prompt Extraction

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks (Zhuang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11459&sa=D&source=editors&ust=1779048535894632&usg=AOvVaw0g2zWK0wJ5pqbDRYQ0zfbR) |
| 3 | [▪ A First Look at the Security Issues in the Model Context Protocol Ecosystem (Li and Gao, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16558&sa=D&source=editors&ust=1779048535894809&usg=AOvVaw1ctzYAhGtJBMCETwSTE-sk) |
| 4 | [▪ Peering Behind the Shield: Guardrail Identification in Large Language Models (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.01241&sa=D&source=editors&ust=1779048535894925&usg=AOvVaw26gYH_iNsb-LBpwoihILGL) |
| 5 | [▪ The System Prompt Is the Attack Surface: How LLM Agent Configuration Shapes Security and Creates Exploitable Vulnerabilities (Litvak, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25056&sa=D&source=editors&ust=1779048535895071&usg=AOvVaw0NHSRpoxYXkt5ZwzKIP9I_) |
| 6 | [▪ Arbiter: Detecting Interference in LLM Agent System Prompts (Mason, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.08993&sa=D&source=editors&ust=1779048535895177&usg=AOvVaw3TL8p8ID7F_iqhatIRBroG) |
| 7 | [▪ OptiLeak: Efficient Prompt Reconstruction via Reinforcement Learning in Multi-tenant LLM Services (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20595&sa=D&source=editors&ust=1779048535895290&usg=AOvVaw1-Ehe7vfJJa1cTaO5rTmtO) |
| 8 | [▪ Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models (Zhao et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06837&sa=D&source=editors&ust=1779048535895382&usg=AOvVaw2ncbsn_nRhYFjwJym5q_Og) |
| 9 | [▪ Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security (Chua, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19399&sa=D&source=editors&ust=1779048535895469&usg=AOvVaw2fkk0CTORB4WBhIzwjCILJ) |
| 10 | [▪ When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16773&sa=D&source=editors&ust=1779048535895570&usg=AOvVaw3_8pP4a70v8dbXQlVNojZF) |
| 11 | [▪ Slot: Provenance-Driven APT Detection through Graph Reinforcement Learning (Qiao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.17910&sa=D&source=editors&ust=1779048535895676&usg=AOvVaw1XLDReMe06WmvT7d02qkTA) |
| 12 | [▪ ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space (Chen, Dolan-Gavitt, and Lin, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10323&sa=D&source=editors&ust=1779048535895775&usg=AOvVaw3cpjJT2TAYuou9HVnGc9e4) |
| 13 | [▪ LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07056&sa=D&source=editors&ust=1779048535895870&usg=AOvVaw3EaZ24TUdw3ygm2BjiN0TR) |
| 14 | [▪ LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users (Hilel et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02850&sa=D&source=editors&ust=1779048535895975&usg=AOvVaw1a2oLloVD9QOgQWd5OaAJR) |
| 15 | [▪ Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models (Liang et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02416&sa=D&source=editors&ust=1779048535896131&usg=AOvVaw1GB9EFI-zH7HEOGDH63i-J) |
| 16 | [▪ Prompt Obfuscation for Large Language Models (Pape, Eisenhofer, and Schönherr, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11026&sa=D&source=editors&ust=1779048535896221&usg=AOvVaw2ckBQqxLPhTbI8K3TRjEeV) |
| 17 | [▪ Prompt Leakage effect and defense strategies for multi-turn LLM interactions (Agarwal et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.16251&sa=D&source=editors&ust=1779048535896313&usg=AOvVaw2yJC8M0gol_UttYGrLg8jA) |

|     |
| --- |
| System and Meta Prompt Extraction |

**>**

**<**

‍

#### Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts

Covers:

- MITRE ATLAS Resource Development

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11086&sa=D&source=editors&ust=1779048536947460&usg=AOvVaw2NrtD1ELMvB-svH4zzezN3) |
| 3 | [▪ Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents (Probst, Happe, and Cito, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27143&sa=D&source=editors&ust=1779048536947688&usg=AOvVaw3Ol8FotmhmqDglmejJXcF0) |
| 4 | [▪ xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models (Luong et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.13021&sa=D&source=editors&ust=1779048536947818&usg=AOvVaw3rbOgC-YioCfDPKzny2Sc8) |
| 5 | [▪ Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks (Merves et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17159&sa=D&source=editors&ust=1779048536947943&usg=AOvVaw10hAEhHL4xDIj5wQh3h0cw) |
| 6 | [▪ Stand-Alone Complex or Vibercrime? Exploring the adoption and innovation of GenAI tools, coding assistants, and agents within cybercrime ecosystems (Hughes, Collier, and Thomas, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29545&sa=D&source=editors&ust=1779048536948072&usg=AOvVaw17vlLyBylujk9ZS3_6MUNa) |
| 7 | [▪ MalTool: Malicious Tool Attacks on LLM Agents (Hu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12194&sa=D&source=editors&ust=1779048536948197&usg=AOvVaw2BmBZBAObZ1wEuIiFE5-Yk) |
| 8 | [▪ QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery (Tsigkourakos and Patsakis, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09774&sa=D&source=editors&ust=1779048536948302&usg=AOvVaw1bBzEeqyL5nRXJRPA2SXm7) |
| 9 | [▪ AICrypto: Evaluating Cryptography Capabilities of Large Language Models (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09580&sa=D&source=editors&ust=1779048536948434&usg=AOvVaw2DneiJM8kyyqdQn0MAvpGq) |
| 10 | [▪ Zer0n: An AI-Assisted Vulnerability Discovery and Blockchain-Backed Integrity Framework (Parmar et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07019&sa=D&source=editors&ust=1779048536948568&usg=AOvVaw2TcEPY92MG-ZRBwZ8GcEPQ) |
| 11 | [▪ AutoPatch: Multi-Agent Framework for Patching Real-World CVE Vulnerabilities (Seo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.04195&sa=D&source=editors&ust=1779048536948693&usg=AOvVaw1oLsZd8YXg70GSsZTSkvT1) |
| 12 | [▪ SeedAIchemy: LLM-Driven Seed Corpus Generation for Fuzzing (Wen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12448&sa=D&source=editors&ust=1779048536948846&usg=AOvVaw3AiL7DqIUKxLdqbhBGXjfU) |
| 13 | [▪ RulePilot: An LLM-Powered Agent for Security Rule Generation (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12224&sa=D&source=editors&ust=1779048536948978&usg=AOvVaw3re626vDS9Vs1AxFNvxV-q) |
| 14 | [▪ AICrypto: A Comprehensive Benchmark for Evaluating Cryptography Capabilities of Large Language Models (Wang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09580&sa=D&source=editors&ust=1779048536949076&usg=AOvVaw0YcjBP_9oTgmQt5TsMMkm2) |
| 15 | [▪ Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search (Moss, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08899&sa=D&source=editors&ust=1779048536949205&usg=AOvVaw00wUJLKP3qWdZv1ddHgFs_) |

|     |
| --- |
| Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts |

**>**

**<**

‍

#### Jailbreak, Cost Harvesting, or Erode ML Model Integrity

Covers:

- MITRE ATLAS Privilege Escalation, Defense Evasion, and Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Jailbreak, Cost Harvesting, or Erode ML Model Integrity

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ EVA: Editing for Versatile Alignment against Jailbreaks (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14750&sa=D&source=editors&ust=1779048536773554&usg=AOvVaw2p9j_4sCYEPonjqMvrfsUZ) |
| 3 | [▪ The Great Pretender: A Stochasticity Problem in LLM Jailbreak (Monteuuis, Chen, and Petit, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14418&sa=D&source=editors&ust=1779048536773682&usg=AOvVaw2yDu7wu-Z2UwnRQhcb2ZIv) |
| 4 | [▪ Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization (Dotsinski and Eustratiadis, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13359&sa=D&source=editors&ust=1779048536773734&usg=AOvVaw3iE0fnvNmsrF079_uGGaUX) |
| 5 | [▪ Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12673&sa=D&source=editors&ust=1779048536773778&usg=AOvVaw1318uabLpeYUVNpTuxy5Rm) |
| 6 | [▪ MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11002&sa=D&source=editors&ust=1779048536773820&usg=AOvVaw0SpVKk8SULefeiaUNHN_x7) |
| 7 | [▪ Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs (Yoon et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10998&sa=D&source=editors&ust=1779048536773858&usg=AOvVaw1X5Bhy2qh91w-bK48rO9p_) |
| 8 | [▪ LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10779&sa=D&source=editors&ust=1779048536773902&usg=AOvVaw0nUuOc8X3L_4vCknFcSqyo) |
| 9 | [▪ Re-Triggering Safeguards within LLMs for Jailbreak Detection (Lin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10611&sa=D&source=editors&ust=1779048536773939&usg=AOvVaw1Llw9FqQHwIec6uhGx6rHu) |
| 10 | [▪ Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing (Lin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10582&sa=D&source=editors&ust=1779048536773976&usg=AOvVaw1Q_HhtlZ1loxyBITKGIAqa) |
| 11 | [▪ The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring (Hossain et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09225&sa=D&source=editors&ust=1779048536774023&usg=AOvVaw218rj9mYqbP9zNn0TZuRNW) |
| 12 | [▪ Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success (Maple, Kumar, and Tapwal, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09070&sa=D&source=editors&ust=1779048536774063&usg=AOvVaw1K9SV95F5zF-OTNU5mnGB-) |
| 13 | [▪ Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off (Chen, Liu, and Cao, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08878&sa=D&source=editors&ust=1779048536774101&usg=AOvVaw2uYvtdcTP0kU-0InRy0S-0) |
| 14 | [▪ Mitigating Many-shot Jailbreak Attacks with One Single Demonstration (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08277&sa=D&source=editors&ust=1779048536774136&usg=AOvVaw0JhKd6SpwNVBQiX3UylWU_) |
| 15 | [▪ OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07414&sa=D&source=editors&ust=1779048536774171&usg=AOvVaw1LGci1_DEVN2NxlAk_DY7l) |
| 16 | [▪ SoK: Robustness in Large Language Models against Jailbreak Attacks (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05058&sa=D&source=editors&ust=1779048536774203&usg=AOvVaw007WuaRzlFxh2CCVRJects) |
| 17 | [▪ Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization (Fang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04700&sa=D&source=editors&ust=1779048536774237&usg=AOvVaw0uoUk8fa90mullCS8LSYQ0) |
| 18 | [▪ Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses (Derya and Sunar, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03095&sa=D&source=editors&ust=1779048536774272&usg=AOvVaw07JBMyKMMC_b5TxhJfpxNq) |
| 19 | [▪ Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection (Hu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02958&sa=D&source=editors&ust=1779048536774308&usg=AOvVaw3Jc0XFET1nvn6FWQnC8_4O) |
| 20 | [▪ ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming (B\\'ejar et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02647&sa=D&source=editors&ust=1779048536774449&usg=AOvVaw1vav0bkrhl7E7KZMX8CA5e) |
| 21 | [▪ SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00974&sa=D&source=editors&ust=1779048536774519&usg=AOvVaw2ABvqnkZlXunXfXBN05u_u) |
| 22 | [▪ Jailbroken Frontier Models Retain Their Capabilities (Zhu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00267&sa=D&source=editors&ust=1779048536774558&usg=AOvVaw0RgRy_xUubHLLQnzNru8hi) |
| 23 | [▪ Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry (Lan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27019&sa=D&source=editors&ust=1779048536774595&usg=AOvVaw0nwptJbuLtyWrrP1R3ethb) |
| 24 | [▪ TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning (Sun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27861&sa=D&source=editors&ust=1779048536774633&usg=AOvVaw1_y6AAcRnojbpViWueZKmn) |
| 25 | [▪ One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety (Arif et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25921&sa=D&source=editors&ust=1779048536774670&usg=AOvVaw0gwl8vWSIozEWJTAsU24Cw) |
| 26 | [▪ Jailbreaking Frontier Foundation Models Through Intention Deception (Wang, Sycara, and Xie, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24082&sa=D&source=editors&ust=1779048536774705&usg=AOvVaw3_Yts5jutWwKCeKk-uzsqe) |
| 27 | [▪ Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards (Hammadia et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23341&sa=D&source=editors&ust=1779048536774743&usg=AOvVaw0JfN-eKRxNBuEiY_6fzdiu) |
| 28 | [▪ Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem (Lin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17299&sa=D&source=editors&ust=1779048536774778&usg=AOvVaw0TFNes_Agx69fNyTCwdl8C) |
| 29 | [▪ Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4 (Polyakov and Kuznetsov, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19461&sa=D&source=editors&ust=1779048536774815&usg=AOvVaw1LwqTx2lq1z3r3khHsp366) |
| 30 | [▪ ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02356&sa=D&source=editors&ust=1779048536774853&usg=AOvVaw1eKqv-jrJEBXmFDQY7AKSH) |
| 31 | [▪ Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks (Kabir and Tiganj, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18510&sa=D&source=editors&ust=1779048536774889&usg=AOvVaw3-ui7JIXSmv333pbVwd3w6) |
| 32 | [▪ HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17093&sa=D&source=editors&ust=1779048536774924&usg=AOvVaw00lV-UvvMsr7U1C1yYm63W) |
| 33 | [▪ SafeDream: Safety World Model for Proactive Early Jailbreak Detection (Yan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16824&sa=D&source=editors&ust=1779048536774959&usg=AOvVaw0bLnRCRAOHlJiES3k-Bzlg) |
| 34 | [▪ Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization (Tang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15022&sa=D&source=editors&ust=1779048536774999&usg=AOvVaw2_86uQLdJ6uU4J7Y_2XB0u) |
| 35 | [▪ TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs (Shen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12232&sa=D&source=editors&ust=1779048536775035&usg=AOvVaw2qYhAo5QCY_P5d8bJ-TGHL) |
| 36 | [▪ Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion (Pramanik et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10326&sa=D&source=editors&ust=1779048536775070&usg=AOvVaw2VJn4jiIx3bgL-KyCtyIeg) |
| 37 | [▪ TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07727&sa=D&source=editors&ust=1779048536775106&usg=AOvVaw3JqHSv01kuoeVFHWPJP9zC) |
| 38 | [▪ Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents (Ding et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07809&sa=D&source=editors&ust=1779048536775143&usg=AOvVaw1EA0ZtX-1ShD778hOrbFya) |
| 39 | [▪ SelfGrader: Stable Jailbreak Detection for Large Language Models using Token-Level Logits (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01473&sa=D&source=editors&ust=1779048536775178&usg=AOvVaw2ej5yYppVx2imPUbxlMxK6) |
| 40 | [▪ Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses (Mishra and Varshney, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12185&sa=D&source=editors&ust=1779048536775214&usg=AOvVaw32n3wfUhpbfd0S1npTWKoW) |
| 41 | [▪ LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops (Fu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.14493&sa=D&source=editors&ust=1779048536775251&usg=AOvVaw3TaBmLj0NcKQhruU5WTI0o) |
| 42 | [▪ Metaphor-based Jailbreak Attacks on Text-to-Image Models (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10766&sa=D&source=editors&ust=1779048536775286&usg=AOvVaw0EtJI4YxqsHR5Bkkk72jPV) |
| 43 | [▪ Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23269&sa=D&source=editors&ust=1779048536775322&usg=AOvVaw24uU6jNefa3haM53PZ4l0B) |
| 44 | [▪ Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models (Hong et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20122&sa=D&source=editors&ust=1779048536775357&usg=AOvVaw33R7Aj_BMGbxcaXU1O14Sm) |
| 45 | [▪ Resource Consumption Threats in Large Language Models (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16068&sa=D&source=editors&ust=1779048536775390&usg=AOvVaw2Ua109sVS7Un7SoruF_8Sk) |
| 46 | [▪ Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt (Jenny et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14278&sa=D&source=editors&ust=1779048536775427&usg=AOvVaw3FCbB2QYQhPF3BzM66NdOK) |
| 47 | [▪ Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs (Ling et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13847&sa=D&source=editors&ust=1779048536775463&usg=AOvVaw2_zSKPElhQQ76dNOPmcP2j) |
| 48 | [▪ Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13420&sa=D&source=editors&ust=1779048536775497&usg=AOvVaw0pO8yY0LoAFv8D1fhHDW9v) |
| 49 | [▪ Colluding LoRA: A Composite Attack on LLM Safety Alignment (Ding, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12681&sa=D&source=editors&ust=1779048536775531&usg=AOvVaw3gljJ3mvKzFfmEZCgm5pm7) |
| 50 | [▪ Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks (Geng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16765&sa=D&source=editors&ust=1779048536775566&usg=AOvVaw3zoXdzFSbTQNocOQOiBlEU) |
| 51 | [▪ Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models (Wang, Balashankar, and Chandrasekaran, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11149&sa=D&source=editors&ust=1779048536775602&usg=AOvVaw1Fh8YRKyTAS9qtGF0PJIuk) |
| 52 | [▪ Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.11106&sa=D&source=editors&ust=1779048536775637&usg=AOvVaw0T3vLYFCtndv57Ah3BCYUj) |
| 53 | [▪ Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models (Zou et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09246&sa=D&source=editors&ust=1779048536775673&usg=AOvVaw370fJ7bFVOogsWxlxDo3He) |
| 54 | [▪ PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17277&sa=D&source=editors&ust=1779048536775707&usg=AOvVaw3AvKKil2jUkEIgWlAh0o1U) |
| 55 | [▪ Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07028&sa=D&source=editors&ust=1779048536775742&usg=AOvVaw1O1iG83qfWoMg2CF5rgxmx) |
| 56 | [▪ SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge (Ying et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13127&sa=D&source=editors&ust=1779048536775777&usg=AOvVaw3Qw01yRN-Xw6HBlhl90Thv) |
| 57 | [▪ Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05772&sa=D&source=editors&ust=1779048536775811&usg=AOvVaw21SOKtGfR0szSI9J7XqEFX) |
| 58 | [▪ When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.20376&sa=D&source=editors&ust=1779048536775847&usg=AOvVaw2oiMUq9hi9musZPUPw3BEb) |
| 59 | [▪ BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage (Nakka and Saxena, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.02479&sa=D&source=editors&ust=1779048536775882&usg=AOvVaw2vzzZhv91pAwVJ_vQxefS1) |
| 60 | [▪ Quantifying Frontier LLM Capabilities for Container Sandbox Escape (Marchand et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02277&sa=D&source=editors&ust=1779048536775917&usg=AOvVaw1s8YREuWEIOz0zkEhDsz-f) |
| 61 | [▪ MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00565&sa=D&source=editors&ust=1779048536775952&usg=AOvVaw0oHUeeb2MfrkUJfCVSPZP8) |
| 62 | [▪ Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents (Dong, Feng, and Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00902&sa=D&source=editors&ust=1779048536775992&usg=AOvVaw3MeKWvXQdkY-TvAiZ-oIay) |
| 63 | [▪ Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking (Fang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.24009&sa=D&source=editors&ust=1779048536776027&usg=AOvVaw17338mh6DTRjvHs1M4aG7Z) |
| 64 | [▪ Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search (Huang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22983&sa=D&source=editors&ust=1779048536776062&usg=AOvVaw0V80FK60wMLkiSYedAqnuz) |
| 65 | [▪ Analysis of LLMs Against Prompt Injection and Jailbreak Attacks (Jaiswal et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22242&sa=D&source=editors&ust=1779048536776094&usg=AOvVaw2zLj9eQc3B_PSwMWUiXoSN) |
| 66 | [▪ A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness (Luo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.14297&sa=D&source=editors&ust=1779048536776129&usg=AOvVaw033b6I9c1_1SJVyS-93dZ6) |
| 67 | [▪ Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs (Coalson, Fang, and Hong, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.17778&sa=D&source=editors&ust=1779048536776164&usg=AOvVaw0ozx8RE8o_rpi7P0AhP_w1) |
| 68 | [▪ Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents (Shavit, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16520&sa=D&source=editors&ust=1779048536776199&usg=AOvVaw1ZnN12P63QeKqekhOlR468) |
| 69 | [▪ Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks (Hu, Robey, and Liu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.00187&sa=D&source=editors&ust=1779048536776235&usg=AOvVaw1TCcRDPyLVY1gtX_l_c73W) |
| 70 | [▪ AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks (Song, Xie, and Yin, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13547&sa=D&source=editors&ust=1779048536776270&usg=AOvVaw2N0AfTYwYroHLmvYMJ0UGO) |
| 71 | [▪ Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models (Kadali and Papalexakis, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11495&sa=D&source=editors&ust=1779048536776307&usg=AOvVaw0tVDOmMfdjP1xV_mUw69A0) |
| 72 | [▪ from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors (Yan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.00038&sa=D&source=editors&ust=1779048536776341&usg=AOvVaw0_qPKJU-my3J7klm6DGwad) |
| 73 | [▪ Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks (Yan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10148&sa=D&source=editors&ust=1779048536776378&usg=AOvVaw11QZWt0mV8uTNco0vIrw5y) |
| 74 | [▪ Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing (Lintelo, Wu, and Picek, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08741&sa=D&source=editors&ust=1779048536776413&usg=AOvVaw21IGLi2bZN1Q9z17CEjQUW) |
| 75 | [▪ ShallowJail: Steering Jailbreaks against Large Language Models (Liu, Pei, and Liu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07107&sa=D&source=editors&ust=1779048536776447&usg=AOvVaw3McbMhcSToLNO50J6OuHHj) |
| 76 | [▪ TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking (Yoon et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06440&sa=D&source=editors&ust=1779048536776482&usg=AOvVaw2vKXdSYFO7oifZjFqPJH6z) |
| 77 | [▪ TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking (Du et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06630&sa=D&source=editors&ust=1779048536776516&usg=AOvVaw3AmIShdtaBBPR-MPT5FVdP) |
| 78 | [▪ A Causal Perspective for Enhancing Jailbreak Attack and Defense (Pan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04893&sa=D&source=editors&ust=1779048536776550&usg=AOvVaw2P3s0CLLJUhll61Iijg40O) |
| 79 | [▪ Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models (Xiong et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04896&sa=D&source=editors&ust=1779048536776586&usg=AOvVaw0F_XhJpGK9Ns4OZ_P6y9LS) |
| 80 | [▪ When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs (Dingeto et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.03365&sa=D&source=editors&ust=1779048536776621&usg=AOvVaw3vSWdH7R8FR57H5xckhf-x) |
| 81 | [▪ How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04294&sa=D&source=editors&ust=1779048536776655&usg=AOvVaw0VSYHbO86RQbdvQYr0uzLI) |
| 82 | [▪ Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence (Fu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.04204&sa=D&source=editors&ust=1779048536776692&usg=AOvVaw264lPjpK3CsY7Lrna5f2TD) |
| 83 | [▪ David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning (Nellessen and Kachman, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02395&sa=D&source=editors&ust=1779048536776727&usg=AOvVaw3JjzBqJYjCFvzTwUMJo0Tu) |
| 84 | [▪ Jailbreaking LLMs via Calibration (Lu, Guo, and Kong, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00619&sa=D&source=editors&ust=1779048536776759&usg=AOvVaw14_5dIWza3cljpWt47jQPT) |
| 85 | [▪ Text is All You Need for Vision-Language Model Jailbreaking (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00420&sa=D&source=editors&ust=1779048536776792&usg=AOvVaw1L-XFcLQelm8tJg8F_PAAe) |
| 86 | [▪ ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack (Lin et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20903&sa=D&source=editors&ust=1779048536776826&usg=AOvVaw2MU5PQg865vYI5QR2RVTSh) |
| 87 | [▪ Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks? (Mu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06350&sa=D&source=editors&ust=1779048536776860&usg=AOvVaw2xtK8dEZoJTh2DO0CVjiAG) |
| 88 | [▪ Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning (Yang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.19180&sa=D&source=editors&ust=1779048536776894&usg=AOvVaw28dgIW1NoSusEPEdXTn6Cd) |
| 89 | [▪ Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models (Liang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09201&sa=D&source=editors&ust=1779048536776927&usg=AOvVaw0srrAiBODaOdbzz5eGEBBA) |
| 90 | [▪ LLMs Can Unlearn Refusal with Only 1,000 Benign Samples (Guo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19231&sa=D&source=editors&ust=1779048536776959&usg=AOvVaw0qkA6t6WcocRLtYo6HZ9a-) |
| 91 | [▪ LLM Jailbreak Detection for (Almost) Free! (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.14558&sa=D&source=editors&ust=1779048536776995&usg=AOvVaw3E9fx0GeKm8cU3nCb-SrA4) |
| 92 | [▪ GCG Attack On A Diffusion LLM (Neyroud and Corley, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14266&sa=D&source=editors&ust=1779048536777028&usg=AOvVaw1EyWZmg40CwYRKgye0UZeg) |
| 93 | [▪ Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs (Kaunismaa et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13528&sa=D&source=editors&ust=1779048536777061&usg=AOvVaw2_zdWjvizKZTFmKUHcEIT_) |
| 94 | [▪ TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning (Xie, Song, and Luo, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12460&sa=D&source=editors&ust=1779048536777095&usg=AOvVaw0EEnv3sKN5KO64s-6dkK4y) |
| 95 | [▪ AJAR: Adaptive Jailbreak Architecture for Red-teaming (Dou and Yang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10971&sa=D&source=editors&ust=1779048536777127&usg=AOvVaw28gqDyUHBHznXQVo7G8cHF) |
| 96 | [▪ SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails (Mou et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09321&sa=D&source=editors&ust=1779048536777161&usg=AOvVaw2HBZoXKgYd1R9niUWqaMwp) |
| 97 | [▪ From static to adaptive: immune memory-based jailbreak detection for large language models (Leng et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03356&sa=D&source=editors&ust=1779048536777196&usg=AOvVaw1aaVJhXeYHm7-rM4xJRkn-) |
| 98 | [▪ MacPrompt: Maraconic-guided Jailbreak against Text-to-Image Models (Ye et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07141&sa=D&source=editors&ust=1779048536777229&usg=AOvVaw0rH998Y0j8Y1JOigu7L7EZ) |
| 99 | [▪ PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline (Rao et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19011&sa=D&source=editors&ust=1779048536777266&usg=AOvVaw2uDcfFL9K0xK3We7PHaASP) |
| 100 | [▪ The Echo Chamber Multi-Turn LLM Jailbreak (Alobaid et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05742&sa=D&source=editors&ust=1779048536777298&usg=AOvVaw0HA99Bcf30tgGqdikE4IcO) |
| 101 | [▪ Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05466&sa=D&source=editors&ust=1779048536777333&usg=AOvVaw338f-FdoLy7qIGPxQfPC0j) |
| 102 | [▪ Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05445&sa=D&source=editors&ust=1779048536777365&usg=AOvVaw3lOXo9za5Oh8jVPEL967ri) |
| 103 | [▪ Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models (Das et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05339&sa=D&source=editors&ust=1779048536777398&usg=AOvVaw1ROvRWDjztLDLGp4qwhwkv) |
| 104 | [▪ Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs (Xing et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.10029&sa=D&source=editors&ust=1779048536777435&usg=AOvVaw3AkxWZd44LUl4CU4YCDTam) |
| 105 | [▪ Jailbreaking Safeguarded Text-to-Image Models via Large Language Models (Jiang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.01839&sa=D&source=editors&ust=1779048536777469&usg=AOvVaw03pLiles9jagJgO_b9vCXi) |
| 106 | [▪ $PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models (Choi et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05150&sa=D&source=editors&ust=1779048536777520&usg=AOvVaw31BTiq9Pn7PAXLi4scW72m) |
| 107 | [▪ Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks (Cunningham et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04603&sa=D&source=editors&ust=1779048536777558&usg=AOvVaw0oB94LqyFmGhwNOhN_0J31) |
| 108 | [▪ Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks (Nurlanov, Schmidt, and Bernard, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03420&sa=D&source=editors&ust=1779048536777593&usg=AOvVaw2Z25gTl_4pJYClLp0KnNlp) |
| 109 | [▪ Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models (Hu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03265&sa=D&source=editors&ust=1779048536777627&usg=AOvVaw3Vntj0bfGS985NLyl76jW_) |
| 110 | [▪ Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03594&sa=D&source=editors&ust=1779048536777662&usg=AOvVaw0G_8I6Sc8paRw8P4IIDozM) |
| 111 | [▪ TRYLOCK: Defense-in-Depth Against LLM Jailbreaks via Layered Preference and Representation Engineering (Thornton, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03300&sa=D&source=editors&ust=1779048536777697&usg=AOvVaw3MWbZtR8Nnnul2Eh_UVhFH) |
| 112 | [▪ How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03288&sa=D&source=editors&ust=1779048536777731&usg=AOvVaw2Ygx12s4tPEMdx1ILaOCfn) |
| 113 | [▪ JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03005&sa=D&source=editors&ust=1779048536777767&usg=AOvVaw3MijEea5oxM38517ered9d) |
| 114 | [▪ Beyond Prompts: Space-Time Decoupling Control-Plane Jailbreaks in LLM Structured Output (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.24191&sa=D&source=editors&ust=1779048536777801&usg=AOvVaw1cQ88fIHI4Cfp_bs6SQg5p) |
| 115 | [▪ Emoji-Based Jailbreaking of Large Language Models (Gopinadh and Hussain, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00936&sa=D&source=editors&ust=1779048536777834&usg=AOvVaw0BkYnwdYhg2KOVdFF39uRE) |
| 116 | [▪ Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak (Gu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00213&sa=D&source=editors&ust=1779048536777870&usg=AOvVaw0-2o5dPZ29YwridnOR2RFb) |
| 117 | [▪ RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.12897&sa=D&source=editors&ust=1779048536777907&usg=AOvVaw0XiiQZveS_5sAtk0wxpf6L) |
| 118 | [▪ Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks (Gohil, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.08990&sa=D&source=editors&ust=1779048536777940&usg=AOvVaw3sehmKOoyAu1NMZjkHWsGa) |
| 119 | [▪ Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service (Zhang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.24415&sa=D&source=editors&ust=1779048536777975&usg=AOvVaw0pjBPig__vbITlPlRL81e1) |
| 120 | [▪ Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race? (Xin et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.24044&sa=D&source=editors&ust=1779048536778026&usg=AOvVaw13gRVnfCFRi78Ski6Hya9j) |
| 121 | [▪ Involuntary Jailbreak: On Self-Prompting Attacks (Guo, Li, and Kankanhalli, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.13246&sa=D&source=editors&ust=1779048536778061&usg=AOvVaw2w9M5E6rbuiLka99vIaEFr) |
| 122 | [▪ EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion (Liang, Huang, and Chen, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23173&sa=D&source=editors&ust=1779048536778098&usg=AOvVaw1JHrAPA1NZcRz3nImAjMoX) |
| 123 | [▪ X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability (Lu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.09990&sa=D&source=editors&ust=1779048536778135&usg=AOvVaw1JSffs94sOVPBxF6EuqCSc) |
| 124 | [▪ Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses (Shang, Wei, and Bai, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.02080&sa=D&source=editors&ust=1779048536778174&usg=AOvVaw0jmc4hfVeKoeU8HZIE4hOT) |
| 125 | [▪ Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking (Huang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21236&sa=D&source=editors&ust=1779048536778208&usg=AOvVaw0UmWaKpDX5cPVbQ51rFyyZ) |
| 126 | [▪ GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs (Wu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21008&sa=D&source=editors&ust=1779048536778241&usg=AOvVaw3uOsgqY9DtQiQAt0lP8NFT) |
| 127 | [▪ Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20168&sa=D&source=editors&ust=1779048536778276&usg=AOvVaw3HHjbX7l9IjWxr5H8KvhOs) |
| 128 | [▪ Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17231&sa=D&source=editors&ust=1779048536778310&usg=AOvVaw3qEw20u1BvnT6CSEHV8Vss) |
| 129 | [▪ Universal Jailbreak Suffixes Are Strong Attention Hijackers (Ben-Tov, Geva, and Sharif, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12880&sa=D&source=editors&ust=1779048536778344&usg=AOvVaw2LwFkYg7zGFqRDzDT_7Wey) |
| 130 | [▪ Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.11185&sa=D&source=editors&ust=1779048536778379&usg=AOvVaw1-0F0A8SdGQMcopZzzBXAA) |
| 131 | [▪ Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline (Rao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19011&sa=D&source=editors&ust=1779048536778414&usg=AOvVaw1UVfjYU9ZC-9MFabGj_des) |
| 132 | [▪ Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation (Liu and Lin, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18244&sa=D&source=editors&ust=1779048536778450&usg=AOvVaw33Z0w33GY-9fCykkis9GJq) |
| 133 | [▪ One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs (Tan, Yu, and Sakuma, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14751&sa=D&source=editors&ust=1779048536778485&usg=AOvVaw1UwtqUNOura1WZeMCj7wZq) |
| 134 | [▪ Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models (Yang, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13703&sa=D&source=editors&ust=1779048536778518&usg=AOvVaw3hgFZOByCKOu5zMNUe331E) |
| 135 | [▪ Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring (Hua et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12069&sa=D&source=editors&ust=1779048536778553&usg=AOvVaw3hJizkooT8oiCMCO_Bxqki) |
| 136 | [▪ Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously (Adiletta et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11783&sa=D&source=editors&ust=1779048536778588&usg=AOvVaw2At7euuV1BXXN6AsMPh-6W) |
| 137 | [▪ Metaphor-based Jailbreaking Attacks on Text-to-Image Models (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10766&sa=D&source=editors&ust=1779048536778621&usg=AOvVaw0HRJJOa9TrEwToCOcM9WxG) |
| 138 | [▪ A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties (Wang, Zhang, and Yagemann, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08185&sa=D&source=editors&ust=1779048536778660&usg=AOvVaw1yQn_eGiM8icPkaTT5pFf9) |
| 139 | [▪ OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation (Jia et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06589&sa=D&source=editors&ust=1779048536778696&usg=AOvVaw26M9DagxINzNpivnMRo71S) |
| 140 | [▪ Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence (Huang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06387&sa=D&source=editors&ust=1779048536778733&usg=AOvVaw3iaSJTixuDRKsAglNi23Mf) |
| 141 | [▪ TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.05485&sa=D&source=editors&ust=1779048536778768&usg=AOvVaw3ZHfWb4APFMd1P0us2Ns5q) |
| 142 | [▪ The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search (Wei et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01353&sa=D&source=editors&ust=1779048536778803&usg=AOvVaw06uCnMQOTawvlcDk1m_HZG) |
| 143 | [▪ SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.01513&sa=D&source=editors&ust=1779048536778839&usg=AOvVaw2I59FF6b0nH_6EvZ38S7zy) |
| 144 | [▪ Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.03111&sa=D&source=editors&ust=1779048536778873&usg=AOvVaw0AoNnb6A0iLg1p1gS9TCXp) |
| 145 | [▪ Immunity memory-based jailbreak detection: multi-agent adaptive guard for large language models (Leng, Zhang, and Zhang, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03356&sa=D&source=editors&ust=1779048536778907&usg=AOvVaw2KRWmZ2HXmlE038f_-8E0a) |
| 146 | [▪ Lockpicking LLMs: A Logit-Based Jailbreak Using Token-level Manipulation (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.13068&sa=D&source=editors&ust=1779048536778940&usg=AOvVaw3QbrSOXS8baFBxYnGv5er-) |
| 147 | [▪ Involuntary Jailbreak (Guo, Li, and Kankanhalli, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.13246&sa=D&source=editors&ust=1779048536778972&usg=AOvVaw1tk4ZvRDs-15K6iTZaVICJ) |
| 148 | [▪ A Wolf in Sheep's Clothing: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search (Wei et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01353&sa=D&source=editors&ust=1779048536779014&usg=AOvVaw3_jiaICmDvJTC6P1llhrYg) |
| 149 | [▪ DefenSee: Dissecting Threat from Sight and Text - A Multi-View Defensive Pipeline for Multi-modal Jailbreaks (Wang, Fok, and Thing, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01185&sa=D&source=editors&ust=1779048536779051&usg=AOvVaw0exRyoXn11jTuCDtP2srzp) |
| 150 | [▪ Jailbreaking and Mitigation of Vulnerabilities in Large Language Models (Peng et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.15236&sa=D&source=editors&ust=1779048536779085&usg=AOvVaw1bLS3BSsE2zIid4uruQ10L) |
| 151 | [▪ Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations (Wong et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18933&sa=D&source=editors&ust=1779048536779120&usg=AOvVaw3SBHfS31BqgOX3eDrvPqGg) |
| 152 | [▪ TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18581&sa=D&source=editors&ust=1779048536779153&usg=AOvVaw3YwBXqwZR8XYGOqBMquLR-) |
| 153 | [▪ Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17874&sa=D&source=editors&ust=1779048536779188&usg=AOvVaw2BHhHRZT-Ljw-a5FoDzdCX) |
| 154 | [▪ Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.17987&sa=D&source=editors&ust=1779048536779222&usg=AOvVaw1SUOk8ZP0tl5ZjjKnBb0qd) |
| 155 | [▪ Practical and Stealthy Touch-Guided Jailbreak Attacks on Deployed Mobile Vision-Language Agents (Ding et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07809&sa=D&source=editors&ust=1779048536779256&usg=AOvVaw2XVsMiP_kFqtYsSrpRZCdT) |
| 156 | [▪ LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution (Yang and Zhang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.01533&sa=D&source=editors&ust=1779048536779291&usg=AOvVaw2BH1qGS2LXLiSGcW7tR3sZ) |
| 157 | [▪ The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16347&sa=D&source=editors&ust=1779048536779326&usg=AOvVaw3b6jl9P4aJfPdfVL78IwF9) |
| 158 | [▪ "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios (Sun et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16278&sa=D&source=editors&ust=1779048536779361&usg=AOvVaw2lKrNcvXUNXGigsGfP0YB9) |
| 159 | [▪ Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments (Nathanson, Williams, and Matuszek, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13788&sa=D&source=editors&ust=1779048536779397&usg=AOvVaw2KnuGjSSkKvuNXdKQ2cCla) |
| 160 | [▪ Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security (Kim, Na, and Choi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14140&sa=D&source=editors&ust=1779048536779431&usg=AOvVaw1M1tLFeOSmMCtTvrrzdmzW) |
| 161 | [▪ VEIL: Jailbreaking Text-to-Video Models via Visual Exploitation from Implicit Language (Ying et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13127&sa=D&source=editors&ust=1779048536779465&usg=AOvVaw0b8cwxd-7rH_Chh6GpHjuI) |
| 162 | [▪ Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs (Chen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12710&sa=D&source=editors&ust=1779048536779502&usg=AOvVaw1rHwDM7IaIP-0Vyl5TIbfJ) |
| 163 | [▪ ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models (Cheng et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13548&sa=D&source=editors&ust=1779048536779536&usg=AOvVaw3kTlFf_MgDru44WWuqvukB) |
| 164 | [▪ NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks (Sleem et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11784&sa=D&source=editors&ust=1779048536779569&usg=AOvVaw3QxbDYBseevImov9buc8V3) |
| 165 | [▪ Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation (Heiding and Lermen, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11759&sa=D&source=editors&ust=1779048536779603&usg=AOvVaw0W2aJKdTAjlEULKSpHW2uY) |
| 166 | [▪ GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events (Reworr, Petrov, and Volkov, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04860&sa=D&source=editors&ust=1779048536779638&usg=AOvVaw02EXky9eL9jLwmKc4DGDJh) |
| 167 | [▪ Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives (Chang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17519&sa=D&source=editors&ust=1779048536779673&usg=AOvVaw1mk79POsPKwvfw8inSUHS9) |
| 168 | [▪ Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation (Schwartz et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18638&sa=D&source=editors&ust=1779048536779709&usg=AOvVaw0DhWaZfUsdsqozUhn3B2Fn) |
| 169 | [▪ UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning (Zhang, Yang, and Li, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.01908&sa=D&source=editors&ust=1779048536779745&usg=AOvVaw1wfDD_po744p1mASpkQ-kw) |
| 170 | [▪ Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs (Zhou et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08367&sa=D&source=editors&ust=1779048536779778&usg=AOvVaw2zomUWTj9DWcH2cB1TCGhQ) |
| 171 | [▪ KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07480&sa=D&source=editors&ust=1779048536779812&usg=AOvVaw3ZsWQ2G16PjSkBiuzPajDT) |
| 172 | [▪ HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.13677&sa=D&source=editors&ust=1779048536779845&usg=AOvVaw3Ujz6cpfdKOPLQ2LQwxZoC) |
| 173 | [▪ JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models (Jin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.01599&sa=D&source=editors&ust=1779048536779880&usg=AOvVaw1H2T9M-yYacepeEhKpuQSa) |
| 174 | [▪ JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework (Zhou et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07315&sa=D&source=editors&ust=1779048536779913&usg=AOvVaw3x26PyV530mIA9SG9mYOSZ) |
| 175 | [▪ Jailbreaking in the Haystack (Shah et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04707&sa=D&source=editors&ust=1779048536779944&usg=AOvVaw2AU-OVH5sphMOYdPURWhS9) |
| 176 | [▪ GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs (Basani and Zhang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14133&sa=D&source=editors&ust=1779048536779978&usg=AOvVaw1Ia-L9AvwNchAADrW1tiDM) |
| 177 | [▪ VERA: Variational Inference Framework for Jailbreaking Large Language Models (Lochab et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.22666&sa=D&source=editors&ust=1779048536780015&usg=AOvVaw3uFgJ5W2fNoNp4pOXIWfG5) |
| 178 | [▪ Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models (Yang and Fu, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.23558&sa=D&source=editors&ust=1779048536780051&usg=AOvVaw3DHmXYh7FLAGfj8IucZ2y0) |
| 179 | [▪ Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs (Liu, Hou, and Sui, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03271&sa=D&source=editors&ust=1779048536780086&usg=AOvVaw3HPu0iij5V1Y-BYB-LxgQi) |
| 180 | [▪ AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models (Reddy, Zagula, and Saban, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02376&sa=D&source=editors&ust=1779048536780122&usg=AOvVaw0px-C8xXCZ2ELJNoMzuOSE) |
| 181 | [▪ An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02356&sa=D&source=editors&ust=1779048536780156&usg=AOvVaw1OAUxSWhm9ULBqajnKGLHp) |
| 182 | [▪ Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.16406&sa=D&source=editors&ust=1779048536780191&usg=AOvVaw3IcjqpK4j8s3vjXZWW70DR) |
| 183 | [▪ What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks (Kirch et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.03343&sa=D&source=editors&ust=1779048536780225&usg=AOvVaw2JDC4n-At8sIOupUNUMBxm) |
| 184 | [▪ Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks (Paim et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00346&sa=D&source=editors&ust=1779048536780271&usg=AOvVaw3jNMyyz8fuEjzm8Mz_T_7o) |
| 185 | [▪ Sentra-Guard: A Multilingual Human-AI Framework for Real-Time Defense Against Adversarial LLM Jailbreaks (Hasan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22628&sa=D&source=editors&ust=1779048536780306&usg=AOvVaw0lH4Ac5qLPK49JLqYgBKW6) |
| 186 | [▪ Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models (Ntais, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22085&sa=D&source=editors&ust=1779048536780340&usg=AOvVaw10LhPW5x93G3UEvoJU-31y) |
| 187 | [▪ Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses (Zhong, Fok, and Thing, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21214&sa=D&source=editors&ust=1779048536780373&usg=AOvVaw2-GFHt-SGbH3S4rTmVYQWZ) |
| 188 | [▪ The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21190&sa=D&source=editors&ust=1779048536780410&usg=AOvVaw1-gWHTCGyD5_SUXH_49xzh) |
| 189 | [▪ Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency (Jiang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21189&sa=D&source=editors&ust=1779048536780445&usg=AOvVaw0He_f3NFEyfOMRx-1UNbjn) |
| 190 | [▪ Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training (Yong and Bach, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20956&sa=D&source=editors&ust=1779048536780480&usg=AOvVaw3JmWKxQ3RlqC_ms66RLHbW) |
| 191 | [▪ Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations (Kumar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20223&sa=D&source=editors&ust=1779048536780516&usg=AOvVaw1Par6-CsgGJPugnyD9f5dB) |
| 192 | [▪ Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models (Liang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09201&sa=D&source=editors&ust=1779048536780549&usg=AOvVaw2Fl_mplx4KK7BAvcuyj_LP) |
| 193 | [▪ HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models (Narula et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18728&sa=D&source=editors&ust=1779048536780583&usg=AOvVaw17HZ0pVKmDEoWE2jParfiM) |
| 194 | [▪ BreakFun: Jailbreaking LLMs via Schema Exploitation (Oskooei and Aktas, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17904&sa=D&source=editors&ust=1779048536780615&usg=AOvVaw0Y8cFMNSY55U7oXj7BwGiS) |
| 195 | [▪ VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models (Liao, Lochab, and Zhang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17759&sa=D&source=editors&ust=1779048536780649&usg=AOvVaw0gNUd3HDe4nwM6B2nicVHF) |
| 196 | [▪ Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17277&sa=D&source=editors&ust=1779048536780684&usg=AOvVaw2ZRlJ9MX2ZlWVrdjX08uRZ) |
| 197 | [▪ PRISON: Unmasking the Criminal Potential of Large Language Models (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.16150&sa=D&source=editors&ust=1779048536780717&usg=AOvVaw3CkmVLLLuQE1h1v7TTLhQx) |
| 198 | [▪ Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15068&sa=D&source=editors&ust=1779048536780752&usg=AOvVaw2jhI8rkf6RJguLwoBijobI) |
| 199 | [▪ Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks (Wu, Wang, and Liao, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15017&sa=D&source=editors&ust=1779048536780787&usg=AOvVaw2flqkGa_-p44ycug_6-MAs) |
| 200 | [▪ SoK: Evaluating Jailbreak Guardrails for Large Language Models (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10597&sa=D&source=editors&ust=1779048536780819&usg=AOvVaw3J9Mz59dhaqbkCTtOiHLnH) |
| 201 | [▪ Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.10390&sa=D&source=editors&ust=1779048536780853&usg=AOvVaw2dRNq3zhwLaL8VvL_Nc_qS) |
| 202 | [▪ Bag of Tricks for Subverting Reasoning-based Safety Guardrails (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11570&sa=D&source=editors&ust=1779048536780886&usg=AOvVaw26wnjEnaVmAZH3MBYBnEZl) |
| 203 | [▪ ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10281&sa=D&source=editors&ust=1779048536780920&usg=AOvVaw0Wa6fNZ-S5XR0S6UrEn2yg) |
| 204 | [▪ MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation (Zhu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10271&sa=D&source=editors&ust=1779048536780954&usg=AOvVaw342KfYogxi_t9PUTLlSBn2) |
| 205 | [▪ The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections (Nasr et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09023&sa=D&source=editors&ust=1779048536780994&usg=AOvVaw3LLVSCyfbBEJaJrR5AVWHA) |
| 206 | [▪ Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models (Nihal et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08859&sa=D&source=editors&ust=1779048536781030&usg=AOvVaw0-pTlNJsN8WcIkWSlcVNiA) |
| 207 | [▪ PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.01444&sa=D&source=editors&ust=1779048536781064&usg=AOvVaw3xB5KW03hbKNycdz14esGi) |
| 208 | [▪ MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation (Jiang and Pan, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07835&sa=D&source=editors&ust=1779048536781097&usg=AOvVaw0DWbf3RvW8H49C-JPcewZM) |
| 209 | [▪ Effective and Stealthy One-Shot Jailbreaks on Deployed Mobile Vision-Language Agents (Ding et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07809&sa=D&source=editors&ust=1779048536781134&usg=AOvVaw1e_goq_UXVYWwwo7ZIPGlA) |
| 210 | [▪ Jailbreak Attack Initializations as Extractors of Compliance Directions (Levi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.09755&sa=D&source=editors&ust=1779048536781167&usg=AOvVaw2A4kCtGMFLQRGqPaEKFpya) |
| 211 | [▪ AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling (Liu and Xiao, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05379&sa=D&source=editors&ust=1779048536781201&usg=AOvVaw1kJ0GMn6dgjHnUjIrdk5xb) |
| 212 | [▪ Auditing Pay-Per-Token in Large Language Models (Velasco, Tsirtsis, and Gomez-Rodriguez, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05181&sa=D&source=editors&ust=1779048536781235&usg=AOvVaw0F7Dg3WubVurO9gPf6tuzx) |
| 213 | [▪ DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization (Huang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.18564&sa=D&source=editors&ust=1779048536781270&usg=AOvVaw1pDp7SF9gsLGYI3jUjwhvy) |
| 214 | [▪ Imperceptible Jailbreaking against Large Language Models (Gao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05025&sa=D&source=editors&ust=1779048536781304&usg=AOvVaw2FNxD3OYEvxoZsYRpTiD4w) |
| 215 | [▪ NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks (Asl et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03417&sa=D&source=editors&ust=1779048536781341&usg=AOvVaw2jYWCNaQemXwrOMRLukzRB) |
| 216 | [▪ JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models (Peng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17568&sa=D&source=editors&ust=1779048536781377&usg=AOvVaw0oWVlJrz-aD6veTb10-Cjo) |
| 217 | [▪ XBreaking: Explainable Artificial Intelligence for Jailbreaking LLMs (Arazzi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21700&sa=D&source=editors&ust=1779048536781411&usg=AOvVaw0NBfCMHxD5jL9DjH8ml8A4) |
| 218 | [▪ PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips (Coalson et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.07192&sa=D&source=editors&ust=1779048536781446&usg=AOvVaw3cLvjCDODNy-bJSHrsZoqY) |
| 219 | [▪ Untargeted Jailbreak Attack (Huang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02999&sa=D&source=editors&ust=1779048536781479&usg=AOvVaw3gaqU2KelyGfFRlaYhdNfi) |
| 220 | [▪ Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs (Xie, Song, and Luo, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02833&sa=D&source=editors&ust=1779048536781513&usg=AOvVaw1Vg_wfM3AWuX125c_KraLy) |
| 221 | [▪ Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks (Saha et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01359&sa=D&source=editors&ust=1779048536781548&usg=AOvVaw2S5Ah13j-JUH55AEUvkqM6) |
| 222 | [▪ Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach (Li, Wang, and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01342&sa=D&source=editors&ust=1779048536781584&usg=AOvVaw1wkYF72qqtuqOdQ0HlViH9) |
| 223 | [▪ Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge (Dou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01223&sa=D&source=editors&ust=1779048536781619&usg=AOvVaw3A8udthajTbUjFsLj-ITF-) |
| 224 | [▪ AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models (Liang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.10846&sa=D&source=editors&ust=1779048536781653&usg=AOvVaw0J67wu1B_GxaeS2C0cYf7m) |
| 225 | [▪ STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25624&sa=D&source=editors&ust=1779048536781687&usg=AOvVaw3mbew5N_-jKHD_cs_aWNPz) |
| 226 | [▪ Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents (Das, Sandler, and Fioretto, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10171&sa=D&source=editors&ust=1779048536781722&usg=AOvVaw3WScMEdCOWpQPzXRpEN2Hn) |
| 227 | [▪ Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning (Wang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23558&sa=D&source=editors&ust=1779048536781756&usg=AOvVaw2MzrSkIt-vG2r6IQ8d9sz3) |
| 228 | [▪ Takedown: How It's Done in Modern Coding Agent Exploits (Lee et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24240&sa=D&source=editors&ust=1779048536781790&usg=AOvVaw1OklsDYGXqpC90ercfKryJ) |
| 229 | [▪ Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks (Tong et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22732&sa=D&source=editors&ust=1779048536781825&usg=AOvVaw3D0GzgBXB-r-X9sCXhArX3) |
| 230 | [▪ Activation-Guided Local Editing for Jailbreaking Attacks (Wang et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00555&sa=D&source=editors&ust=1779048536781858&usg=AOvVaw2QarX8WxWoJAz-MUKvM9nM) |
| 231 | [▪ Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application (Xu et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.06989&sa=D&source=editors&ust=1779048536781893&usg=AOvVaw1ufQliQ05KxPnfNLV_OjUu) |
| 232 | [▪ Enhancing Jailbreak Attacks on LLMs via Persona Prompts (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22171&sa=D&source=editors&ust=1779048536781925&usg=AOvVaw2V3K0LPzkw-Ge_xgOYwoCb) |
| 233 | [▪ PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking (Zou et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21540&sa=D&source=editors&ust=1779048536781960&usg=AOvVaw2PLE2NXMC6gP8wABFxygsi) |
| 234 | [▪ NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples (Shi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21483&sa=D&source=editors&ust=1779048536781996&usg=AOvVaw0ZaEe_TBUMacAyhaErIR8Q) |
| 235 | [▪ How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework (Liang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19219&sa=D&source=editors&ust=1779048536782036&usg=AOvVaw2lpO24P2-KmoSRpjGSjyrp) |
| 236 | [▪ From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem (Mao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15170&sa=D&source=editors&ust=1779048536782072&usg=AOvVaw2pEpk0lmmmo9XQLJQPmEle) |
| 237 | [▪ Exploiting Jailbreaking Vulnerabilities in Generative AI to Bypass Ethical Safeguards for Facilitating Phishing Attacks (Mishra and Varshney, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12185&sa=D&source=editors&ust=1779048536782107&usg=AOvVaw3L3D9a9jSuvWCeyxBLTWNp) |
| 238 | [▪ Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility (Murphy et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11630&sa=D&source=editors&ust=1779048536782140&usg=AOvVaw36FH82IrpC6C2r65D88u8k) |
| 239 | [▪ Representation Bending for Large Language Model Safety (Yousefpour et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.01550&sa=D&source=editors&ust=1779048536782173&usg=AOvVaw2EFfCKulBVPaVMR8VOfdvm) |
| 240 | [▪ GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07735&sa=D&source=editors&ust=1779048536782207&usg=AOvVaw1IAVQM0freX-Cmekf6VWEh) |
| 241 | [▪ Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07139&sa=D&source=editors&ust=1779048536782242&usg=AOvVaw0wgkz4vrcrpCdtUQh9gRZt) |
| 242 | [▪ GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods (Huang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16903&sa=D&source=editors&ust=1779048536782277&usg=AOvVaw2ZWGQUTCbgyt_bIvnblUIe) |
| 243 | [▪ On Jailbreaking Quantized Language Models Through Fault Injection Attacks (Zahran et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.03236&sa=D&source=editors&ust=1779048536782311&usg=AOvVaw0Sw7LiWt_lrW27xIHzoa6F) |
| 244 | [▪ Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs (Pu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.16327&sa=D&source=editors&ust=1779048536782345&usg=AOvVaw26nz9eoXN-jHm3u8ukWuBI) |
| 245 | [▪ CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06043&sa=D&source=editors&ust=1779048536782381&usg=AOvVaw1sSUhtrkPFaYVgEpYqpdv8) |
| 246 | [▪ Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers\<br>\<br> (Nguyen et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16241&sa=D&source=editors&ust=1779048536782417&usg=AOvVaw242YaQBHUjTJezNkPMDcWZ) |
| 247 | [▪ Prompt Inference Attack on Distributed Large Language Model Inference Frameworks (Luo, Yu, and Xiao, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.09291&sa=D&source=editors&ust=1779048536782453&usg=AOvVaw2MnUvyNYRlzdW1IWgPaMBn) |
| 248 | [▪ JULI: Jailbreak Large Language Models by Self-Introspection (Wang, Hu, and Wagnetr, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11790&sa=D&source=editors&ust=1779048536782486&usg=AOvVaw1AmlfbyX7CURuu70MK2S38) |
| 249 | [▪ AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models (Chen et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.14103&sa=D&source=editors&ust=1779048536782520&usg=AOvVaw23XWzArdJJMjpcB_6BpgD1) |
| 250 | [▪ Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search (Arel and Zhou, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.10619&sa=D&source=editors&ust=1779048536782555&usg=AOvVaw3wX-yeG5wnniK35hswhBGm) |
| 251 | [▪ Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate (Qi et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16489&sa=D&source=editors&ust=1779048536782589&usg=AOvVaw1Qj4wc2QSK9csfENv2Cwvx) |
| 252 | [▪ Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models (Ke et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.20320&sa=D&source=editors&ust=1779048536782623&usg=AOvVaw0o7qGj1qDquky8BDQhup78) |
| 253 | [▪ h4rm3l: A language for Composable Jailbreak Attack Synthesis (Doumbouya et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.04811&sa=D&source=editors&ust=1779048536782655&usg=AOvVaw2iF1RIrV6QJBV-l8pcci9n) |
| 254 | [▪ Siege: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search (Zhou, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.10619&sa=D&source=editors&ust=1779048536782689&usg=AOvVaw3QQFobOY7eIFhjouCi3GJM) |
| 255 | [▪ Adversarial Training for Multimodal Large Language Models against Jailbreak Attacks (Lu et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.04833&sa=D&source=editors&ust=1779048536782722&usg=AOvVaw19QIODdygMlHdTQIJyNh48) |
| 256 | [▪ Dagger Behind Smile: Fool LLMs with a Happy Ending Story (Song et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.13115&sa=D&source=editors&ust=1779048536782755&usg=AOvVaw2DsNrjbFkpp4lEkBmIVwHr) |
| 257 | [▪ Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks (Fang et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.04234&sa=D&source=editors&ust=1779048536782790&usg=AOvVaw3lnNqTJrfqkw3Sxly0rQ02) |
| 258 | [▪ An Engorgio Prompt Makes Large Language Model Babble on (Dong et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.19394&sa=D&source=editors&ust=1779048536782823&usg=AOvVaw2dZ8b43Q3nX6qHQQnPk-T6) |
| 259 | [▪ Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models (Ying et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.11054&sa=D&source=editors&ust=1779048536782857&usg=AOvVaw1okeXw6O3aJA9PUM1bAeLl) |
| 260 | [▪ LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models (Lin et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.16205&sa=D&source=editors&ust=1779048536782892&usg=AOvVaw1qZ6--SmA_MABbLXNeEpcV) |
| 261 | [▪ Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation (Armstrong et al, Jan 2025)](https://www.google.com/url?q=https://www.researchgate.net/publication/388555790_Defense_Against_the_Dark_Prompts_Mitigating_Best-of-N_Jailbreaking_with_Prompt_Evaluation&sa=D&source=editors&ust=1779048536782926&usg=AOvVaw3WIQ4LNzwUc7AgUCmK2eJj) |
| 262 | [▪ Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI (Ye et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16671&sa=D&source=editors&ust=1779048536782967&usg=AOvVaw2QVbfCp5CJ6hLSfLFgn4aW) |
| 263 | [▪ GreedyPixel: Fine-Grained Black-Box Adversarial Attack Via Greedy Algorithm (Wang et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.14230&sa=D&source=editors&ust=1779048536783018&usg=AOvVaw0Ri7Mtf6-V56ADHagU094T) |
| 264 | [▪ MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework (Guo et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.07251&sa=D&source=editors&ust=1779048536783053&usg=AOvVaw2KjhTH_Y0_zWCuBGbzxoAu) |
| 265 | [▪ Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors (Zhao and Zhang, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.14250&sa=D&source=editors&ust=1779048536783089&usg=AOvVaw2pk4_keAd07AIl0xBFyjqj) |
| 266 | [▪ Gandalf the Red: Adaptive Security for LLMs (Pfister et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.07927&sa=D&source=editors&ust=1779048536783122&usg=AOvVaw2UzMhWTQKIQUQsV7van7C0) |
| 267 | [▪ DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models (Yin, Liu, and Xu, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10561&sa=D&source=editors&ust=1779048536783158&usg=AOvVaw05WyVI4xOQJ8aajrp9Wayv) |
| 268 | [▪ Infecting Generative AI With Viruses (Noever and McKee, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.05542&sa=D&source=editors&ust=1779048536783191&usg=AOvVaw1AQ76Phy45DNRRmbpdChJm) |
| 269 | [▪ BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger (Chen et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09093&sa=D&source=editors&ust=1779048536783226&usg=AOvVaw3J2FrSpiuVhSWmnUIl88Px) |
| 270 | [▪ Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment (Ghosal et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.18688&sa=D&source=editors&ust=1779048536783261&usg=AOvVaw2NC0R_B8JhDcPC-7lMXZG8) |
| 271 | [▪ Automated Progressive Red Teaming (Jiang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.03876&sa=D&source=editors&ust=1779048536783292&usg=AOvVaw38OA4JJC1tFtqwFF3vLdT7) |
| 272 | [▪ Crabs: Consuming Resrouce via Auto-generation for LLM-DoS Attack under Black-box Settings (Zhang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.13879&sa=D&source=editors&ust=1779048536783326&usg=AOvVaw3KVQH8TsTYtfR70TAKW0A1) |
| 273 | [▪ LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses (Lin et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.04755&sa=D&source=editors&ust=1779048536783358&usg=AOvVaw0U3rQfJSIVg_YR1OjjfDoG) |
| 274 | [▪ When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search (Chen et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.08705&sa=D&source=editors&ust=1779048536783391&usg=AOvVaw28My5583i3cJcyt3LJReS1) |
| 275 | [▪ Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks (Liu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02220&sa=D&source=editors&ust=1779048536783425&usg=AOvVaw3u_N67JqihSQtXBRiY11ik) |
| 276 | [▪ DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints (Zhao et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.19026&sa=D&source=editors&ust=1779048536783461&usg=AOvVaw3rPmiE_mNpztHlJBP4qdsl) |
| 277 | [▪ SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage (Dong et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15289&sa=D&source=editors&ust=1779048536783494&usg=AOvVaw0_1vwREsBF3e5h1efjxi2P) |
| 278 | [▪ JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs (Li et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15623&sa=D&source=editors&ust=1779048536783528&usg=AOvVaw0Wa1pHtP66bVtTB8N2z3SL) |
| 279 | [▪ Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation (Gao et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15924&sa=D&source=editors&ust=1779048536783563&usg=AOvVaw2uNIYdgABjML8alqvgcRzZ) |
| 280 | [▪ CAMH: Advancing Model Hijacking Attack in Machine Learning (He et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.13741&sa=D&source=editors&ust=1779048536783595&usg=AOvVaw3ALkOz5qjTNPOjlVTxHKWM) |
| 281 | [▪ Best-of-N Jailbreaking (Hughes et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.03556&sa=D&source=editors&ust=1779048536783626&usg=AOvVaw3JC4s-K-cIGlZdxu5kuZ5K) |
| 282 | [▪ AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization (Li et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.11940&sa=D&source=editors&ust=1779048536783660&usg=AOvVaw2VmIsmk9fM2-SNYQX4ZfVX) |
| 283 | [▪ AdvPrefix: An Objective for Nuanced LLM Jailbreaks (Zhu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.10321%23&sa=D&source=editors&ust=1779048536783692&usg=AOvVaw3C5AigHVHKa8VKDYoUjl__) |
| 284 | [▪ Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach (Wang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.02159&sa=D&source=editors&ust=1779048536783730&usg=AOvVaw2U5TU0g83Eys7HaK-FDybt) |
| 285 | [▪ Stealthy Multi-Task Adversarial Attacks (Guo et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.17936&sa=D&source=editors&ust=1779048536783761&usg=AOvVaw1BAqlJjgFJyCxlCGM4UakU) |
| 286 | [▪ AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs (Liu et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.05295&sa=D&source=editors&ust=1779048536783795&usg=AOvVaw1jQTvYUKidTazzT_CWpuce) |
| 287 | [▪ BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models (Wang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.09804&sa=D&source=editors&ust=1779048536783830&usg=AOvVaw0_mI_VbRsXt9TwlamSRFDK) |
| 288 | [▪ JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks (Luo et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.03027&sa=D&source=editors&ust=1779048536783865&usg=AOvVaw2a3mVWMMxRuuAaqRTFc23U) |
| 289 | [▪ SQL Injection Jailbreak: a structural disaster of large language models (Zhao et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.01565&sa=D&source=editors&ust=1779048536783898&usg=AOvVaw0L-NW2JhXsr4x9MispKjnL) |
| 290 | [▪ LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs (Jha, Arora, and Ganesh, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.08862&sa=D&source=editors&ust=1779048536783931&usg=AOvVaw3FIoevGOZYO6YSTze_WaT4) |
| 291 | [▪ AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks (Zeng et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.04783&sa=D&source=editors&ust=1779048536783963&usg=AOvVaw3Uv3mVjfNVPWT4rV6Ingdc) |
| 292 | [▪ DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers (Li et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.16914&sa=D&source=editors&ust=1779048536784000&usg=AOvVaw10EtaR4Smnw9AYhFFvP_Yt) |
| 293 | [▪ SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains (Saiem et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.06426&sa=D&source=editors&ust=1779048536784036&usg=AOvVaw3ckZ722MmaWsU639v8rRuB) |
| 294 | [▪ Transferable Ensemble Black-box Jailbreak Attacks on Large Language Models (Yang and Fu, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.23558&sa=D&source=editors&ust=1779048536784070&usg=AOvVaw3PS12sRFNtZXYGNq1wQHrl) |
| 295 | [▪ Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning (Hasan, Rugina, and Wang, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2401.10862&sa=D&source=editors&ust=1779048536784105&usg=AOvVaw2Lty_uWG9LhqRL5-XvbSpf) |
| 296 | [▪ Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (Mehrotra et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.02119&sa=D&source=editors&ust=1779048536784137&usg=AOvVaw0NdrxoGmfgduf6Z4VnlUoK) |
| 297 | [▪ Fight Back Against Jailbreaking via Prompt Adversarial Tuning (Mo et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.06255&sa=D&source=editors&ust=1779048536784170&usg=AOvVaw2n8DUrTIv_sSjAhqhODMJj) |
| 298 | [▪ HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models (Zhang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22832&sa=D&source=editors&ust=1779048536784204&usg=AOvVaw3diURNw1kCsTfnGXrT7JHx) |
| 299 | [▪ Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector (Huang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22888&sa=D&source=editors&ust=1779048536784238&usg=AOvVaw1xDPsVV8lh7CXctpBOg7J1) |
| 300 | [▪ Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses (Zheng et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.01288&sa=D&source=editors&ust=1779048536784272&usg=AOvVaw1RWIT_8eIMUT_7v6Ln47xQ) |
| 301 | [▪ Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents (Yang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.11208&sa=D&source=editors&ust=1779048536784306&usg=AOvVaw04GTv2EHoxDM9XHO5Yq-_h) |
| 302 | [▪ Transferable Adversarial Attacks on SAM and Its Downstream Models (Xia et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.20197&sa=D&source=editors&ust=1779048536784338&usg=AOvVaw1yQHSZZslBQjXf_huRZc6i) |
| 303 | [▪ Remote Timing Attacks on Efficient Language Model Inference (Carlini and Nasr, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.17175&sa=D&source=editors&ust=1779048536784371&usg=AOvVaw2xsfxPb1y3XmxgwUgzgxJD) |
| 304 | [▪ MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models (Weng et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08464&sa=D&source=editors&ust=1779048536784406&usg=AOvVaw3kpSCR820hHMD28W_31rf6) |
| 305 | [▪ RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process (Wang, Liu, and Xiao, Oct 2024)](https://www.google.com/url?q=http://Wang,+Liu,+and+Xiao&sa=D&source=editors&ust=1779048536784442&usg=AOvVaw3_IKVqyq9FBn1ka2oiBl-j) |
| 306 | [▪ Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents (Kumar et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.13886&sa=D&source=editors&ust=1779048536784478&usg=AOvVaw1nGFM-8STHpOD1D2nnXHK5) |
| 307 | [▪ Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning (Xu, Gumaste, and Singh, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.09695&sa=D&source=editors&ust=1779048536784512&usg=AOvVaw2b6DExneS-RX8xvRycD7iq) |
| 308 | [▪ Adversarial Attacks on Large Language Models Using Regularized Relaxation (Chacko et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19160&sa=D&source=editors&ust=1779048536784546&usg=AOvVaw2AKrACLLKsGATZh1giesFW) |
| 309 | [▪ Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models (Li et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.15362&sa=D&source=editors&ust=1779048536784580&usg=AOvVaw0DLqGychySt4a3F18wHgSg) |
| 310 | [▪ Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables (Chen et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2309.06223&sa=D&source=editors&ust=1779048536784615&usg=AOvVaw1BHW5Z_gXBKZtcotQjVzys) |
| 311 | [▪ Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks (Peng et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.08087&sa=D&source=editors&ust=1779048536784649&usg=AOvVaw0pQ5ky7nww4P-T1NDbDCb9) |
| 312 | [▪ Effective and Evasive Fuzz Testing-Driven Jailbreaking Attacks against LLMs (Gong et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14866&sa=D&source=editors&ust=1779048536784682&usg=AOvVaw2kzdu6dxwq7TeKskdgootS) |
| 313 | [▪ Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks (Wang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.04234&sa=D&source=editors&ust=1779048536784717&usg=AOvVaw3ZAPnPZZ_hCOSpjpv4BDeS) |
| 314 | [▪ Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models (Shen et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02298&sa=D&source=editors&ust=1779048536784753&usg=AOvVaw0Vr4DsiiREcihK31isUIMl) |
| 315 | [▪ PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach (Lin et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14177&sa=D&source=editors&ust=1779048536784788&usg=AOvVaw0UljMuAEVKfijrvRHu8U2V) |
| 316 | [▪ Rethinking and Defending Protective Perturbation in Personalized Diffusion Models (Liu et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.18944&sa=D&source=editors&ust=1779048536784821&usg=AOvVaw2Vbe861nS7sT9YnNMC9Jki) |
| 317 | [▪ Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models (Yu et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.17336&sa=D&source=editors&ust=1779048536784856&usg=AOvVaw24nJiYCVTC2wU3lAgTGUP6) |
| 318 | [▪ Read Over the Lines: Attacking LLMs and Toxicity Detection Systems with ASCII Art to Mask Profanity (Berezin, Farahbakhsh, Crespi, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.18708&sa=D&source=editors&ust=1779048536784891&usg=AOvVaw2zpKSSnsJjV331uVY2FJ7T) |
| 319 | [▪ Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI (Rawat et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.15398&sa=D&source=editors&ust=1779048536784926&usg=AOvVaw0ErF7V1LAtHIP_x4DGqmbN) |
| 320 | [▪ Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction (Zhang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.16783&sa=D&source=editors&ust=1779048536784961&usg=AOvVaw1Dx_EKBsNQ4ONwtxmUK6Qj) |
| 321 | [▪ Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation (Shahariar et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.15381&sa=D&source=editors&ust=1779048536785000&usg=AOvVaw1GyZvf45q_Iw2V7SIRGSSx) |
| 322 | [▪ Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack (Russinovich, Salem, and Eldan, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.01833&sa=D&source=editors&ust=1779048536785036&usg=AOvVaw2rhz7gizim8dMjrntRI153) |
| 323 | [▪ VulZoo: A Comprehensive Vulnerability Intelligence Dataset (Ruan et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.16347&sa=D&source=editors&ust=1779048536785068&usg=AOvVaw2NwYG_XOjf2BcXG_jMTS8e) |
| 324 | [▪ Adversarial Attacks on Machine Learning-Aided Visualizations (Fujiwara et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.02485&sa=D&source=editors&ust=1779048536785101&usg=AOvVaw1T7wfQZFoxujoQ2CpAT5ND) |
| 325 | [▪ Jailbreaking Large Language Models with Symbolic Mathematics (Bethany et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11445&sa=D&source=editors&ust=1779048536785133&usg=AOvVaw2ILKXXYt2wuLeHp1pcOKdZ) |
| 326 | [▪ Image Hijacks: Adversarial Images can Control Generative Models at Runtime (Bailey et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2309.00236&sa=D&source=editors&ust=1779048536785167&usg=AOvVaw3XXtk5amQrPyLVIlYCek8b) |
| 327 | [▪ LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses (Lin et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.04755&sa=D&source=editors&ust=1779048536785200&usg=AOvVaw1pkcTik4PNjAr4eVuZZ41N) |
| 328 | [▪ Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents (Shafran, Schuster, and Shmatikov, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05870&sa=D&source=editors&ust=1779048536785234&usg=AOvVaw10pEvbuSrNWcV2s4RLhq9s) |
| 329 | [▪ Security Attacks on LLM-based Code Completion Tools (Cheng et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.11006&sa=D&source=editors&ust=1779048536785268&usg=AOvVaw2Yy-fp8eKZsgWDH1_Q-Oqx) |
| 330 | [▪ Adversarial Attacks to Multi-Modal Models (Dou et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.06793&sa=D&source=editors&ust=1779048536785300&usg=AOvVaw2XQ6beeMnm0OKGr07xlEM6) |
| 331 | [▪ HSF: Defending against Jailbreak Attacks with Hidden State Filtering (Qian et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03788&sa=D&source=editors&ust=1779048536785333&usg=AOvVaw1MMRTrKEpuv0ilE9NoWSLb) |
| 332 | [▪ Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models (Kalavasis et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05660&sa=D&source=editors&ust=1779048536785368&usg=AOvVaw3ZWtyDt83ead5Li2MvJpZD) |
| 333 | [▪ Recent Advances in Attack and Defense Approaches of Large Language Models (Cui et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03274&sa=D&source=editors&ust=1779048536785402&usg=AOvVaw0kQLYtrweA0Qe_yLJLhqwD) |
| 334 | [▪ SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner (Wang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05498&sa=D&source=editors&ust=1779048536785436&usg=AOvVaw1sCk5P63bkstAHBnC7_Fde) |
| 335 | [▪ LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet (Li et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.15221&sa=D&source=editors&ust=1779048536785469&usg=AOvVaw2MKYJpF0PPHigCTK9vBgnn) |
| 336 | [▪ Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models (Ma et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.02928&sa=D&source=editors&ust=1779048536785502&usg=AOvVaw1NW_FMvxovphDNEo3nnIjM) |
| 337 | [▪ Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models (An et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00598&sa=D&source=editors&ust=1779048536785536&usg=AOvVaw0avpGVZeoYj-xRewaZLYAE) |
| 338 | [▪ The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models (Wu et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.17915&sa=D&source=editors&ust=1779048536785569&usg=AOvVaw3wQWn1_9plcDWs-ENxlrDu) |
| 339 | [▪ Detecting AI Flaws: Target-Driven Attacks on Internal Faults in Language Models (Du et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.14853&sa=D&source=editors&ust=1779048536785602&usg=AOvVaw32o4EdNtpGed-aT1whahkZ) |
| 340 | [▪ LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet (Li et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.15221&sa=D&source=editors&ust=1779048536785634&usg=AOvVaw0rFRmbFpW6tTQRY3uzI--c) |
| 341 | [▪ Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything (Zou et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.02534&sa=D&source=editors&ust=1779048536785667&usg=AOvVaw0cDzem6t0c4LEGzVUE5Hik) |
| 342 | [▪ A StrongREJECT for Empty Jailbreaks (Souly et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.10260&sa=D&source=editors&ust=1779048536785698&usg=AOvVaw1gF2I1dRxu6yq3QsRyrcR5) |
| 343 | [▪ RT-Attack: Jailbreaking Text-to-Image Models via Random Token (Gao et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.13896&sa=D&source=editors&ust=1779048536785731&usg=AOvVaw3FQ9MnD2KN3P6uC3h7UmJP) |
| 344 | [▪ CAMH: Advancing Model Hijacking Attack in Machine Learning (He et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.13741%23&sa=D&source=editors&ust=1779048536785763&usg=AOvVaw2z9oWWnMMwFqyRw5ct6QVL) |
| 345 | [▪ RT-Attack: Jailbreaking Text-to-Image Models via Random Token (Gao et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.13896&sa=D&source=editors&ust=1779048536785796&usg=AOvVaw2FGOFk3XgknFAxxMgmaDKn) |
| 346 | [▪ BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger (Chen et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09093%23&sa=D&source=editors&ust=1779048536785831&usg=AOvVaw3VVjGrrmNdYzy8CX5lk1dQ) |
| 347 | [▪ Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks (Zhao et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08924&sa=D&source=editors&ust=1779048536785867&usg=AOvVaw1yMoLB2xF0LveZqc8azmyq) |
| 348 | [▪ A Survey of Trojan Attacks and Defenses to Deep Neural Networks (Jin et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08920&sa=D&source=editors&ust=1779048536785914&usg=AOvVaw3IJXnorI6NCkyr8BI6Y-XU) |
| 349 | [▪ MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defensesfor Vision Language Models (Weng et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08464&sa=D&source=editors&ust=1779048536785955&usg=AOvVaw2UgypDu0kx4vO2dxyGQbwl) |
| 350 | [▪ Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment (Wang and Shu, Aug 2023)](https://www.google.com/url?q=https://arxiv.org/abs/2311.09433&sa=D&source=editors&ust=1779048536785995&usg=AOvVaw1HvTl1lgQKnX49PpHQf3i2) |
| 351 | [▪ Resilience in Online Federated Learning: Mitigating Model-Poisoning Attacks via Partial Sharing (Lari et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.13108&sa=D&source=editors&ust=1779048536786040&usg=AOvVaw34cm8M8haMK4f_Bp7M0Ant) |
| 352 | [▪ EnJa: Ensemble Jailbreak on Large Language Models (Zhang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.03603&sa=D&source=editors&ust=1779048536786075&usg=AOvVaw1WRj03MIhtYjmZ0PWs3RNV) |
| 353 | [▪ Can Reinforcement Learning Unlock the Hidden Dangers in Aligned Large Language Models? (Bahrami, Vishwamitra, and Najafirad, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02651&sa=D&source=editors&ust=1779048536786119&usg=AOvVaw0Yjx6nF2bO5MRaSIg7ABGd) |
| 354 | [▪ Jailbreaking Text-to-Image Models with LLM-Based Agents (Dong et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.00523&sa=D&source=editors&ust=1779048536786155&usg=AOvVaw1hutiSLY4GCTu_m4TPh2G1) |
| 355 | [▪ Can LLMs be Fooled? Investigating Vulnerabilities in LLMs (Abdali et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20529&sa=D&source=editors&ust=1779048536786190&usg=AOvVaw1-Y15PWrpDzj8qtl5LYv0b) |
| 356 | [▪ Figure it Out: Analyzing-based Jailbreak Attack on Large Language Models (Lu et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.16205&sa=D&source=editors&ust=1779048536786224&usg=AOvVaw0olUCK7oBZ6ImwbLqf67C-) |
| 357 | [▪ Vera Verto: Multimodal Hijacking Attack (Zhang et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.00129%23&sa=D&source=editors&ust=1779048536786260&usg=AOvVaw3qtvExSZoBJuVc8n2XI-dF) |

|     |
| --- |
| Jailbreak, Cost Harvesting, or Erode ML Model Integrity |

**>**

**<**

‍

#### Proxy AI ML Model (Simulations)

Covers:

- MITRE ATLAS ML Attack Staging

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Proxy AI ML Model (Simulations)

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior (Flemings and Annavaram, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12147&sa=D&source=editors&ust=1779048536739260&usg=AOvVaw3nfEMyyyy8nKFXarFcMp7z) |
| 3 | [▪ Searching for Privacy Risks in LLM Agents via Simulation (Zhang and Yang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.10880&sa=D&source=editors&ust=1779048536739455&usg=AOvVaw1ZyZxFjHDrM3YKtaJgAzbj) |
| 4 | [▪ HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion (Ferrel, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07472&sa=D&source=editors&ust=1779048536739624&usg=AOvVaw2vSTnmfxg9ReqmAe7-bcsw) |
| 5 | [▪ Threat-Oriented Digital Twinning for Security Evaluation of Autonomous Platforms (Neubert, Kandel, and Pek\\"oz, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25757&sa=D&source=editors&ust=1779048536739766&usg=AOvVaw2jQzC2-SHTbKyQRG80iJHl) |
| 6 | [▪ Text-Based Personas for Simulating User Privacy Decisions (Fawaz et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19791&sa=D&source=editors&ust=1779048536739899&usg=AOvVaw3CeudKyMsWAleePDxC29yb) |
| 7 | [▪ Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents (Luk\\'a\\v{s} et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10041&sa=D&source=editors&ust=1779048536740051&usg=AOvVaw2y29BCJEXJAx8fdhtZ8lOS) |
| 8 | [▪ How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors? (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18464&sa=D&source=editors&ust=1779048536740229&usg=AOvVaw0kJgR2aQQWroXkfPMQUn4p) |
| 9 | [▪ Towards Production-Worthy Simulation for Autonomous Cyber Operations (Tholl et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.19278&sa=D&source=editors&ust=1779048536740389&usg=AOvVaw0hxd-xUuy_GjAABFGNU_U6) |
| 10 | [▪ CyberExplorer: Benchmarking LLM Offensive Security Capabilities in a Real-World Attacking Simulation Environment (Rani et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08023&sa=D&source=editors&ust=1779048536740530&usg=AOvVaw3EpXM5KIzb91JTa7JtMxmS) |
| 11 | [▪ Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent (Xuan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06325&sa=D&source=editors&ust=1779048536740659&usg=AOvVaw1--dga4LEw9ecZ_tNCEWN8) |
| 12 | [▪ Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks (Isbarov and Kantarcioglu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05066&sa=D&source=editors&ust=1779048536740772&usg=AOvVaw1a4j0MPCh1Ol1Okv6N4RIE) |
| 13 | [▪ VirtualCrime: Evaluating Criminal Potential of Large Language Models via Sandbox Simulation (Tang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13981&sa=D&source=editors&ust=1779048536740896&usg=AOvVaw2CSQurBwLgRIZoWglf9Ej_) |
| 14 | [▪ MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers (Ahmadi, Sharif, and Banad, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.08999&sa=D&source=editors&ust=1779048536741037&usg=AOvVaw0xI1NK-fdwZ1aI6ul7P907) |
| 15 | [▪ The Imitation Game: Using Large Language Models as Chatbots to Combat Chat-Based Cybercrimes (Yao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21371&sa=D&source=editors&ust=1779048536741158&usg=AOvVaw0OkEkxDKP6ETzlZnc7WlM0) |
| 16 | [▪ Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation (Yu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.07745&sa=D&source=editors&ust=1779048536741317&usg=AOvVaw2C7gSJVMzaOyEjlKUDDiR6) |
| 17 | [▪ ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls (Badhe, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06457&sa=D&source=editors&ust=1779048536741471&usg=AOvVaw0YmbRUeJDJaq9SpmWyZJvV) |
| 18 | [▪ HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities (Ren et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12200&sa=D&source=editors&ust=1779048536741606&usg=AOvVaw36H-jr-4LkiyyWRMcA1tAa) |
| 19 | [▪ Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems (Peng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10937&sa=D&source=editors&ust=1779048536741724&usg=AOvVaw03l5rqHRYghp9YQQemC_cK) |
| 20 | [▪ SCART: Simulation of Cyber Attacks for Real-Time (Rahimi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2304.03657&sa=D&source=editors&ust=1779048536741824&usg=AOvVaw02Q_y1X_TyneUa33jEK93R) |
| 21 | [▪ LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits (Badhe, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03405&sa=D&source=editors&ust=1779048536741921&usg=AOvVaw0CFYAXGzdHLq-vwElsGx2Z) |
| 22 | [▪ Secret Collusion among AI Agents: Multi-Agent Deception via Steganography (Motwani et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.07510&sa=D&source=editors&ust=1779048536742053&usg=AOvVaw3_uzKDhG6SaS9uic1JgswB) |

|     |
| --- |
| Proxy AI ML Model (Simulations) |

**>**

**<**

‍

#### Verify Attack (Efficacy)

Covers:

- MITRE ATLAS ML Attack Staging

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Verify Attack (Efficacy)

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis (Topol, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12869&sa=D&source=editors&ust=1779048535895681&usg=AOvVaw2eVs4jJuH8CSuCJdiJzxvt) |
| 3 | [▪ CTFusion: A CTF-based Benchmark for LLM Agent Evaluation (Lee, Bae, and Yun, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11504&sa=D&source=editors&ust=1779048535896019&usg=AOvVaw28JNGLxt5PHynLC95j7p_q) |
| 4 | [▪ Behavioral Integrity Verification for AI Agent Skills (Wu, Li, and Liu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11770&sa=D&source=editors&ust=1779048535896220&usg=AOvVaw12WGgxmgOC6SMoHay1RG3Z) |
| 5 | [▪ SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response (Barbieri et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05501&sa=D&source=editors&ust=1779048535896401&usg=AOvVaw22Zk7MNcB084KPrtL-D1nU) |
| 6 | [▪ Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges (Al-Kaswan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19354&sa=D&source=editors&ust=1779048535896580&usg=AOvVaw1O5RIN3Cggj2n9POBfZWqH) |
| 7 | [▪ Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection (Kulkarni, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.28129&sa=D&source=editors&ust=1779048535896752&usg=AOvVaw3iCqUypq5zV1JxP9r1uaVM) |
| 8 | [▪ Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode (Ji et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04978&sa=D&source=editors&ust=1779048535896910&usg=AOvVaw16JtMdBnR_vN2wAlnkFJED) |
| 9 | [▪ The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive (Bogdan and Valois-Franklin, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25634&sa=D&source=editors&ust=1779048535897049&usg=AOvVaw2kS1Vtjfnfb52Xw33QdfbF) |
| 10 | [▪ MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25152&sa=D&source=editors&ust=1779048535897222&usg=AOvVaw3XgOL_57MdwVIMtaLdr3kA) |
| 11 | [▪ Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges (Al-Kaswan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19354&sa=D&source=editors&ust=1779048535897405&usg=AOvVaw0pbd2WD0eg6pGOytp8qc6w) |
| 12 | [▪ Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery (Agarwal, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19049&sa=D&source=editors&ust=1779048535897547&usg=AOvVaw3Xgq_V0NcYSZA4HZX0LISh) |
| 13 | [▪ RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code (Pellew and Raza, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13764&sa=D&source=editors&ust=1779048535897699&usg=AOvVaw13M0yWHTeEFAcSkzGxwtxx) |
| 14 | [▪ Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis (Gajjar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10800&sa=D&source=editors&ust=1779048535897863&usg=AOvVaw0YkhB41UOw3OzBXkk0ojIH) |
| 15 | [▪ PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy (Duy et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06618&sa=D&source=editors&ust=1779048535898064&usg=AOvVaw1ynG3AIMAaf-UQwjdQe-k1) |
| 16 | [▪ Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery (Shafiuzzaman et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06506&sa=D&source=editors&ust=1779048535898333&usg=AOvVaw3Lqz0vLqCXOn8Uy7_ErQlC) |
| 17 | [▪ Dynamic Free-Rider Detection in Federated Learning via Simulated Attack Patterns (Nakamura, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04611&sa=D&source=editors&ust=1779048535898563&usg=AOvVaw0vkMjpE5WddHoI4TIpHJua) |
| 18 | [▪ OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection (Flynt, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22499&sa=D&source=editors&ust=1779048535898783&usg=AOvVaw20mg8E2TV6sY7Kt4eS_5tf) |
| 19 | [▪ SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration (Guo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21019&sa=D&source=editors&ust=1779048535899071&usg=AOvVaw1KsHzOEBf3m1Z7pQT6AFkd) |
| 20 | [▪ Towards Verifiable AI with Lightweight Cryptographic Proofs of Inference (Anchuri et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19025&sa=D&source=editors&ust=1779048535899345&usg=AOvVaw0JV4FIuJrjbHyOye0CRYSS) |
| 21 | [▪ When Scanners Lie: Evaluator Instability in LLM Red-Teaming (Erez et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14633&sa=D&source=editors&ust=1779048535899581&usg=AOvVaw2JCOvPK4OztP7-T16OhfiY) |
| 22 | [▪ Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis (Kharma et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.01853&sa=D&source=editors&ust=1779048535899733&usg=AOvVaw3UQNFfkZ5HYn3NIAq3XML6) |
| 23 | [▪ Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis (Jajodia et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.06422&sa=D&source=editors&ust=1779048535899885&usg=AOvVaw2rup4lcBYJ12YOaW8T29Rl) |
| 24 | [▪ Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection (Wei et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04469&sa=D&source=editors&ust=1779048535900044&usg=AOvVaw15sAPNqaBygGFeQImWDZtZ) |
| 25 | [▪ AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding (Ma et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.03170&sa=D&source=editors&ust=1779048535900212&usg=AOvVaw3t9ppugGR3F73-fQOl3tYO) |
| 26 | [▪ ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense (Lau et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02297&sa=D&source=editors&ust=1779048535900371&usg=AOvVaw0I0SdPlMF3sexsXR48X9jN) |
| 27 | [▪ Can LLMs Hack Enterprise Networks? -- Replicated Computational Results (RCR) Report (Happe and Cito, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01789&sa=D&source=editors&ust=1779048535900503&usg=AOvVaw0NfC9hFOM1NSDpXAKmCivP) |
| 28 | [▪ DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern (Pang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01574&sa=D&source=editors&ust=1779048535900627&usg=AOvVaw2vI6fA8Batg8MFEKrCW4LB) |
| 29 | [▪ vEcho: A Paradigm Shift from Vulnerability Verification to Proactive Discovery with Large Language Models (Jiang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01154&sa=D&source=editors&ust=1779048535900758&usg=AOvVaw3aLeUw_BnqPJb5qfuUDMgU) |
| 30 | [▪ IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation (Guo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22700&sa=D&source=editors&ust=1779048535900858&usg=AOvVaw0MN3i6jIuxfq9khkTQPoBj) |
| 31 | [▪ Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study (Khatiwala, Addai, and Xu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20202&sa=D&source=editors&ust=1779048535900964&usg=AOvVaw0rGjvxk40pcxOo5syV9LFY) |
| 32 | [▪ Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments (Mukherjee, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19450&sa=D&source=editors&ust=1779048535901081&usg=AOvVaw2J3y6k88PGU_zkHabhGUAq) |
| 33 | [▪ Mind the Gap: Evaluating LLMs for High-Level Malicious Package Detection vs. Fine-Grained Indicator Identification (Ryan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16304&sa=D&source=editors&ust=1779048535901222&usg=AOvVaw0Kw7RQl241fosNbn1rGxQf) |
| 34 | [▪ Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection (Corll, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11247&sa=D&source=editors&ust=1779048535901355&usg=AOvVaw2C4gUA4NkQs9VbSBRr6YBe) |
| 35 | [▪ Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models (Wei et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20677&sa=D&source=editors&ust=1779048535901471&usg=AOvVaw0vFO7dZreZ0OCR-yR4iyue) |
| 36 | [▪ VideoSTF: Stress-Testing Output Repetition in Video Large Language Models (Cao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10639&sa=D&source=editors&ust=1779048535901571&usg=AOvVaw3_s41RnU3QkYaqbIT6A9T5) |
| 37 | [▪ Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation (Qi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09319&sa=D&source=editors&ust=1779048535901676&usg=AOvVaw3Z20d3NutCZUH3nifboSuF) |
| 38 | [▪ Capability-Based Scaling Trends for LLM-Based Red-Teaming (Panfilov et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.20162&sa=D&source=editors&ust=1779048535901778&usg=AOvVaw3RZNXlF6E2WAipl6QFZGh-) |
| 39 | [▪ Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4 (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08384&sa=D&source=editors&ust=1779048535901880&usg=AOvVaw2TNfsBjv50_MY6lBACNAVO) |
| 40 | [▪ Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs (Chagna and Goldschmidt, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06777&sa=D&source=editors&ust=1779048535902002&usg=AOvVaw2iQXn99E46n4OTxcv8Ho1L) |
| 41 | [▪ On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection (Hannan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27675&sa=D&source=editors&ust=1779048535902130&usg=AOvVaw2LaA2hm52G2Wg1D6DpJSvI) |
| 42 | [▪ SVIP: Towards Verifiable Inference of Open-source Large Language Models (Sun et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22307&sa=D&source=editors&ust=1779048535902227&usg=AOvVaw0rzKbDrTLnRWJa0wL07HsP) |
| 43 | [▪ GradingAttack: Attacking Large Language Models Towards Short Answer Grading Ability (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00979&sa=D&source=editors&ust=1779048535902336&usg=AOvVaw1Ho337bJxkYaC0UtMFwgWi) |
| 44 | [▪ Evaluating Large Language Models for Security Bug Report Prediction (Soltaniani, Razzaq, and Ghafari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22921&sa=D&source=editors&ust=1779048535902492&usg=AOvVaw30cUilLdUWrMr2tgWV4iZx) |
| 45 | [▪ Lightweight LLMs for Network Attack Detection in IoT Networks (Sudasinghe, Liyanage, and Pussewalage, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15269&sa=D&source=editors&ust=1779048535902658&usg=AOvVaw35iuudGzrvLkT3XZVcNDrv) |
| 46 | [▪ Holmes: An Evidence-Grounded LLM Agent for Auditable DDoS Investigation in Cloud Networks (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14601&sa=D&source=editors&ust=1779048535902783&usg=AOvVaw0kNYlDHMLVMc53_7Nt2FAC) |
| 47 | [▪ Proactively Detecting Threats: A Novel Approach Using LLMs (Chawla and Prasad, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09029&sa=D&source=editors&ust=1779048535902915&usg=AOvVaw1-3fFr3RmDjhPiqt7aeSXo) |
| 48 | [▪ LLMs as verification oracles for Solidity (Bartoletti, Lipparini, and Pompianu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.19153&sa=D&source=editors&ust=1779048535903055&usg=AOvVaw3lJYttcC5R4NQNIGP8WFO7) |
| 49 | [▪ Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection (Pushkar et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22306&sa=D&source=editors&ust=1779048535903179&usg=AOvVaw0n8KMfQfyM6pnciDRtC3Ai) |
| 50 | [▪ Evaluating Large Language Models for Line-Level Vulnerability Localization (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2404.00287&sa=D&source=editors&ust=1779048535903306&usg=AOvVaw1MITpq9-6RTCBnjJZhfa9y) |
| 51 | [▪ VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces (Grigor et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15892&sa=D&source=editors&ust=1779048535903477&usg=AOvVaw2iDEpIrQ0Jq1WIaQZuGTuJ) |
| 52 | [▪ Exact Verification of Graph Neural Networks with Incremental Constraint Solving (Liu, Lu, and Kwiatkowska, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09320&sa=D&source=editors&ust=1779048535903605&usg=AOvVaw1yQosIpRoq50KPj3ELjYIO) |
| 53 | [▪ Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification (Zhou et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11087&sa=D&source=editors&ust=1779048535903739&usg=AOvVaw28AGNg7jutUWBIpjeIDpsx) |
| 54 | [▪ Verifying LLM Inference to Detect Model Weight Exfiltration (Rinberg et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02620&sa=D&source=editors&ust=1779048535903856&usg=AOvVaw1Ph5qFM9ETRy7S2hz0MF5B) |
| 55 | [▪ From Description to Score: Can LLMs Quantify Vulnerabilities? (Jafarikhah et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06781&sa=D&source=editors&ust=1779048535903970&usg=AOvVaw3r7natz2UlauYJPpnIUb0H) |
| 56 | [▪ Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation (Fei et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01255&sa=D&source=editors&ust=1779048535904088&usg=AOvVaw15nXF_0CFbdRCNhk7Ahzis) |
| 57 | [▪ Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22044&sa=D&source=editors&ust=1779048535904206&usg=AOvVaw2umP63kffOsqyuG1UkqkeN) |
| 58 | [▪ Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning (Gravereaux and Islam, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19654&sa=D&source=editors&ust=1779048535904339&usg=AOvVaw3mj_j_oSJJtr6EXCVdooNh) |
| 59 | [▪ AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19536&sa=D&source=editors&ust=1779048535904508&usg=AOvVaw0PDuphFAoEL989wmskKqDy) |
| 60 | [▪ LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models (Shahid, Ahmed, and Ranjan, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18966&sa=D&source=editors&ust=1779048535904625&usg=AOvVaw0sOSYe1_dX31DDg3zVByAz) |
| 61 | [▪ Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees (Nakano et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.07939&sa=D&source=editors&ust=1779048535904731&usg=AOvVaw1c0GsSXb4LMXLOvd-XV6gs) |
| 62 | [▪ CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning (Deason et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.20166&sa=D&source=editors&ust=1779048535904846&usg=AOvVaw1VDhtUeaj2qatKq00hJ3ud) |
| 63 | [▪ Verifying LLM Inference to Prevent Model Weight Exfiltration (Rinberg et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02620&sa=D&source=editors&ust=1779048535904946&usg=AOvVaw1u4YkmhLVu2bKOF3oyETfr) |
| 64 | [▪ AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence (Alam et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01144&sa=D&source=editors&ust=1779048535905033&usg=AOvVaw0MuzMR0AaTLSNEZ9J41xAk) |
| 65 | [▪ Scalable GPU-Based Integrity Verification for Large Machine Learning Models (Spoczynski and Melara, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23938&sa=D&source=editors&ust=1779048535905116&usg=AOvVaw17GGyk8jM8sRLHGexZmdf-) |
| 66 | [▪ Floating-Point Neural Network Verification at the Software Level (Manino et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23389&sa=D&source=editors&ust=1779048535905196&usg=AOvVaw1isyvT05b0iKMjByxXoMJB) |
| 67 | [▪ A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs (H{\\o}st, Lison, and Moonen, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.18439&sa=D&source=editors&ust=1779048535905286&usg=AOvVaw1RekU-s-QqaNTI8dxRvTOO) |
| 68 | [▪ Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts (Downey-Webb, Jogunola, and Ajao, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15973&sa=D&source=editors&ust=1779048535905372&usg=AOvVaw2xxGO30RkNy2eQUhjj4yfl) |
| 69 | [▪ LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet? (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14700&sa=D&source=editors&ust=1779048535905561&usg=AOvVaw0ZKl0USybR9zz4gknMTCTe) |
| 70 | [▪ CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.02548&sa=D&source=editors&ust=1779048535905676&usg=AOvVaw1X4mHgmFW144g1gR3phF8U) |
| 71 | [▪ VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference (Wang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24257&sa=D&source=editors&ust=1779048535905762&usg=AOvVaw0pVrQlG98BH2TuiKPLnCDE) |
| 72 | [▪ Automated Vulnerability Validation and Verification: A Large Language Model Approach (Lotfi, Katsis, and Bertino, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24037&sa=D&source=editors&ust=1779048535905847&usg=AOvVaw1ZleBde9PN6iyO72p775Fp) |
| 73 | [▪ DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16329&sa=D&source=editors&ust=1779048535905931&usg=AOvVaw0olkbJO7qmbOdElWjzOste) |
| 74 | [▪ LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification (Tang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11310&sa=D&source=editors&ust=1779048535906018&usg=AOvVaw0pNw_vJge2Q1G318Awddtg) |

|     |
| --- |
| Verify Attack (Efficacy) |

**>**

**<**

‍

#### Insecure Output Handling

Covers:

- OWASP LLM 02: Insecure Output Handling

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Insecure Output Handling

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ "Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs (Lyu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06759&sa=D&source=editors&ust=1779048535893605&usg=AOvVaw2ucgeTtOyJts30jJrhItwm) |
| 3 | [▪ Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs (Yuan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20211&sa=D&source=editors&ust=1779048535893764&usg=AOvVaw0myJzsDHmSd-XdrL-EyN0N) |
| 4 | [▪ Surgical Repair of Insecure Code Generation in LLMs (Sandoval, Dolan-Gavitt, and Garg, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16697&sa=D&source=editors&ust=1779048535893843&usg=AOvVaw28hoSq2dK8YWuTK67nkoRu) |
| 5 | [▪ From IOCs to Regex: Automating CTI Operationalization for SOC with LLMs (University et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12228&sa=D&source=editors&ust=1779048535893917&usg=AOvVaw0jeb25jV8P39gisnryTEs-) |
| 6 | [▪ How Secure is Code Generated by ChatGPT? (Khoury et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2304.09655&sa=D&source=editors&ust=1779048535893991&usg=AOvVaw3SiYgPjCxNqAqzlxGol0zC) |
| 7 | [▪ Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents (Chinaei, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04035&sa=D&source=editors&ust=1779048535894060&usg=AOvVaw3RX0MFYyGUy1kcJr3FPOCV) |
| 8 | [▪ Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing (Langiu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28972&sa=D&source=editors&ust=1779048535894134&usg=AOvVaw3z81dCsz-WdL2aF8AL3bU0) |
| 9 | [▪ How Vulnerable Are Edge LLMs? (Ding et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23822&sa=D&source=editors&ust=1779048535894206&usg=AOvVaw3dSLWDabMoP_moYTPcGvZB) |
| 10 | [▪ FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09046&sa=D&source=editors&ust=1779048535894280&usg=AOvVaw10Ia6q2Tm5EGSkDwhUpWNo) |
| 11 | [▪ Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise (Saini, Jiang, and Liu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11088&sa=D&source=editors&ust=1779048535894366&usg=AOvVaw3sUDQ1guEwXlZXdxXxX3zQ) |
| 12 | [▪ CryptoGen: Secure Transformer Generation with Encrypted KV-Cache Reuse (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08798&sa=D&source=editors&ust=1779048535894435&usg=AOvVaw1kwQv3c12xQXz_B323Td4K) |
| 13 | [▪ "Tab, Tab, Bug'': Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs (Lyu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06759&sa=D&source=editors&ust=1779048535894518&usg=AOvVaw0fanZFNLpRxbypc1QydOO6) |
| 14 | [▪ Supporting Students in Navigating LLM-Generated Insecure Code (Park et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20878&sa=D&source=editors&ust=1779048535894597&usg=AOvVaw0wtCgoED0BSBmnNADJjvGy) |
| 15 | [▪ Think Fast: Real-Time IoT Intrusion Reasoning Using IDS and LLMs at the Edge Gateway (Jamshidi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18230&sa=D&source=editors&ust=1779048535894677&usg=AOvVaw1vcMmyLZmmu9S8ANYhTibo) |
| 16 | [▪ GenSIaC: Toward Security-Aware Infrastructure-as-Code Generation with Large Language Models (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12385&sa=D&source=editors&ust=1779048535894766&usg=AOvVaw05k44BLt6r1tLTX6W2j0NS) |
| 17 | [▪ DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agents (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.23804&sa=D&source=editors&ust=1779048535894847&usg=AOvVaw0hoKUXFhiiCNMVmKExh7cG) |
| 18 | [▪ TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code (Sternfeld, Kucharavy, and Dolamic, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11151&sa=D&source=editors&ust=1779048535894929&usg=AOvVaw1vSfZ8Yxzv1xh8gDw-kuzU) |
| 19 | [▪ Semantic Encryption: Secure and Effective Interaction with Cloud-based Large Language Models via Semantic Transformation (Chen et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01638&sa=D&source=editors&ust=1779048535895018&usg=AOvVaw1k3QbTxL_6D-LhXknOIwjp) |
| 20 | [▪ eX-NIDS: A Framework for Explainable Network Intrusion Detection Leveraging Large Language Models (Houssel et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16241&sa=D&source=editors&ust=1779048535895118&usg=AOvVaw2NeGvrqjbde3ACb3Tqh30D) |
| 21 | [▪ Accelerating Automatic Program Repair with Dual Retrieval-Augmented Fine-Tuning and Patch Generation on Large Language Models (Guo et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10103&sa=D&source=editors&ust=1779048535895225&usg=AOvVaw3_iELzZcvLULy6isCQjMes) |
| 22 | [▪ ETrace:Event-Driven Vulnerability Detection in Smart Contracts via LLM-Based Trace Analysis (Peng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15790&sa=D&source=editors&ust=1779048535895311&usg=AOvVaw0LkoYEDoG16qOaaHkIYtdr) |
| 23 | [▪ DESIGN: Encrypted GNN Inference via Server-Side Input Graph Pruning (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05649&sa=D&source=editors&ust=1779048535895390&usg=AOvVaw0-gTEJlWgviULHzup42bzA) |
| 24 | [▪ Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning (Suo et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21680&sa=D&source=editors&ust=1779048535895471&usg=AOvVaw1WsFq2MqBg9jZzukeWQvaQ) |
| 25 | [▪ Bayes-Nash Generative Privacy Against Membership Inference Attacks (Zhang et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07414&sa=D&source=editors&ust=1779048535895578&usg=AOvVaw1Oj8YRc3vW5tvogJop35xl) |
| 26 | [▪ A hierarchical approach for assessing the vulnerability of tree-based classification models to membership inference attack (Preen and Smith, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.09396&sa=D&source=editors&ust=1779048535895669&usg=AOvVaw3xG_6ZdzAXPkDBHobeX0QG) |
| 27 | [▪ The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems (Song et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.20002&sa=D&source=editors&ust=1779048535895804&usg=AOvVaw23hbYU_yNoN23A1jvlPZOj) |
| 28 | [▪ LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states (Ibanez-Lissen et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.19876&sa=D&source=editors&ust=1779048535895906&usg=AOvVaw2WowN9Sr4a3vFgChi0s8xF) |
| 29 | [▪ Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models (Zhang, Saileshwar, and Lie, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15431&sa=D&source=editors&ust=1779048535896000&usg=AOvVaw1Z7V7SbrHu18ecNlCmfm2R) |
| 30 | [▪ Privacy-Preserving Low-Rank Adaptation against Membership Inference Attacks for Latent Diffusion Models (Luo et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.11989&sa=D&source=editors&ust=1779048535896099&usg=AOvVaw2QVC_T0xTN8_YShOpuiJnc) |
| 31 | [▪ Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration (Fu et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2311.06062&sa=D&source=editors&ust=1779048535896182&usg=AOvVaw3PZfQcsgVLHl9yXWCGeK0o) |
| 32 | [▪ InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks (Zheng et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.18191&sa=D&source=editors&ust=1779048535896268&usg=AOvVaw1dejDFyakCDWDdt2sRCiDR) |
| 33 | [▪ Towards Black-Box Membership Inference Attack for Diffusion Models (Li et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20771&sa=D&source=editors&ust=1779048535896362&usg=AOvVaw1hX_flPn3egdUEPxuBmmVd) |
| 34 | [▪ Protection against Source Inference Attacks in Federated Learning using Unary Encoding and Shuffling (Athanasiou, Jung, and Palmidessi, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.06458&sa=D&source=editors&ust=1779048535896457&usg=AOvVaw0r60-tLRbCdBAqFQIE6EPU) |
| 35 | [▪ OSLO: One-Shot Label-Only Membership Inference Attacks (Peng et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.16978&sa=D&source=editors&ust=1779048535896554&usg=AOvVaw3EYsFl1ZklS5_sMsqQ6gFT) |
| 36 | [▪ Detecting Training Data of Large Language Models via Expectation Maximization (Kim et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07582&sa=D&source=editors&ust=1779048535896640&usg=AOvVaw2fuyJuyK9KkwTob4WV_XLH) |
| 37 | [▪ Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models (Suriyakumar et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.08074&sa=D&source=editors&ust=1779048535896724&usg=AOvVaw1QFzLWAY6W3zd9UaXqpYzQ) |
| 38 | [▪ Black-box Membership Inference Attacks against Fine-tuned Diffusion Models (Pang and Wang, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.08207&sa=D&source=editors&ust=1779048535896806&usg=AOvVaw1PE0jbg8-clnxWaFXgufIj) |
| 39 | [▪ Is Difficulty Calibration All We Need? Towards More Practical Membership Inference Attacks (He et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00426&sa=D&source=editors&ust=1779048535896886&usg=AOvVaw3TadT0dGqyr4GnUjHmpOWB) |
| 40 | [▪ Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies (Wang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.07728&sa=D&source=editors&ust=1779048535896977&usg=AOvVaw0KT2_dbJmOG-0eZ_QMBtf0) |
| 41 | [▪ Membership Inference Attack Against Masked Image Modeling (Li et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.06825&sa=D&source=editors&ust=1779048535897051&usg=AOvVaw0X3mBadZZuZh7WkeRsEjzz) |
| 42 | [▪ Pathway to Secure and Trustworthy 6G for LLMs: Attacks, Defense, and Opportunities (Khowaja et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/html/2408.00722v1&sa=D&source=editors&ust=1779048535897141&usg=AOvVaw2D2OXHL0ecOL1tbyMHjpG6) |
| 43 | [▪ Synthetic Image Learning: Preserving Performance and Preventing Membership Inference Attacks (Lomurno and Matteucci, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.15526&sa=D&source=editors&ust=1779048535897226&usg=AOvVaw2aPwltwZotd_gSB8nE6lEQ) |
| 44 | [▪ Thermometer: Towards Universal Calibration for Large Language Models (Shen et al, Jun 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.08819&sa=D&source=editors&ust=1779048535897297&usg=AOvVaw1aY8uQU0PDRuANINhUFGWV) |

|     |
| --- |
| Insecure Output Handling |

**>**

**<**

‍

#### Sensitive Information Disclosure

Covers:

- OWASP LLM 06: Sensitive Information Disclosure

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Sensitive Information Disclosure

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Do Skill Descriptions Tell the Truth? Detecting Undisclosed Security Behaviors in Code-Backed LLM Skills (He et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12875&sa=D&source=editors&ust=1779048535943164&usg=AOvVaw0C4EfGGPRND5t1xc-duua0) |
| 3 | [▪ Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG) (Mori et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06719&sa=D&source=editors&ust=1779048535943358&usg=AOvVaw30DIUaYlwZoMAv2Bf65KGZ) |
| 4 | [▪ Reconstruction of Personally Identifiable Information from Supervised Finetuned Models (Furukawa and Oprea, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12264&sa=D&source=editors&ust=1779048535943457&usg=AOvVaw2QfcDyR0-d70Yu-MA3CzXh) |
| 5 | [▪ Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing (Holtzman and West, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10794&sa=D&source=editors&ust=1779048535943548&usg=AOvVaw1e3-PjshEwpRapq8H62OSV) |
| 6 | [▪ Deep Learning under Fractional-Order Differential Privacy (Partohaghighi and Marcia, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09890&sa=D&source=editors&ust=1779048535943644&usg=AOvVaw3RTLzO6CCm2sSg4SlIA9w6) |
| 7 | [▪ Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06232&sa=D&source=editors&ust=1779048535943746&usg=AOvVaw0pMfrqHQuwobG1WB0LI8cY) |
| 8 | [▪ How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05340&sa=D&source=editors&ust=1779048535943833&usg=AOvVaw3CxuSKGzgD7_X21hyRw3xJ) |
| 9 | [▪ From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists (Cecconello et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04724&sa=D&source=editors&ust=1779048535943959&usg=AOvVaw2EESA7M8c5gfLpcg9jVzx0) |
| 10 | [▪ Graph Reconstruction from Differentially Private GNN Explanations (Sahoo, Shivottam, and Mishra, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03388&sa=D&source=editors&ust=1779048535944047&usg=AOvVaw3fWMG5MqJSfWEr9OMwIN8j) |
| 11 | [▪ Evaluating Retrieval-Augmented Generation for Explainable Malware Analysis (Ng and Fard, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03140&sa=D&source=editors&ust=1779048535944129&usg=AOvVaw1af2s5SmRGvYu9c5oIO96j) |
| 12 | [▪ PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization (Liu, Zha, and Chen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03129&sa=D&source=editors&ust=1779048535944211&usg=AOvVaw1A1QaHwQb_lqS_CGGEPfGR) |
| 13 | [▪ Metric-Normalized Posterior Leakage (mPL): Attacker-Aligned Privacy for Joint Consumption (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01137&sa=D&source=editors&ust=1779048535944297&usg=AOvVaw2T5mKh61JGYPrORW1mfvfC) |
| 14 | [▪ Revisiting Privacy Leakage in Machine Unlearning: Membership Inference Beyond the Forgotten Set (Fu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01129&sa=D&source=editors&ust=1779048535944379&usg=AOvVaw0bAvRQ3r5-gF9UOfvzqiNp) |
| 15 | [▪ E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems (Guan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00955&sa=D&source=editors&ust=1779048535944461&usg=AOvVaw0JA8HhzKsFm7vSBhsGyMbD) |
| 16 | [▪ When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI (Madrid-Garc\\'ia and Rujas, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00796&sa=D&source=editors&ust=1779048535944547&usg=AOvVaw1bkuv4bBnj11Cq2i9F_-54) |
| 17 | [▪ Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC) (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08921&sa=D&source=editors&ust=1779048535944630&usg=AOvVaw0mSFEhQujTadgynXZjapcO) |
| 18 | [▪ Tracking Conversations: Measuring Content and Identity Exposure on AI Chatbots (Jazlan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27438&sa=D&source=editors&ust=1779048535944718&usg=AOvVaw3m9kjulrTd7f_E4jDvsKiK) |
| 19 | [▪ Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing (Cui et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23711&sa=D&source=editors&ust=1779048535944804&usg=AOvVaw3NcQtzwHbye5Oyul7DUP8R) |
| 20 | [▪ Privacy Leakage via Output Label Space and Differentially Private Continual Learning (Tobaben et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.04680&sa=D&source=editors&ust=1779048535944887&usg=AOvVaw1cK0O8uheDIlCurYBfN8oG) |
| 21 | [▪ Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning (Chen, Yuan, and Kairouz, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22191&sa=D&source=editors&ust=1779048535944970&usg=AOvVaw04UpErCBjbO5xntQtX9K1B) |
| 22 | [▪ Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints (Pennekamp et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20927&sa=D&source=editors&ust=1779048535945057&usg=AOvVaw35Y9ESJ0X3mEhvqhREXrfr) |
| 23 | [▪ DP-FlogTinyLLM: Differentially private federated log anomaly detection using Tiny LLMs (Thompson, Sen, and Bhattacharya, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19118&sa=D&source=editors&ust=1779048535945142&usg=AOvVaw2PD2MRThnrV06xB6Ha6K9H) |
| 24 | [▪ Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks (Zhao, Kne\\v{z}evi\\'c, and K\\"aser, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18660&sa=D&source=editors&ust=1779048535945235&usg=AOvVaw2NDu08aj1qoIP9digW33Gi) |
| 25 | [▪ Breaking Euston: Recovering Private Inputs from Secure Inference by Exploiting Subspace Leakage (Zhao and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17238&sa=D&source=editors&ust=1779048535945322&usg=AOvVaw0hLkMr7dlSA5FJe4qm0pGw) |
| 26 | [▪ PolicyGapper: Automated Detection of Inconsistencies Between Google Play Data Safety Sections and Privacy Policies Using LLMs (Ferrari et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16128&sa=D&source=editors&ust=1779048535945406&usg=AOvVaw3v-EpBUc_3ojz-C-A-f5Jv) |
| 27 | [▪ Too Private to Tell: Practical Token Theft Attacks on Apple Intelligence (Zhou et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15637&sa=D&source=editors&ust=1779048535945487&usg=AOvVaw16NXBTWX4Rj3oL1meN6pPe) |
| 28 | [▪ Automated Profile Inference with Language Model Agents (Du et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.12402&sa=D&source=editors&ust=1779048535945567&usg=AOvVaw2040IIr8T5_9eGEd4tEEyb) |
| 29 | [▪ A2-DIDM: Privacy-preserving Accumulator-enabled Auditing for Distributed Identity of DNN Model (Xie et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2405.04108&sa=D&source=editors&ust=1779048535945666&usg=AOvVaw0eJZIYF-Xg_geW2SXKSzq-) |
| 30 | [▪ CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12342&sa=D&source=editors&ust=1779048535945775&usg=AOvVaw1XHpnO0akiyzYD5Axh4gUU) |
| 31 | [▪ Fully Homomorphic Encryption on Llama 3 model for privacy preserving LLM inference (Abdennebi, Kara, and Lahlou, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12168&sa=D&source=editors&ust=1779048535945860&usg=AOvVaw2MInjXO53JKd76RanIxjUE) |
| 32 | [▪ LLM-Redactor: An Empirical Evaluation of Eight Techniques for Privacy-Preserving LLM Requests (Agyemang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12064&sa=D&source=editors&ust=1779048535945944&usg=AOvVaw35UN7H_LJ6wB8JVbbVrGwZ) |
| 33 | [▪ Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game (Xie et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10717&sa=D&source=editors&ust=1779048535946025&usg=AOvVaw2qbp9TM2RShvhTURE2ySms) |
| 34 | [▪ ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying (Lyu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09747&sa=D&source=editors&ust=1779048535946107&usg=AOvVaw1_Fi6d528qA3XIklSQsghO) |
| 35 | [▪ Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout (Jiang and Wu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09056&sa=D&source=editors&ust=1779048535946188&usg=AOvVaw2-uTpxzdKn5qIWzM-Yez-7) |
| 36 | [▪ Security Concerns in Generative AI Coding Assistants: Insights from Online Discussions on GitHub Copilot (Ferreyra et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08352&sa=D&source=editors&ust=1779048535946275&usg=AOvVaw1tivpRRUR8crJSQyJzanFf) |
| 37 | [▪ ConfusionPrompt: Practical Private Inference for Online Large Language Models (Mai et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2401.00870&sa=D&source=editors&ust=1779048535946358&usg=AOvVaw2jGYBCqhE1bEY9ysn7Kctw) |
| 38 | [▪ Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation (Yoon et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06831&sa=D&source=editors&ust=1779048535946439&usg=AOvVaw1BmvVuqQ1sSsBw_46EKuP-) |
| 39 | [▪ Evaluating LLM-based Personal Information Extraction and Countermeasures (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2408.07291&sa=D&source=editors&ust=1779048535946519&usg=AOvVaw3hdK0Atk9fm_4kKlfC_fHQ) |
| 40 | [▪ Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics (Zhou et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.00446&sa=D&source=editors&ust=1779048535946602&usg=AOvVaw29bNYAQP7c4yAuHgP0Ww5k) |
| 41 | [▪ No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01350&sa=D&source=editors&ust=1779048535946703&usg=AOvVaw2KHVzs4XUlmPGOlBReF91_) |
| 42 | [▪ Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22751&sa=D&source=editors&ust=1779048535946794&usg=AOvVaw3XR5dTIvaVT_WkZoIpbyZR) |
| 43 | [▪ Towards Privacy-Preserving LLM Inference via Covariant Obfuscation (Technical Report) (Lin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01499&sa=D&source=editors&ust=1779048535946876&usg=AOvVaw0Mk_wnjUjrWGWJqEzgJSek) |
| 44 | [▪ "Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27204&sa=D&source=editors&ust=1779048535946960&usg=AOvVaw1kVmJY5C1spez0IjUMFafu) |
| 45 | [▪ Not All Entities are Created Equal: A Dynamic Anonymization Framework for Privacy-Preserving Retrieval-Augmented Generation (Zhu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26074&sa=D&source=editors&ust=1779048535947046&usg=AOvVaw2ldF1fGh2Iz7a1wJT2KZAg) |
| 46 | [▪ Computing Maximal Per-Record Leakage and Leakage-Distortion Functions for Privacy Mechanisms under Entropy-Constrained Adversaries (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00689&sa=D&source=editors&ust=1779048535947131&usg=AOvVaw1CcJb0Mn69QYHE8ToWU_td) |
| 47 | [▪ Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information (Zhan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.11680&sa=D&source=editors&ust=1779048535947213&usg=AOvVaw0x3dBWYWeISjGJ-bUBMrag) |
| 48 | [▪ A Critical Review on the Effectiveness and Privacy Threats of Membership Inference Attacks (Jebreel, S\\'anchez, and Domingo-Ferrer, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22987&sa=D&source=editors&ust=1779048535947323&usg=AOvVaw0g1Ir9MqAQr8e3iXhFXgM3) |
| 49 | [▪ Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22968&sa=D&source=editors&ust=1779048535947421&usg=AOvVaw2ZJgTkZ-GGnkJjD3e60sGs) |
| 50 | [▪ CIPL: A Target-Independent Framework for Channel-Inversion Privacy Leakage in Agents (Huang, Hou, and Meng, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22751&sa=D&source=editors&ust=1779048535947504&usg=AOvVaw3VmratG2ZLwEic6xw7VjQx) |
| 51 | [▪ RedacBench: Can AI Erase Your Secrets? (Jeon, Kim, and Shin, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20208&sa=D&source=editors&ust=1779048535947582&usg=AOvVaw1IzFNeoR3mKdu47gXzgLIC) |
| 52 | [▪ PlanTwin: Privacy-Preserving Planning Abstractions for Cloud-Assisted LLM Agents (Yu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18377&sa=D&source=editors&ust=1779048535947663&usg=AOvVaw2SG0db_NyIajE5Rb2Voqk3) |
| 53 | [▪ Revisiting Label Inference Attacks in Vertical Federated Learning: Why They Are Vulnerable and How to Defend (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18680&sa=D&source=editors&ust=1779048535947752&usg=AOvVaw20rbVGpqih47vuOuo4bGEz) |
| 54 | [▪ WebPII: Benchmarking Visual PII Detection for Computer-Use Agents (Zhao, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17357&sa=D&source=editors&ust=1779048535947832&usg=AOvVaw2N960TS0WAv-r877jENMG4) |
| 55 | [▪ SEAL-Tag: Self-Tag Evidence Aggregation with Probabilistic Circuits for PII-Safe Retrieval-Augmented Generation (Xie, Li, and Cheng, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17292&sa=D&source=editors&ust=1779048535947916&usg=AOvVaw0up0emMWZel_a0qgiChTxY) |
| 56 | [▪ VisualLeakBench: Auditing the Fragility of Large Vision-Language Models against PII Leakage and Social Engineering (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13385&sa=D&source=editors&ust=1779048535948000&usg=AOvVaw2t5fhlTrCth8k7TigmjfW0) |
| 57 | [▪ AEX: Non-Intrusive Multi-Hop Attestation and Provenance for LLM APIs (Guan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14283&sa=D&source=editors&ust=1779048535948081&usg=AOvVaw3uPgNi7ycXn6Yhruo2yzmp) |
| 58 | [▪ Membership Inference for Contrastive Pre-training Models with Text-only PII Queries (Cheng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14222&sa=D&source=editors&ust=1779048535948163&usg=AOvVaw1L70QHhftnecrf0TiG9hfp) |
| 59 | [▪ A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring (Anwar et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23163&sa=D&source=editors&ust=1779048535948258&usg=AOvVaw0ot9mPOzTXteNL85lt07Ol) |
| 60 | [▪ Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights (Fang and Kim, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13186&sa=D&source=editors&ust=1779048535948356&usg=AOvVaw2rPkpzawoZrHiVxJKlZYCJ) |
| 61 | [▪ STAMP: Selective Task-Aware Mechanism for Text Privacy (Tian et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12237&sa=D&source=editors&ust=1779048535948438&usg=AOvVaw0pEGxvolWp-zw7Vec1HKZl) |
| 62 | [▪ WebWeaver: Breaking Topology Confidentiality in LLM Multi-Agent Systems with Stealthy Context-Based Inference (Xiong et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11132&sa=D&source=editors&ust=1779048535948522&usg=AOvVaw2F32HuwL93RV9URPnSgGpH) |
| 63 | [▪ CLIOPATRA: Extracting Private Information from LLM Insights (Annamalai, Cristofaro, and Kairouz, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09781&sa=D&source=editors&ust=1779048535948602&usg=AOvVaw3g3eL9iv6uhgwJU_2N8cfl) |
| 64 | [▪ Automated TEE Adaptation with LLMs: Identifying, Transforming, and Porting Sensitive Functions in Programs (Han et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.13379&sa=D&source=editors&ust=1779048535948683&usg=AOvVaw0OQHiOfJVRJzkdkbc_xefX) |
| 65 | [▪ Exposing Citation Vulnerabilities in Generative Engines (Mochizuki et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06823&sa=D&source=editors&ust=1779048535948767&usg=AOvVaw1E0zuBjJL35eBVPVxMypwu) |
| 66 | [▪ BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints (Gill, Isak, and Dressman, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05608&sa=D&source=editors&ust=1779048535948852&usg=AOvVaw2JySQSwFpbV6SwBaKFt-oJ) |
| 67 | [▪ Real Money, Fake Models: Deceptive Model Claims in Shadow APIs (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01919&sa=D&source=editors&ust=1779048535948931&usg=AOvVaw2X9X1U-__89YspmjQyFIYW) |
| 68 | [▪ Towards Privacy-Preserving LLM Inference via Collaborative Obfuscation (Technical Report) (Lin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01499&sa=D&source=editors&ust=1779048535949010&usg=AOvVaw1457MbYHp6qMiWkUEUfMA7) |
| 69 | [▪ Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00196&sa=D&source=editors&ust=1779048535949090&usg=AOvVaw3ha_RvXWkOvDHFbILgHlfR) |
| 70 | [▪ Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent (Zhang and Zhang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23079&sa=D&source=editors&ust=1779048535949190&usg=AOvVaw0lD5Ss0qpi4dUgK7raH9ex) |
| 71 | [▪ Personal Information Parroting in Language Models (Subramani, Ghate, and Diab, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20580&sa=D&source=editors&ust=1779048535949321&usg=AOvVaw0Vo4Ogjzvu2iNPDXjJDZId) |
| 72 | [▪ A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage (Xin et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21035&sa=D&source=editors&ust=1779048535949420&usg=AOvVaw21YTI6oKkaWM-QCMs7ds3O) |
| 73 | [▪ FeatureBleed: Inferring Private Enriched Attributes From Sparsity-Optimized AI Accelerators (Asher et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18304&sa=D&source=editors&ust=1779048535949504&usg=AOvVaw1khOxUCpXhlqHzp7yczWj6) |
| 74 | [▪ Discovering Universal Activation Directions for PII Leakage in Language Models (Marchyok et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16980&sa=D&source=editors&ust=1779048535949602&usg=AOvVaw2nM7UHGHTtyIWL4GidOPiV) |
| 75 | [▪ Large-scale online deanonymization with LLMs (Lermen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16800&sa=D&source=editors&ust=1779048535949684&usg=AOvVaw3G3Hsa1ycZh5J7OCdq3iNQ) |
| 76 | [▪ NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (Goswami, Kumar, and Das, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15866&sa=D&source=editors&ust=1779048535949785&usg=AOvVaw0EcglIERj60PcQwzxJ-ES3) |
| 77 | [▪ PII-Bench: Evaluating Query-Aware Privacy Protection Systems (Shen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.18545&sa=D&source=editors&ust=1779048535949869&usg=AOvVaw0d83DQSrMMht6AjXOazNoZ) |
| 78 | [▪ SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs (Shaaban and Elmahallawy, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13529&sa=D&source=editors&ust=1779048535949952&usg=AOvVaw2FvLgT_TuFhQ5pkmNxMqrL) |
| 79 | [▪ Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs (Yan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11528&sa=D&source=editors&ust=1779048535950033&usg=AOvVaw2Qs3CkCwDappzz1gP7M6mK) |
| 80 | [▪ CAPID: Context-Aware PII Detection for Question-Answering Systems (Ponomarenko et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10074&sa=D&source=editors&ust=1779048535950119&usg=AOvVaw1V-766bjqyMxiR3gKgk_Hs) |
| 81 | [▪ Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs (Salem, Paverd, and Abdelnabi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08563&sa=D&source=editors&ust=1779048535950204&usg=AOvVaw1EkWyZf5Df9ur7ZkaYwQIo) |
| 82 | [▪ Retrieval Pivot Attacks in Hybrid RAG: Measuring and Mitigating Amplified Leakage from Vector Seeds to Graph Expansion (Thornton, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08668&sa=D&source=editors&ust=1779048535950292&usg=AOvVaw1IUcHmFvoVxmvNKa96X7C1) |
| 83 | [▪ Taipan: A Query-free Transfer-based Multiple Sensitive Attribute Inference Attack Solely from Publicly Released Graphs (Song and Palanisamy, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06700&sa=D&source=editors&ust=1779048535950377&usg=AOvVaw3NmON96qimghLLfIiKOZf5) |
| 84 | [▪ Do Vision-Language Models Respect Contextual Integrity in Location Disclosure? (Yang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05023&sa=D&source=editors&ust=1779048535950459&usg=AOvVaw3gOoNBSHDIfujtYILvtJJQ) |
| 85 | [▪ PriMod4AI: Lifecycle-Aware Privacy Threat Modeling for AI Systems using LLM (Savaliya et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04927&sa=D&source=editors&ust=1779048535950541&usg=AOvVaw00AM_Bfd1hFYbCZaFRh0rw) |
| 86 | [▪ Evaluating the Vulnerability Landscape of LLM-Generated Smart Contracts (Do, Sohrabi, and Hassan, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04039&sa=D&source=editors&ust=1779048535950628&usg=AOvVaw2qz3bRcrXTdYdLU91qZUuO) |
| 87 | [▪ No More Hidden Pitfalls? Exposing Smart Contract Bad Practices with LLM-Powered Hybrid Analysis (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15179&sa=D&source=editors&ust=1779048535950714&usg=AOvVaw2QQumcRr31Mz3gVN5x7gXe) |
| 88 | [▪ Security Analysis of ChatGPT: Threats and Privacy Risks (Xiang, Li, and Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09426&sa=D&source=editors&ust=1779048535950796&usg=AOvVaw3JBCCfqLoqZwey1_lpN42O) |
| 89 | [▪ Enhancing Smart Contract Vulnerability Detection in DApps Leveraging Fine-Tuned LLM (Bu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.05006&sa=D&source=editors&ust=1779048535950876&usg=AOvVaw1ltBUGQPhpmKXAh1r7S9mF) |
| 90 | [▪ Decoupling Generalizability and Membership Privacy Risks in Neural Networks (Fang and Kim, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02296&sa=D&source=editors&ust=1779048535950976&usg=AOvVaw3W0D3CIQjIAum0942fE0RL) |
| 91 | [▪ Semantic Leakage from Image Embeddings (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22929&sa=D&source=editors&ust=1779048535951057&usg=AOvVaw0nqc_CJ775MAujh1OY57NQ) |
| 92 | [▪ Protecting Private Code in IDE Autocomplete using Differential Privacy (Grigorenko et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22935&sa=D&source=editors&ust=1779048535951141&usg=AOvVaw2dyzRma4CIzfvPjW1jIR8J) |
| 93 | [▪ Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models (Westphal, Navaie, and Rosas, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22818&sa=D&source=editors&ust=1779048535951228&usg=AOvVaw2y7QPAqmkdQm8G2iUI9wso) |
| 94 | [▪ Okara: Detection and Attribution of TLS Man-in-the-Middle Vulnerabilities in Android Apps with Foundation Models (Yang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22770&sa=D&source=editors&ust=1779048535951322&usg=AOvVaw2J0AMGOpTLJye_YDrDwkJo) |
| 95 | [▪ AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs (Kim and Kang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22710&sa=D&source=editors&ust=1779048535951410&usg=AOvVaw2SPNzPHavFQrPbeydVMDBq) |
| 96 | [▪ OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions (Huang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.04328&sa=D&source=editors&ust=1779048535951492&usg=AOvVaw0dLhSCt4MHwYWn7vCUCKJY) |
| 97 | [▪ Enhancing Membership Inference Attacks on Diffusion Models from a Frequency-Domain Perspective (Lian et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.20955&sa=D&source=editors&ust=1779048535951572&usg=AOvVaw3DrXySulyCpFNrki6Q7mMH) |
| 98 | [▪ Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models (Luo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.19373&sa=D&source=editors&ust=1779048535951653&usg=AOvVaw2lFLiMuXNekCShfaJapATt) |
| 99 | [▪ Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning (Gu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05739&sa=D&source=editors&ust=1779048535951746&usg=AOvVaw13sC1UJFGIKcWrtXHtKgwF) |
| 100 | [▪ CanaryBench: Stress Testing Privacy Leakage in Cluster-Level Conversation Summaries (Mehta, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.18834&sa=D&source=editors&ust=1779048535951831&usg=AOvVaw1O2YD4vIsq5doM2I5DDFBD) |
| 101 | [▪ Beyond Data Privacy: New Privacy Risks for Large Language Models (Du et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.14278&sa=D&source=editors&ust=1779048535951908&usg=AOvVaw1RHPmjR1a2t0w4lPPRChAJ) |
| 102 | [▪ Connect the Dots: Knowledge Graph-Guided Crawler Attack on Retrieval-Augmented Generation Systems (Yao et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15678&sa=D&source=editors&ust=1779048535951990&usg=AOvVaw2wlY-GU7tsPXMymI6AEvoJ) |
| 103 | [▪ NeuroFilter: Privacy Guardrails for Conversational LLM Agents (Das and Fioretto, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14660&sa=D&source=editors&ust=1779048535952070&usg=AOvVaw1vr7Ts1CRuxpXgmJAB4UPP) |
| 104 | [▪ VidLeaks: Membership Inference Attacks Against Text-to-Video Models (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11210&sa=D&source=editors&ust=1779048535952150&usg=AOvVaw3S9AVkC-OhXtMdo0vLgA-a) |
| 105 | [▪ On Membership Inference Attacks in Knowledge Distillation (Cui, Zhang, and Pei, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11837&sa=D&source=editors&ust=1779048535952244&usg=AOvVaw2BE38kQshFfI5J-KKcw1C7) |
| 106 | [▪ Automated Generation of Accurate Privacy Captions From Android Source Code Using Large Language Models (Jain et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06276&sa=D&source=editors&ust=1779048535952329&usg=AOvVaw1NCbjbXlClqjo2XMMmGMaS) |
| 107 | [▪ Leveraging Membership Inference Attacks for Privacy Measurement in Federated Learning for Remote Sensing Images (Duong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06200&sa=D&source=editors&ust=1779048535952412&usg=AOvVaw1oTg3WJ7uJtKFL1gOTDRc5) |
| 108 | [▪ PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (Shahariar et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05739&sa=D&source=editors&ust=1779048535952497&usg=AOvVaw26Z5M3ZuSNfY5geF7-k_nG) |
| 109 | [▪ Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset (Li, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05918&sa=D&source=editors&ust=1779048535952581&usg=AOvVaw25i60a0H5iT2JFlzJny2J_) |
| 110 | [▪ DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04641&sa=D&source=editors&ust=1779048535952666&usg=AOvVaw1ugfgAOlhsgWgcTlMRGVCI) |
| 111 | [▪ SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems (Bodea et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03979&sa=D&source=editors&ust=1779048535952755&usg=AOvVaw0vFRWfYBIrkygNOLcq37E1) |
| 112 | [▪ MemHunter: Automated and Verifiable Memorization Detection at Dataset-scale in LLMs (Wu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2412.07261&sa=D&source=editors&ust=1779048535952839&usg=AOvVaw39laRBrEQhoD2MovPFNI1F) |
| 113 | [▪ I Large Language Models possono nascondere un testo in un altro testo della stessa lunghezza (Norelli and Bronstein, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20075&sa=D&source=editors&ust=1779048535952921&usg=AOvVaw1Wid3sxCPbJsElYfQH596Q) |
| 114 | [▪ InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split Inference (Deng, Lu, and Duan, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13365&sa=D&source=editors&ust=1779048535953003&usg=AOvVaw3QGvmYA6_bOV5iKh-fcYp9) |
| 115 | [▪ Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric (Shvartzshnaider and Duddu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03735&sa=D&source=editors&ust=1779048535953085&usg=AOvVaw32bL8p5q9P_L7D7boIPGus) |
| 116 | [▪ Detecting Malicious Entra OAuth Apps with LLM-Based Permission Risk Scoring (Mahara, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15781&sa=D&source=editors&ust=1779048535953179&usg=AOvVaw1qm9pVuKEnuP7TvtqJJihB) |
| 117 | [▪ PerProb: Indirectly Evaluating Memorization in Large Language Models (Liao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14600&sa=D&source=editors&ust=1779048535953262&usg=AOvVaw2rwmb14fkwMOqQiKRQsPAy) |
| 118 | [▪ CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs (Arachchige et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12914&sa=D&source=editors&ust=1779048535953342&usg=AOvVaw1no8LKu7R9tW2nJRbHPRSa) |
| 119 | [▪ Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection (Nguyen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12313&sa=D&source=editors&ust=1779048535953423&usg=AOvVaw31Uk22nFquybbf_6_Va2Fq) |
| 120 | [▪ Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering (Rashid et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2310.16152&sa=D&source=editors&ust=1779048535953505&usg=AOvVaw0YVi6diHMPZ1FBQtZ3xk_8) |
| 121 | [▪ Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08326&sa=D&source=editors&ust=1779048535953587&usg=AOvVaw1pJ5q_h-Y2cF2zhcmvKanE) |
| 122 | [▪ Exposing and Defending Membership Leakage in Vulnerability Prediction Models (Liao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08291&sa=D&source=editors&ust=1779048535953673&usg=AOvVaw3K7yVe11rr4epyhGQv7854) |
| 123 | [▪ Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07814&sa=D&source=editors&ust=1779048535953758&usg=AOvVaw1bjuSotaEUFJKrXykV33gk) |
| 124 | [▪ CKG-LLM: LLM-Assisted Detection of Smart Contract Access Control Vulnerabilities Based on Knowledge Graphs (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06846&sa=D&source=editors&ust=1779048535953871&usg=AOvVaw2Dd8HGwSFByif2MUyyjWte) |
| 125 | [▪ Sell Data to AI Algorithms Without Revealing It: Secure Data Valuation and Sharing via Homomorphic Encryption (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06033&sa=D&source=editors&ust=1779048535953958&usg=AOvVaw1iGyXh8TkeYzGEyoq2Nywn) |
| 126 | [▪ When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.18874&sa=D&source=editors&ust=1779048535954057&usg=AOvVaw2ZSe1DWnZrnUySGjY5Qti1) |
| 127 | [▪ WildCode: An Empirical Analysis of Code Generated by ChatGPT (Khanmohammadi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04259&sa=D&source=editors&ust=1779048535954182&usg=AOvVaw1vFQBsg0qrgALQpLJSmZkL) |
| 128 | [▪ Towards Contextual Sensitive Data Detection (Telkamp and Hulsebos, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04120&sa=D&source=editors&ust=1779048535954280&usg=AOvVaw1O1LX4v0XNj6Tl6_D36bsG) |
| 129 | [▪ Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review (Allana, Kankanhalli, and Dara, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.02828&sa=D&source=editors&ust=1779048535954369&usg=AOvVaw2ok47droNatUM1tHFEdY6-) |
| 130 | [▪ Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference (Luo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09442&sa=D&source=editors&ust=1779048535954451&usg=AOvVaw2YBvYBfBN8NZP_V_bjkPLj) |
| 131 | [▪ GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22441&sa=D&source=editors&ust=1779048535954537&usg=AOvVaw1RRJhhblXt1oK6-wpZY-7Z) |
| 132 | [▪ Vision Token Masking Alone Cannot Prevent PHI Leakage in Medical Document OCR: A Systematic Evaluation (Young, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18272&sa=D&source=editors&ust=1779048535954621&usg=AOvVaw0ZwFf7gDoKk788Z2Qi_sfJ) |
| 133 | [▪ Privacy Auditing of Multi-domain Graph Pre-trained Model under Membership Inference Attacks (Luo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17989&sa=D&source=editors&ust=1779048535954703&usg=AOvVaw0nILm3CT7V5jlyPiD9Dr5B) |
| 134 | [▪ Membership Inference Attacks Beyond Overfitting (Khalil et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16792&sa=D&source=editors&ust=1779048535954811&usg=AOvVaw2GA12tCxV9axNPMu_igWyj) |
| 135 | [▪ Password Strength Analysis Through Social Network Data Exposure: A Combined Approach Relying on Data Reconstruction and Generative Models (Atzori et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16716&sa=D&source=editors&ust=1779048535954910&usg=AOvVaw2uxuo1WpGIAhOCoRcRHY3g) |
| 136 | [▪ Privacy Preserving In-Context-Learning Framework for Large Language Models (Bhusal et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.13625&sa=D&source=editors&ust=1779048535955012&usg=AOvVaw1h2LDgvn67yuylU8ChwWu3) |
| 137 | [▪ Confidential Prompting: Privacy-preserving LLM Inference on Cloud (Li, Gim, and Zhong, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.19134&sa=D&source=editors&ust=1779048535955099&usg=AOvVaw3SetdLOE7ZHPghWaveg6hO) |
| 138 | [▪ Quantifying Privacy Leakage in Split Inference via Fisher-Approximated Shannon Information Analysis (Deng et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.10016&sa=D&source=editors&ust=1779048535955183&usg=AOvVaw3Y-nT7EiYRuRoVHK2yt24F) |
| 139 | [▪ Observational Auditing of Label Privacy (Kalemaj et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14084&sa=D&source=editors&ust=1779048535955261&usg=AOvVaw31Wz4uSNWDhdBelp0rtdyG) |
| 140 | [▪ Explainable Transformer-Based Email Phishing Classification with Adversarial Robustness (P, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12085&sa=D&source=editors&ust=1779048535955341&usg=AOvVaw1PtvdLtc86w6JQrNctUC_5) |
| 141 | [▪ BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12043&sa=D&source=editors&ust=1779048535955422&usg=AOvVaw31LsUUiXNxe-nyZ6zP0H6C) |
| 142 | [▪ Privacy Challenges and Solutions in Retrieval-Augmented Generation-Enhanced LLMs for Healthcare Chatbots: A Review of Applications, Risks, and Future Directions (Guan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11347&sa=D&source=editors&ust=1779048535955511&usg=AOvVaw26XH_htiDvpVoOGx-WJGur) |
| 143 | [▪ How Worrying Are Privacy Attacks Against Machine Learning? (Domingo-Ferrer, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10516&sa=D&source=editors&ust=1779048535955590&usg=AOvVaw3jSRF5xNwLx7-_gioesG_W) |
| 144 | [▪ Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07637&sa=D&source=editors&ust=1779048535955670&usg=AOvVaw1Q8WtegcMyOKzQRLJIujd6) |
| 145 | [▪ SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought (Batra et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07772&sa=D&source=editors&ust=1779048535955758&usg=AOvVaw32YuszVtVenI_nZ5Xg-zZT) |
| 146 | [▪ $\\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05319&sa=D&source=editors&ust=1779048535955840&usg=AOvVaw02hr_bvOOtfQdBTAb1jLiU) |
| 147 | [▪ Whisper Leak: a side-channel attack on Large Language Models (McDonald and Or, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03675&sa=D&source=editors&ust=1779048535955921&usg=AOvVaw32vCfMLku86u2Wypt-l3Ju) |
| 148 | [▪ Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03248&sa=D&source=editors&ust=1779048535956002&usg=AOvVaw0ESehNJHNg6b1BmYZxtGHg) |
| 149 | [▪ PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases (Bae et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17336&sa=D&source=editors&ust=1779048535956087&usg=AOvVaw13nXFjvgbb4s50poMCll2v) |
| 150 | [▪ Exploring the limits of strong membership inference attacks on large language models (Hayes et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.18773&sa=D&source=editors&ust=1779048535956168&usg=AOvVaw15wwx4bqgB4b83ozLK95Wx) |
| 151 | [▪ FTSmartAudit: A Knowledge Distillation-Enhanced Framework for Automated Smart Contract Auditing Using Fine-Tuned LLMs (Wei et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.13918&sa=D&source=editors&ust=1779048535956250&usg=AOvVaw3wSdMIT4U7843Ciq89VrBK) |
| 152 | [▪ Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents (Grosse and Ebert, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27275&sa=D&source=editors&ust=1779048535956333&usg=AOvVaw2gjSvvUd-mpShJrBRE9vzS) |
| 153 | [▪ NetEcho: From Real-World Streaming Side-Channels to Full LLM Conversation Recovery (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25472&sa=D&source=editors&ust=1779048535956414&usg=AOvVaw1hkkdB6usFuDJzGrW0B9E_) |
| 154 | [▪ Learning to Attack: Uncovering Privacy Risks in Sequential Data Releases (Cui, Zhang, and Pei, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24807&sa=D&source=editors&ust=1779048535956494&usg=AOvVaw3nyb3eiUleuuv7_gr8CTy9) |
| 155 | [▪ Differential Privacy: Gradient Leakage Attacks in Federated Learning Environments (Fernandez-de-Retana et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23931&sa=D&source=editors&ust=1779048535956575&usg=AOvVaw1aCjs04pNnB7aLV3tlzoMb) |
| 156 | [▪ Membership Inference Attacks on Recommender System: A Survey (He et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11080&sa=D&source=editors&ust=1779048535956654&usg=AOvVaw24j9jUitc2baCuA1G7cR0H) |
| 157 | [▪ LLMs can hide text in other text of the same length (Norelli and Bronstein, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20075&sa=D&source=editors&ust=1779048535956751&usg=AOvVaw2oRuR1ErHuDSoygsZ8kzQP) |
| 158 | [▪ LLMs can hide text in other text of the same length.ipynb (Norelli and Bronstein, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20075&sa=D&source=editors&ust=1779048535956835&usg=AOvVaw2V6rIb4hgwIACamhRRDDZp) |
| 159 | [▪ The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models (Dodd et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19773&sa=D&source=editors&ust=1779048535956919&usg=AOvVaw2Zn4cTVpiG9EGV-NnjrTv9) |
| 160 | [▪ CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage (Mashnoor et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19676&sa=D&source=editors&ust=1779048535957000&usg=AOvVaw1tdReLqu_l42l72FlUnPqj) |
| 161 | [▪ Exploring Membership Inference Vulnerabilities in Clinical Large Language Models (Nemecek et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18674&sa=D&source=editors&ust=1779048535957083&usg=AOvVaw1TzBpscNfUJyVU5sHHv7om) |
| 162 | [▪ Evaluating Large Language Models in detecting Secrets in Android Apps (Alecci et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18601&sa=D&source=editors&ust=1779048535957171&usg=AOvVaw1Ak3ZZUX6JDFoSWqhS-pmv) |
| 163 | [▪ The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers (Makroo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16122&sa=D&source=editors&ust=1779048535957264&usg=AOvVaw2q0WaWZvJp-_6ZAz8lmCTW) |
| 164 | [▪ Membership Inference over Diffusion-models-based Synthetic Tabular Data (Cheng and Bahmani, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16037&sa=D&source=editors&ust=1779048535957357&usg=AOvVaw0tMqoFBFuOyyZ42W6yJdLD) |
| 165 | [▪ DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing (Qiao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15303&sa=D&source=editors&ust=1779048535957443&usg=AOvVaw00RpIyD-jGjInxchZB3Muf) |
| 166 | [▪ AndroByte: LLM-Driven Privacy Analysis through Bytecode Summarization and Dynamic Dataflow Call Graph Generation (Khatun et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15112&sa=D&source=editors&ust=1779048535957531&usg=AOvVaw0mfboQgGJxJr0pMEKIuHIA) |
| 167 | [▪ Lost in the Averages: A New Specific Setup to Evaluate Membership Inference Attacks Against Machine Learning Models (Kr\\v{c}o et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.15423&sa=D&source=editors&ust=1779048535957621&usg=AOvVaw1hU4q7zChXUOg9JBZcPodv) |
| 168 | [▪ Early Signs of Steganographic Capabilities in Frontier LLMs (Zolkowski et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02737&sa=D&source=editors&ust=1779048535957726&usg=AOvVaw2hmxvq5Y-VsasvR3usc2Rz) |
| 169 | [▪ Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09655&sa=D&source=editors&ust=1779048535957818&usg=AOvVaw0j_2chhDU1m5auG94zJ4ku) |
| 170 | [▪ The Model's Language Matters: A Comparative Privacy Analysis of LLMs (Mishra, Boutet, and Magnana, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08813&sa=D&source=editors&ust=1779048535957906&usg=AOvVaw1b2KTgB2MUxfyUAIMIi9Tc) |
| 171 | [▪ PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing (Hughes et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07452&sa=D&source=editors&ust=1779048535957991&usg=AOvVaw2aM4Aj3Ja8OTZ9AdPI0rp0) |
| 172 | [▪ Membership Inference Attacks on LLM-based Recommender Systems (He et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.18665&sa=D&source=editors&ust=1779048535958072&usg=AOvVaw2RBO0CzDsv_qAYpjzPf0zM) |
| 173 | [▪ Empirical Comparison of Membership Inference Attacks in Deep Transfer Learning (Bai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05753&sa=D&source=editors&ust=1779048535958159&usg=AOvVaw3Tk04jIANxEK9-vbg9Y3nN) |
| 174 | [▪ Membership Inference Attacks on Tokenizers of Large Language Models (Tong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05699&sa=D&source=editors&ust=1779048535958251&usg=AOvVaw1aaAgyxTUc5p87HLz2-sDo) |
| 175 | [▪ Can We Infer Confidential Properties of Training Data from LLMs? (Huang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10364&sa=D&source=editors&ust=1779048535958343&usg=AOvVaw0HSLpkSoLrIy1lnJRYSnYL) |
| 176 | [▪ Rethinking Exact Unlearning under Exposure: Extracting Forgotten Data under Exact Unlearning in Large Language Model (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.24379&sa=D&source=editors&ust=1779048535958430&usg=AOvVaw3un8Knu8uXZcQBdvkZCz5h) |
| 177 | [▪ Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent (Zhang, Guo, and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.01344&sa=D&source=editors&ust=1779048535958518&usg=AOvVaw3_eH1hgW8io4EYzLvH5t_O) |
| 178 | [▪ MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data (German et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.13046&sa=D&source=editors&ust=1779048535958616&usg=AOvVaw0XbAQOUsRu0SJkCUGsONKT) |
| 179 | [▪ Activation Functions Considered Harmful: Recovering Neural Network Weights through Controlled Channels (Spielman et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.19142&sa=D&source=editors&ust=1779048535958715&usg=AOvVaw0fqUibxFNWdJ85JC2OlosG) |
| 180 | [▪ Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers (Green et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15674&sa=D&source=editors&ust=1779048535958800&usg=AOvVaw1XBj1LV-Z73ZA2-TAd8-Wp) |
| 181 | [▪ An Ethically Grounded LLM-Based Approach to Insider Threat Synthesis and Detection (Gelman, Hastings, and Kenley, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06920&sa=D&source=editors&ust=1779048535958884&usg=AOvVaw0yh27_Frqa2OJD7RdP8sib) |
| 182 | [▪ Inducing Uncertainty on Open-Weight Models for Test-Time Privacy in Image Recognition (Ashiq et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11625&sa=D&source=editors&ust=1779048535958988&usg=AOvVaw3MBDC_0ZGRl8_xp-e93XrK) |
| 183 | [▪ Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25525&sa=D&source=editors&ust=1779048535959095&usg=AOvVaw1H4oFi9V2wC1yEEF8PoC6t) |
| 184 | [▪ When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs (Wei et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.01076&sa=D&source=editors&ust=1779048535959196&usg=AOvVaw2MYH9Efa6uSAMHBa_Jv1OK) |
| 185 | [▪ Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models (Fu et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24488&sa=D&source=editors&ust=1779048535959296&usg=AOvVaw00B-bnV6YVhHgfjWIzEJs3) |
| 186 | [▪ Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts (Roh and Kumar, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24807&sa=D&source=editors&ust=1779048535959385&usg=AOvVaw3CW7qHCR8vO988H9LIHJ64) |
| 187 | [▪ MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction (Waterloo et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23459&sa=D&source=editors&ust=1779048535959469&usg=AOvVaw0IWO2kn7AeWfpTv3c4Qmai) |
| 188 | [▪ RecPS: Privacy Risk Scoring for Recommender Systems (He, Gu, and Chen, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18365&sa=D&source=editors&ust=1779048535959552&usg=AOvVaw2B7Xd83VbyQ_YCOpfTewFv) |
| 189 | [▪ Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents (Fang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21142&sa=D&source=editors&ust=1779048535959640&usg=AOvVaw20jXYHf04gyBAMX0RC7NHo) |
| 190 | [▪ Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents (Ngong et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.18509&sa=D&source=editors&ust=1779048535959747&usg=AOvVaw3wsrjAtkb-DQDBY8kXR2Lg) |
| 191 | [▪ Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14214&sa=D&source=editors&ust=1779048535959839&usg=AOvVaw2D1I0FMKaiD0gCzlFcy-xD) |
| 192 | [▪ Auditing Prompt Caching in Language Model APIs (Gu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.07776&sa=D&source=editors&ust=1779048535959924&usg=AOvVaw0QNTvvXQk--EubcTBckkHA) |
| 193 | [▪ Measuring the Accuracy and Effectiveness of PII Removal Services (He et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.06989&sa=D&source=editors&ust=1779048535960008&usg=AOvVaw3LEtj5CsNOnmEX3rWHU_zq) |
| 194 | [▪ Clio-X: AWeb3 Solution for Privacy-Preserving AI Access to Digital Archives (Lemieux et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08853&sa=D&source=editors&ust=1779048535960092&usg=AOvVaw0HP5LIb3opXShSioO8DCdX) |
| 195 | [▪ PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER) (Mancera et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02966&sa=D&source=editors&ust=1779048535960178&usg=AOvVaw3QJSkCElVvsDpSdeB-9KSb) |
| 196 | [▪ The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation (Xiong et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05578&sa=D&source=editors&ust=1779048535960262&usg=AOvVaw1yjVUzIsPfNqDKyjZFmvRH) |
| 197 | [▪ Real-Time Privacy Risk Measurement with Privacy Tokens for Gradient Leakage (Meng et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.02913&sa=D&source=editors&ust=1779048535960347&usg=AOvVaw1EEnJkjCVFIl_ttTZJXBW8) |
| 198 | [▪ Exploring Privacy and Fairness Risks in Sharing Diffusion Models: An Adversarial Perspective (Luo et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.18607&sa=D&source=editors&ust=1779048535960445&usg=AOvVaw1rWGzuuWyVS93SZFQczW2j) |
| 199 | [▪ LLM-PBE: Assessing Data Privacy in Large Language Models (Li et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.12787&sa=D&source=editors&ust=1779048535960528&usg=AOvVaw2Fiv0K3Jgsv-ARXZTsMtmo) |
| 200 | [▪ PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action (Shao et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00138&sa=D&source=editors&ust=1779048535960609&usg=AOvVaw1C7SKGcnoAxmgBj96OQxzf) |
| 201 | [▪ Privacy-preserving Universal Adversarial Defense for Black-box Models (Li et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.10647&sa=D&source=editors&ust=1779048535960690&usg=AOvVaw1T--E6_9jRJjzEIzUxvUtL) |
| 202 | [▪ DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts (Sun et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08930%23&sa=D&source=editors&ust=1779048535960793&usg=AOvVaw2HNipE5vzfMX1_rySskMij) |
| 203 | [▪ Casper: Prompt Sanitization for Protecting User Privacy in Web-Based Large Language Models (Chong et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.07004&sa=D&source=editors&ust=1779048535960886&usg=AOvVaw01xuDWzZ6MeQAj8rJhj5Sy) |

|     |
| --- |
| Sensitive Information Disclosure |

**>**

**<**

‍

#### Insecure Plugin Design and Plugin Compromise

Covers:

- OWASP LLM 07: Insecure Plugin Design
- MITRE ATLAS Execution & Privilege Escalation

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Insecure Plugin Design and Plugin Compromise

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills (Zhuang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13940&sa=D&source=editors&ust=1779048536772391&usg=AOvVaw3lLnTxpgf-WCWaFfLAyIQg) |
| 3 | [▪ Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09594&sa=D&source=editors&ust=1779048536772598&usg=AOvVaw3yMbzdOv61zXEndKVj9Lut) |
| 4 | [▪ Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows (Ruan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05969&sa=D&source=editors&ust=1779048536772723&usg=AOvVaw17mBxHt-Jn4o5N5ka6WAyB) |
| 5 | [▪ SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents (Ouyang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03353&sa=D&source=editors&ust=1779048536772825&usg=AOvVaw02d32MlCx7XniCGLLxp04n) |
| 6 | [▪ ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24414&sa=D&source=editors&ust=1779048536772925&usg=AOvVaw1cpNwQ30kJZd3FdUGayGbp) |
| 7 | [▪ Paladin: A Policy Framework for Securing Cloud APIs by Combining Application Context with Generative AI (Priya, Stephen, and Natarajan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10228&sa=D&source=editors&ust=1779048536773029&usg=AOvVaw2fFqO6ZyhPrkbCdpZITmt5) |
| 8 | [▪ OpenPort Protocol: A Security Governance Specification for AI Agent Tool Access (Zhu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20196&sa=D&source=editors&ust=1779048536773126&usg=AOvVaw3-3U_ZeDQTsFyZ0VGishMb) |
| 9 | [▪ Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform (Xie et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15528&sa=D&source=editors&ust=1779048536773222&usg=AOvVaw0Q5tr9bm3XNPa-qLdPcJLz) |
| 10 | [▪ Keep the Lights On, Keep the Lengths in Check: Plug-In Adversarial Detection for Time-Series LLMs in Energy Forecasting (Ma et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12154&sa=D&source=editors&ust=1779048536773322&usg=AOvVaw2Qj-_3_Sa7HDdl_NwJrq9U) |
| 11 | [▪ When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins (Kaya et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05797&sa=D&source=editors&ust=1779048536773428&usg=AOvVaw0CCK32_FF9TBVwkBjHX-z9) |
| 12 | [▪ Security study based on the Chatgptplugin system: ldentifying Security Vulnerabilities (Ren, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21128&sa=D&source=editors&ust=1779048536773530&usg=AOvVaw1wHLGC4iHi_R3ej-fwKKkl) |
| 13 | [▪ When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents (Kozak, Moghaddam, and Sivaraman, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09329&sa=D&source=editors&ust=1779048536773628&usg=AOvVaw0g9RLFUiVJnT3STPYOuBW6) |
| 14 | [▪ PromptChain: A Decentralized Web3 Architecture for Managing AI Prompts as Digital Assets (Bara, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09579&sa=D&source=editors&ust=1779048536773723&usg=AOvVaw2-5o5iYoyMiuU7W7vnLfRZ) |
| 15 | [▪ We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06250&sa=D&source=editors&ust=1779048536773818&usg=AOvVaw0XuGe-rcdEU1ZTb_EoI1NJ) |
| 16 | [▪ The Philosopher's Stone: Trojaning Plugins of Large Language Models (Dong et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.00374&sa=D&source=editors&ust=1779048536773911&usg=AOvVaw0OB4zLJ4eV4Ww2ObA4qLmS) |
| 17 | [▪ LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins (Iqbal, Kohno, and Roesner, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2309.10254&sa=D&source=editors&ust=1779048536774015&usg=AOvVaw35nLfJ_vpVZ2VJr77xbL2K) |

|     |
| --- |
| Insecure Plugin Design and Plugin Compromise |

**>**

**<**

‍

#### Hallucination Squatting and Phishing

Covers:

- MITRE ATLAS Initial Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Hallucination Squatting and Phishing

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21457&sa=D&source=editors&ust=1779048535913299&usg=AOvVaw1_FclOUzLb2v4CqBQd5gh2) |
| 3 | [▪ REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations (Liang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12813&sa=D&source=editors&ust=1779048535913426&usg=AOvVaw3H0kf6RqTT3uXv6dPSv9S_) |
| 4 | [▪ Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data (Vafa, Roy, and Nilizadeh, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11268&sa=D&source=editors&ust=1779048535913486&usg=AOvVaw27AEmSlxPEZl2sLnTbthZH) |
| 5 | [▪ LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning (Spracklen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01047&sa=D&source=editors&ust=1779048535913535&usg=AOvVaw1I2uZLKtowacPUb4B4RGKK) |
| 6 | [▪ CyberCane: Neuro-Symbolic RAG for Privacy-Preserving Phishing Detection with Formal Ontology Reasoning (Hakim et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23563&sa=D&source=editors&ust=1779048535913586&usg=AOvVaw126dNhYIK2xMDhViyoZ6Vf) |
| 7 | [▪ GuardPhish: Securing Open-Source LLMs from Phishing Abuse (Mishra, Varshney, and Sahithi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17313&sa=D&source=editors&ust=1779048535913635&usg=AOvVaw24sa1Zl3Vy7AqwkBOFuk-Y) |
| 8 | [▪ Segment-Level Coherence for Robust Harmful Intent Probing in LLMs (He et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14865&sa=D&source=editors&ust=1779048535913684&usg=AOvVaw3DQbQhFhBMOK7pCZ_AH6AD) |
| 9 | [▪ Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR) (Fang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11141&sa=D&source=editors&ust=1779048535913732&usg=AOvVaw02UErzsTErMG5L-pcCjMil) |
| 10 | [▪ Hijacking Text Heritage: Hiding the Human Signature through Homoglyphic Substitution (Dilworth, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10271&sa=D&source=editors&ust=1779048535913784&usg=AOvVaw1IhbDEi_zItGxLb8_Coz0Q) |
| 11 | [▪ The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs (Toth, Bisztray, and Gruschka, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21448&sa=D&source=editors&ust=1779048535913834&usg=AOvVaw2-FB6bvGizo5CmJW-31JVN) |
| 12 | [▪ Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models (Guesmi and Shafique, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15940&sa=D&source=editors&ust=1779048535913882&usg=AOvVaw0jSO8VLksPk8yxEWbpulqJ) |
| 13 | [▪ Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents (Basu, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10060&sa=D&source=editors&ust=1779048535913928&usg=AOvVaw1BByx16-F5ReDy3mmqE_Df) |
| 14 | [▪ PhishDebate: An LLM-Based Multi-Agent Framework for Phishing Website Detection (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15656&sa=D&source=editors&ust=1779048535913976&usg=AOvVaw2TrCQFx8sY6q8qMnTk9SG1) |
| 15 | [▪ Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection (Gao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21593&sa=D&source=editors&ust=1779048535914021&usg=AOvVaw2E3OToCuG3uI59wDGhaH_Y) |
| 16 | [▪ PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20717&sa=D&source=editors&ust=1779048535914067&usg=AOvVaw2hPkHf8wegGvHjq-dUbOzf) |
| 17 | [▪ GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models (Xu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06718&sa=D&source=editors&ust=1779048535914137&usg=AOvVaw1clxGL1sPISNzY_cXv_QyA) |
| 18 | [▪ Hallucination-Resistant Security Planning with a Large Language Model (Hammar, Alpcan, and Lupu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05279&sa=D&source=editors&ust=1779048535914185&usg=AOvVaw2qmE0l-FY_w9P4GltCkco0) |
| 19 | [▪ Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection (Hasan and BusiReddyGari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02641&sa=D&source=editors&ust=1779048535914231&usg=AOvVaw1kDbjXSHciCGNamYXtoXWl) |
| 20 | [▪ When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems (Xu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00428&sa=D&source=editors&ust=1779048535914282&usg=AOvVaw0E7cHgUUI10RWbY2yCsJlp) |
| 21 | [▪ Eliciting Least-to-Most Reasoning for Phishing URL Detection (Trikilis et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20270&sa=D&source=editors&ust=1779048535914329&usg=AOvVaw0Z3OOofoP2zKgKvpY8l9ol) |
| 22 | [▪ Predictive Coding and Information Bottleneck for Hallucination Detection in Large Language Models (Bhatt, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15652&sa=D&source=editors&ust=1779048535914374&usg=AOvVaw0f0g8DBDrxzm5IRD1br3k7) |
| 23 | [▪ CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models (Min et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14310&sa=D&source=editors&ust=1779048535914422&usg=AOvVaw0B4Tavl55pH4c9NVZPiyyM) |
| 24 | [▪ Can Large Language Models Automate Phishing Warning Explanations? A Controlled Experiment on Effectiveness and User Perception (Cau et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07916&sa=D&source=editors&ust=1779048535914471&usg=AOvVaw1W8cgxR5OKRQRkBlRrDWOU) |
| 25 | [▪ SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing Through Examination of Generation Methods, Content Characteristics, and Countermeasures (Chen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21457&sa=D&source=editors&ust=1779048535914523&usg=AOvVaw1XzRHqz2ok7IOcC_RghrkC) |
| 26 | [▪ Trustworthiness Calibration Framework for Phishing Email Detection Using Large Language Models (Ganiuly and Smaiyl, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04728&sa=D&source=editors&ust=1779048535914573&usg=AOvVaw3VpLHNfYZlwk_2QFJXDL-o) |
| 27 | [▪ Cross-Lingual Summarization as a Black-Box Watermark Removal Attack (Ganesan, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24789&sa=D&source=editors&ust=1779048535914618&usg=AOvVaw0sZbM9SSDUQHUIbyxIwy_Q) |
| 28 | [▪ MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs (Zollicoffer, Vu, and Bhattarai, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11741&sa=D&source=editors&ust=1779048535914666&usg=AOvVaw3Qx7XprEstPHGjyuvvjJ-d) |
| 29 | [▪ BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? (Jiang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18003&sa=D&source=editors&ust=1779048535914712&usg=AOvVaw21uyHdttItsaSnAUmn_jNP) |
| 30 | [▪ HauntAttack: When Attack Follows Reasoning as a Shadow (Ma et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.07031&sa=D&source=editors&ust=1779048535914758&usg=AOvVaw3uBd30Rap5StACE2554Tyr) |
| 31 | [▪ Elevating Cyber Threat Intelligence against Disinformation Campaigns with LLM-based Concept Extraction and the FakeCTI Dataset (Cotroneo, Natella, and Orbinato, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.03345&sa=D&source=editors&ust=1779048535914807&usg=AOvVaw37e3W7Gf9GO9A1-9ZnADDx) |
| 32 | [▪ Robust ML-based Detection of Conventional, LLM-Generated, and Adversarial Phishing Emails Using Advanced Text Preprocessing (Kulal et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11915&sa=D&source=editors&ust=1779048535914853&usg=AOvVaw3afyY5WoIPJqZT9vrZdKnR) |
| 33 | [▪ DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation (Ahn, Park, and Han, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10987&sa=D&source=editors&ust=1779048535914902&usg=AOvVaw1lP4go7JWxnJAWjoDR6pQ6) |
| 34 | [▪ The Security Threat of Compressed Projectors in Large Vision-Language Models (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00534&sa=D&source=editors&ust=1779048535914948&usg=AOvVaw3bZats3tYSyvIvX8cuklsW) |
| 35 | [▪ SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations (Liang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04398&sa=D&source=editors&ust=1779048535914991&usg=AOvVaw0ILahWIAYpVdiIAmk8i-9p) |
| 36 | [▪ Talking Like a Phisher: LLM-Based Attacks on Voice Phishing Classifiers (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16291&sa=D&source=editors&ust=1779048535915035&usg=AOvVaw19M3OVxhIKrsfAoD3DB6Zb) |
| 37 | [▪ DomainLynx: Leveraging Large Language Models for Enhanced Domain Squatting Detection (Chiba et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02095&sa=D&source=editors&ust=1779048535915094&usg=AOvVaw3-SDnAJ39RC6KWL_T3ZtsO) |
| 38 | [▪ We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (Spracklen et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.10279&sa=D&source=editors&ust=1779048535915147&usg=AOvVaw3BMA3RgLGPOugW6yI5x144) |

|     |
| --- |
| Hallucination Squatting and Phishing |

**>**

**<**

‍

#### Persistence

Covers:

- MITRE ATLAS Persistence

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Persistence

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents (Leong, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08442&sa=D&source=editors&ust=1779048535972467&usg=AOvVaw0zJwZZ-YjrhOIHrYRTWe-8) |
| 3 | [▪ Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense (Zha and Wang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02812&sa=D&source=editors&ust=1779048535972650&usg=AOvVaw1a1AXJn_kRH1JvvzxzJI-d) |
| 4 | [▪ Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents (Tallam, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14717&sa=D&source=editors&ust=1779048535972745&usg=AOvVaw3AYZGTY6UIs0ZeZQGvXW6J) |
| 5 | [▪ ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore (Zheng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20625&sa=D&source=editors&ust=1779048535972831&usg=AOvVaw14urUAeycJrZjFq6B4OmTa) |
| 6 | [▪ Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections (Yang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15654&sa=D&source=editors&ust=1779048535972917&usg=AOvVaw07OaWhwSTUcwHSxDn3mAxR) |
| 7 | [▪ MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (Srivastava and He, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16962&sa=D&source=editors&ust=1779048535972998&usg=AOvVaw2Za1-DTVDDlt5VGzaE60Pl) |

|     |
| --- |
| Persistence |

**>**

**<**

‍

#### Backdoor ML Model and Craft Adversarial Data

Covers:

- MITRE ATLAS ML Attack Staging

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Backdoor ML Model and Craft Adversarial Data

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning (Ma et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14587&sa=D&source=editors&ust=1779048535969569&usg=AOvVaw1ZLlf4BTuV5uUuxWiOx38X) |
| 3 | [▪ MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs (Wen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.15172&sa=D&source=editors&ust=1779048535969715&usg=AOvVaw2VZPaLhVuJM87CNuKRzgKO) |
| 4 | [▪ To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14291&sa=D&source=editors&ust=1779048535969798&usg=AOvVaw3fUfZtp9_r39BPdECEhAg5) |
| 5 | [▪ Backdoor Threats in Variational Quantum Circuits: Taxonomy, Attacks, and Defenses (Jiang and Chen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13796&sa=D&source=editors&ust=1779048535969872&usg=AOvVaw1Ig7J1UVJo3_eLqwBsRnUn) |
| 6 | [▪ Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks (Eggen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13214&sa=D&source=editors&ust=1779048535969944&usg=AOvVaw0fQX52dZmVJAIe9L99sS_Q) |
| 7 | [▪ LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters (Ahn et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13163&sa=D&source=editors&ust=1779048535970022&usg=AOvVaw3aS-wBWkoUH3tmHYstwTFR) |
| 8 | [▪ DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense (You et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13115&sa=D&source=editors&ust=1779048535970097&usg=AOvVaw13TiqjyJEcw9TdNAIvWGMU) |
| 9 | [▪ BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models (Rachapudi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12529&sa=D&source=editors&ust=1779048535970201&usg=AOvVaw22lxu5mGx9tqm7z370hgaQ) |
| 10 | [▪ FedSurrogate: Backdoor Defense in Federated Learning via Layer Criticality and Surrogate Replacement (Abacha et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11122&sa=D&source=editors&ust=1779048535970286&usg=AOvVaw0qEG2mXvlIyxusuw_WbPxS) |
| 11 | [▪ Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors (Yang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08278&sa=D&source=editors&ust=1779048535970372&usg=AOvVaw0f5QIHQLsJ17STPyrVkU8G) |
| 12 | [▪ BadDLM: Backdooring Diffusion Language Models with Diverse Targets (Zhai et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09397&sa=D&source=editors&ust=1779048535970454&usg=AOvVaw2iTFuGCg_R3a2aC_Np57cH) |
| 13 | [▪ Enhancing Adversarial Robustness in Network Intrusion Detection: A Layer-wise Adaptive Regularization Approach (Nasir et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08910&sa=D&source=editors&ust=1779048535970546&usg=AOvVaw3YnI1BhSLvq0-CVx7xq04T) |
| 14 | [▪ Hammer and Anvil: Toward a Theory of Backdoors in Federated Learning (Fenaux et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.08089&sa=D&source=editors&ust=1779048535970664&usg=AOvVaw0XvOQvs2NfYi6tyfqvDkbd) |
| 15 | [▪ Activation Differences Reveal Backdoors: A Comparison of SAE Architectures (Kumar, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07324&sa=D&source=editors&ust=1779048535970750&usg=AOvVaw14sOs9fAy4Dt_dGNKAHUKj) |
| 16 | [▪ Cross-Modal Backdoors in Multimodal Large Language Models (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07490&sa=D&source=editors&ust=1779048535970868&usg=AOvVaw3d7d_drMRfZHdmXNEGeMa4) |
| 17 | [▪ DeTrigger: A Gradient-Centric Approach to Backdoor Attack Mitigation in Federated Learning (Lee et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.12220&sa=D&source=editors&ust=1779048535970971&usg=AOvVaw387s9GCa7-CI7P4X6KkyYa) |
| 18 | [▪ Backdoor Mitigation in Object Detection via Adversarial Fine-Tuning (Dunnett et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05928&sa=D&source=editors&ust=1779048535971080&usg=AOvVaw1uWZJ0kAI5KSNcyV6uky7Z) |
| 19 | [▪ Stateful Agent Backdoor (Dai et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06158&sa=D&source=editors&ust=1779048535971174&usg=AOvVaw38BxTFFJU-TDV5S7ASyxeV) |
| 20 | [▪ Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02115&sa=D&source=editors&ust=1779048535971277&usg=AOvVaw37XnIFpJiHFV_fCiJsK6UW) |
| 21 | [▪ The Adversarial Discount - AI, Signal Correlation, and the Cybersecurity Arms Race (Bono, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04336&sa=D&source=editors&ust=1779048535971361&usg=AOvVaw1A2AkFE0teRAbbmUgciFuw) |
| 22 | [▪ Laundering AI Authority with Adversarial Examples (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04261&sa=D&source=editors&ust=1779048535971438&usg=AOvVaw1jMIYja7aXNU4gEnW1X-nM) |
| 23 | [▪ Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions (Choudhary et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04209&sa=D&source=editors&ust=1779048535971521&usg=AOvVaw1mwBJXfQPzvZ0IPu2CkUTg) |
| 24 | [▪ Detecting Adversarial Data via Provable Adversarial Noise Amplification (Mumcu and Yilmaz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02109&sa=D&source=editors&ust=1779048535971604&usg=AOvVaw3b997i7-gQJDwEKMs2pill) |
| 25 | [▪ Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration (Das et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01970&sa=D&source=editors&ust=1779048535971707&usg=AOvVaw1J9dR-oaJ_edg4ejohGHlO) |
| 26 | [▪ VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models (Liu and Lao, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01449&sa=D&source=editors&ust=1779048535971808&usg=AOvVaw3OUfp9z4HKabpxqBcU0PZ8) |
| 27 | [▪ Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget (Yang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01298&sa=D&source=editors&ust=1779048535971907&usg=AOvVaw0g5yWWT02lpwNJD9iQ1XJX) |
| 28 | [▪ Attention Is Where You Attack (Srivastava and Panda, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00236&sa=D&source=editors&ust=1779048535972028&usg=AOvVaw2Vbx0adLjAI3UIofPtaz1P) |
| 29 | [▪ One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness (Deguchi, Chousa, and Sakai, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27674&sa=D&source=editors&ust=1779048535972155&usg=AOvVaw03wprfSOV1eN9us4CE3cej) |
| 30 | [▪ Low Rank Adaptation for Adversarial Perturbation (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27487&sa=D&source=editors&ust=1779048535972245&usg=AOvVaw0QyB8vS58lxB3DulywIeUo) |
| 31 | [▪ Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis (Fernandez et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27414&sa=D&source=editors&ust=1779048535972333&usg=AOvVaw3BKM6SNAFTn-Za1wToxvEB) |
| 32 | [▪ Variational Autoencoder-Based Black-Box Adversarial Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01107&sa=D&source=editors&ust=1779048535972448&usg=AOvVaw20QXfKC-XX1uMPm34Xqww8) |
| 33 | [▪ Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights (Almalky et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22700&sa=D&source=editors&ust=1779048535972555&usg=AOvVaw0Lh_NLJVP-IsdLlAIuLH-m) |
| 34 | [▪ CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs (Nahian et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22681&sa=D&source=editors&ust=1779048535972658&usg=AOvVaw03Hrf8QfJ0oUdtrLvHT4SI) |
| 35 | [▪ MEASER: Malware embedding attacks on open-source LLMs (Tan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10486&sa=D&source=editors&ust=1779048535972758&usg=AOvVaw3TKM3AyPidcWKZUMclmk7F) |
| 36 | [▪ Prototype-Guided Robust Learning against Backdoor Attacks (Guo et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.08748&sa=D&source=editors&ust=1779048535972844&usg=AOvVaw3aKzGB9S63y401deUJEGzV) |
| 37 | [▪ Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24350&sa=D&source=editors&ust=1779048535972931&usg=AOvVaw31ffXFqukd4jWUcO2GmNst) |
| 38 | [▪ DETOUR: A Practical Backdoor Attack against Object Detection (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24599&sa=D&source=editors&ust=1779048535973027&usg=AOvVaw1wpsEPE0qyn8Xq5h4ix2e0) |
| 39 | [▪ Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing (Fan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24162&sa=D&source=editors&ust=1779048535973136&usg=AOvVaw3gbUbMuyFeF8yqX2MVerAS) |
| 40 | [▪ Toward Polymorphic Backdoor against Semantic Communication via Intensity-Based Poisoning (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23231&sa=D&source=editors&ust=1779048535973222&usg=AOvVaw3hHYBpQAUz23zWz2PlwFLf) |
| 41 | [▪ Adversarial Malware Generation in Linux ELF Binaries via Semantic-Preserving Transformations (Hrdonka and Jure\\v{c}ek, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22639&sa=D&source=editors&ust=1779048535973318&usg=AOvVaw2CB3NMZpIVYhuY30VICLr1) |
| 42 | [▪ Adversarial Co-Evolution of Malware and Detection Models: A Bilevel Optimization Perspective (Jure\\v{c}kov\\'a et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22569&sa=D&source=editors&ust=1779048535973402&usg=AOvVaw29sMMCWvO5RZc9o_4ICBuz) |
| 43 | [▪ Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers (Wei et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21700&sa=D&source=editors&ust=1779048535973491&usg=AOvVaw0KyujqqLfN3iVQlDq-j918) |
| 44 | [▪ PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20047&sa=D&source=editors&ust=1779048535973592&usg=AOvVaw2jkJDmili5XEBoYZPu-JqH) |
| 45 | [▪ Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients (Yuan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.06335&sa=D&source=editors&ust=1779048535973673&usg=AOvVaw3neCbgVePCsPjZlBK5Qpy0) |
| 46 | [▪ Penny Wise, Pixel Foolish: Bypassing Price Constraints in Multimodal Agents via Visual Adversarial Perturbations (Qian and Kang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16515&sa=D&source=editors&ust=1779048535973767&usg=AOvVaw2lDp63zg7lbt6pww0EIKF3) |
| 47 | [▪ Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15829&sa=D&source=editors&ust=1779048535973855&usg=AOvVaw1-QZFUPawxGNK7NScjx4Cm) |
| 48 | [▪ TopFeaRe: Locating Critical State of Adversarial Resilience for Graphs Regarding Topology-Feature Entanglement (Fan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15370&sa=D&source=editors&ust=1779048535973953&usg=AOvVaw1QPV53nDyZ39YuZUppx7GW) |
| 49 | [▪ Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training (Yi, Lai, and Li, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2408.14728&sa=D&source=editors&ust=1779048535974058&usg=AOvVaw0xcIXx9CVwfLJ8PuY7m_qN) |
| 50 | [▪ Robustness of Vision Foundation Models to Common Perturbations (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14973&sa=D&source=editors&ust=1779048535974151&usg=AOvVaw1-enx3zAH1Y97W7_A15ybx) |
| 51 | [▪ NeuroTrace: Inference Provenance-Based Detection of Adversarial Examples (Hmida et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14457&sa=D&source=editors&ust=1779048535974259&usg=AOvVaw3LXNz9TrWIXND0bZR4ipME) |
| 52 | [▪ INTARG: Informed Real-Time Adversarial Attack Generation for Time-Series Regression (Tokgoz et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11928&sa=D&source=editors&ust=1779048535974371&usg=AOvVaw23Qd7-p6d6U7tS_96lYKuX) |
| 53 | [▪ Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12446&sa=D&source=editors&ust=1779048535974514&usg=AOvVaw11IszJn98Gl9qE5khebymJ) |
| 54 | [▪ Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors (Yin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12359&sa=D&source=editors&ust=1779048535974624&usg=AOvVaw31PnnBLa4xK9KAP0JLLRMA) |
| 55 | [▪ Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version) (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03753&sa=D&source=editors&ust=1779048535974747&usg=AOvVaw21wVxx0pH2Xeh9qyB1xlSV) |
| 56 | [▪ Look Twice before You Leap: A Rational Framework for Localized Adversarial Anonymization (Duan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06713&sa=D&source=editors&ust=1779048535974848&usg=AOvVaw1abWgp-SspF0MHBEWqpuF6) |
| 57 | [▪ Property-Preserving Hashing for $\\ell\_1$-Distance Predicates: Applications to Countering Adversarial Input Attacks (Asghar, Zhang, and Kaafar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16355&sa=D&source=editors&ust=1779048535974939&usg=AOvVaw3LIn6QBQoVTxRSRdwGCnFh) |
| 58 | [▪ Defending against Backdoor Attacks via Module Switching (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.05902&sa=D&source=editors&ust=1779048535975038&usg=AOvVaw1m6BySXLPLnIaqeRHf5jGK) |
| 59 | [▪ Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models (Truong and Le, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10681&sa=D&source=editors&ust=1779048535975131&usg=AOvVaw2pTFVLHwV2HitJtvVvqH0k) |
| 60 | [▪ Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward (Guo et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09748&sa=D&source=editors&ust=1779048535975207&usg=AOvVaw3dNH_oqDcMzW9ZzUbGDa7N) |
| 61 | [▪ BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning (Tie et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09378&sa=D&source=editors&ust=1779048535975281&usg=AOvVaw0LytE_0d8KelWrhN2TKMwd) |
| 62 | [▪ Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor (Chang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09235&sa=D&source=editors&ust=1779048535975374&usg=AOvVaw0yn805VcyCc4tQLUplq2tp) |
| 63 | [▪ CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion (Jindal et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09101&sa=D&source=editors&ust=1779048535975448&usg=AOvVaw3qhQP1EWbRi3cQOJNyyEg8) |
| 64 | [▪ Follow My Eyes: Backdoor Attacks on VLM-based Scanpath Prediction (Romero et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08766&sa=D&source=editors&ust=1779048535975527&usg=AOvVaw3wVLQebiwh4n4akcgTqQvE) |
| 65 | [▪ BoBa: Boosting Backdoor Detection through Data Distribution Inference in Federated Learning (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2407.09658&sa=D&source=editors&ust=1779048535975616&usg=AOvVaw29C7THhi7C3HL0SbpS6u1U) |
| 66 | [▪ BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack (Khan, Miah, and Bi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15474&sa=D&source=editors&ust=1779048535975695&usg=AOvVaw26chKI09ZEyDheMxtTwDju) |
| 67 | [▪ CAAP: Capture-Aware Adversarial Patch Attacks on Palmprint Recognition Models (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06987&sa=D&source=editors&ust=1779048535975782&usg=AOvVaw1EKo4Ngdm07ANVhZWZ4rlh) |
| 68 | [▪ MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning (Zeng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06840&sa=D&source=editors&ust=1779048535975875&usg=AOvVaw1elMjk2mfUL-m2ndymTSjn) |
| 69 | [▪ SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems (Feng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06811&sa=D&source=editors&ust=1779048535975960&usg=AOvVaw1g_zuIKjZoxD-5kk7K8lPR) |
| 70 | [▪ Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models (Fortier et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01157&sa=D&source=editors&ust=1779048535976048&usg=AOvVaw3WrfvVC9P2zZbmqyOeZW7t) |
| 71 | [▪ Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05809&sa=D&source=editors&ust=1779048535976125&usg=AOvVaw1j1H2A2EWENBoPygrWkWMc) |
| 72 | [▪ Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use (Zhang and Pei, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05432&sa=D&source=editors&ust=1779048535976203&usg=AOvVaw2gyjf9m2comH0cfcKBm47G) |
| 73 | [▪ FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models (Deng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.12167&sa=D&source=editors&ust=1779048535976278&usg=AOvVaw3NbqVKZmx7o6O7f__6n71_) |
| 74 | [▪ Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs (Wheeler, Aryal, and Gupta, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03843&sa=D&source=editors&ust=1779048535976360&usg=AOvVaw3m8NLpQaJJHmbADE0Tsdmc) |
| 75 | [▪ ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning (Shui et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03595&sa=D&source=editors&ust=1779048535976438&usg=AOvVaw3_jXrFAGGDj6lR44i4OFAS) |
| 76 | [▪ S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.13891&sa=D&source=editors&ust=1779048535976529&usg=AOvVaw3VA9a_FM7wSLw0ltTAV6oA) |
| 77 | [▪ Street-Legal Physical-World Adversarial Rim for License Plates (Kalidasu and Ganapathy, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02457&sa=D&source=editors&ust=1779048535976608&usg=AOvVaw3EPcRH_FTpvRFBUG2EsX3-) |
| 78 | [▪ Backdoor Attacks on Decentralised Post-Training (Ersoy et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02372&sa=D&source=editors&ust=1779048535976693&usg=AOvVaw3LWxUGVipLUhyfUBeXATrJ) |
| 79 | [▪ Towards Physically Realizable Adversarial Attenuation Patch against SAR Object Detection (Zhang, Qin, and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00887&sa=D&source=editors&ust=1779048535976783&usg=AOvVaw2m4eIjrVIWu0xUCJox1riZ) |
| 80 | [▪ Spike-PTSD: A Bio-Plausible Adversarial Example Attack on Spiking Neural Networks via PTSD-Inspired Spike Scaling (Jin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01750&sa=D&source=editors&ust=1779048535976868&usg=AOvVaw2ZoKLbYdq7dWHxFLA6_K6r) |
| 81 | [▪ Diffusion-Guided Adversarial Perturbation Injection for Generalizable Defense Against Facial Manipulations (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01635&sa=D&source=editors&ust=1779048535976958&usg=AOvVaw1oXsH2VoNRjUxgaunGvqbg) |
| 82 | [▪ Adversarial Attenuation Patch Attack for SAR Object Detection (Zhang, Qin, and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00887&sa=D&source=editors&ust=1779048535977047&usg=AOvVaw3fxWrn8bEjwGUd2Gl311Vt) |
| 83 | [▪ Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses (Yu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29182&sa=D&source=editors&ust=1779048535977133&usg=AOvVaw1hJaEtJrgzAe4ZWTu9hjBD) |
| 84 | [▪ Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning (Herath, Zhao, and Bagchi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29328&sa=D&source=editors&ust=1779048535977222&usg=AOvVaw0E9oId0oBklgTnvuyiAa1h) |
| 85 | [▪ Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning (Sel et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29038&sa=D&source=editors&ust=1779048535977305&usg=AOvVaw2aGnfcjeLhBQlBuAKQWqYI) |
| 86 | [▪ SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28824&sa=D&source=editors&ust=1779048535977392&usg=AOvVaw1ZZLk1ej26JkHgLq90jee7) |
| 87 | [▪ Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs (Luo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20339&sa=D&source=editors&ust=1779048535977473&usg=AOvVaw2Rg1kbFlCw2lRYPJB6qOcZ) |
| 88 | [▪ FlowPure: Continuous Normalizing Flows for Adversarial Purification (Collaert et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.13280&sa=D&source=editors&ust=1779048535977565&usg=AOvVaw02m7NQOn_Yj3F0torAB4T7) |
| 89 | [▪ Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation (Zheng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16474&sa=D&source=editors&ust=1779048535977661&usg=AOvVaw2z8vG13RklBmQaVW1wgZ0y) |
| 90 | [▪ FL-PBM: Pre-Training Backdoor Mitigation for Federated Learning (Wehbi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28673&sa=D&source=editors&ust=1779048535977756&usg=AOvVaw3jF5TKv-jAFJgGE-Ryj_jg) |
| 91 | [▪ Mitigating Backdoor Attacks in Federated Learning Using PPA and MiniMax Game Theory (Wehbi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28652&sa=D&source=editors&ust=1779048535977848&usg=AOvVaw3q2mhNkno78cc5PDwrwXrk) |
| 92 | [▪ Detection of Adversarial Attacks in Robotic Perception (Sharawy, Nakshbandiand, and Grigorescu, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28594&sa=D&source=editors&ust=1779048535977950&usg=AOvVaw2AY1CK5tDBqk1h9LR-EJoD) |
| 93 | [▪ Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models (Yao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27522&sa=D&source=editors&ust=1779048535978045&usg=AOvVaw3TJlCp22owMcUm5tI8B_oH) |
| 94 | [▪ Attacking AI Accelerators by Leveraging Arithmetic Properties of Addition (Heidary and Joardar, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27439&sa=D&source=editors&ust=1779048535978130&usg=AOvVaw1drNtLp1knbm6zKvqe7iAU) |
| 95 | [▪ A Channel-Triggered Backdoor Attack on Wireless Semantic Image Reconstruction (Wan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.23866&sa=D&source=editors&ust=1779048535978215&usg=AOvVaw2e9umTtA_P_ADmPApebTVA) |
| 96 | [▪ On the Vulnerability of Deep Automatic Modulation Classifiers to Explainable Backdoor Threats (Salmi and Bogucka, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25310&sa=D&source=editors&ust=1779048535978301&usg=AOvVaw3QgCOQfIUkIyAPXANbG-Ad) |
| 97 | [▪ Physical Backdoor Attack Against Deep Learning-Based Modulation Classification (Salmi and Bogucka, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25304&sa=D&source=editors&ust=1779048535978390&usg=AOvVaw1vfrfUG62c87k7o42ynxG1) |
| 98 | [▪ IrisFP: Adversarial-Example-based Model Fingerprinting with Enhanced Uniqueness and Robustness (Geng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24996&sa=D&source=editors&ust=1779048535978479&usg=AOvVaw2wkRuwoXdzlUZuMJ2WN_OU) |
| 99 | [▪ Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach (Liao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.15778&sa=D&source=editors&ust=1779048535978568&usg=AOvVaw1MeJ9YH7K_gDnOTr44fBJ0) |
| 100 | [▪ Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs (Panfilov et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24511&sa=D&source=editors&ust=1779048535978651&usg=AOvVaw3XnJi9IoJwdjBWd4H-D002) |
| 101 | [▪ Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks (Pizarro, Narasimhan, and Fischer, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22590&sa=D&source=editors&ust=1779048535978731&usg=AOvVaw0ReQ8HPxXL50qx-cypLSLW) |
| 102 | [▪ Adversarial Vulnerabilities in Neural Operator Digital Twins: Gradient-Free Attacks on Nuclear Thermal-Hydraulic Surrogates (Roy et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22525&sa=D&source=editors&ust=1779048535978805&usg=AOvVaw2GEq0kLz3VwM2VTAHuez-m) |
| 103 | [▪ TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23117&sa=D&source=editors&ust=1779048535978871&usg=AOvVaw0uyChFlSehEVcTZky9YoIX) |
| 104 | [▪ AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents (Luo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23007&sa=D&source=editors&ust=1779048535978942&usg=AOvVaw0Wr98WDfyRLiUbySGetYal) |
| 105 | [▪ Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation (K\\"uchler et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.18323&sa=D&source=editors&ust=1779048535979018&usg=AOvVaw3neQ3Yg-GuMLzYJkGuVLgo) |
| 106 | [▪ Adversarial Attacks on Locally Private Graph Neural Networks (Kharagpur et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20746&sa=D&source=editors&ust=1779048535979091&usg=AOvVaw0-ez49dO1bTYFYxN2DCNhE) |
| 107 | [▪ Graph-Aware Text-Only Backdoor Poisoning for Text-Attributed Graphs (Luo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20339&sa=D&source=editors&ust=1779048535979161&usg=AOvVaw2B85p9KyDpl0NApWVh051w) |
| 108 | [▪ Trojan horse hunt in deep forecasting models: Insights from the European Space Agency competition (Kotowski et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20108&sa=D&source=editors&ust=1779048535979240&usg=AOvVaw1sEOfP1wDIqspkqx23c5-C) |
| 109 | [▪ Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19974&sa=D&source=editors&ust=1779048535979318&usg=AOvVaw2RM46726AOlDyx7EMS09qf) |
| 110 | [▪ Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks (Zhang, Wang, and Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18570&sa=D&source=editors&ust=1779048535979395&usg=AOvVaw0PyRgy-PWBzKpAqtsDM4wT) |
| 111 | [▪ MAED: Mathematical Activation Error Detection for Mitigating Physical Fault Attacks in DNN Inference (Ahmadi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18120&sa=D&source=editors&ust=1779048535979474&usg=AOvVaw3oCI18ZmhQPVsJYxMvFDDf) |
| 112 | [▪ STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18103&sa=D&source=editors&ust=1779048535979559&usg=AOvVaw27Q9b2yJeoU0MYblrgyDEN) |
| 113 | [▪ Adversarial attacks against Modern Vision-Language Models (Torre, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16960&sa=D&source=editors&ust=1779048535979660&usg=AOvVaw0nB1R1u-EWKFpwCcvTxWUA) |
| 114 | [▪ Coded Robust Aggregation for Distributed Learning under Byzantine Attacks (Li, Xiao, and Skoglund, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.01989&sa=D&source=editors&ust=1779048535979732&usg=AOvVaw2EmeEfU9CoJD0JCmOsDVxJ) |
| 115 | [▪ Poisoning the Pixels: Revisiting Backdoor Attacks on Semantic Segmentation (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16405&sa=D&source=editors&ust=1779048535979804&usg=AOvVaw2o0B6aHXC7Uv6PRpw4rwVf) |
| 116 | [▪ BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15692&sa=D&source=editors&ust=1779048535979872&usg=AOvVaw23M4BATxS7x0jPtwHPz1L9) |
| 117 | [▪ Inevitable Encounters: Backdoor Attacks Involving Lossy Compression (Li, Chen, and Chen, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13864&sa=D&source=editors&ust=1779048535979940&usg=AOvVaw12n7aXIpnnCPnHG4QrmQmL) |
| 118 | [▪ Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference (Li and Kim, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13461&sa=D&source=editors&ust=1779048535980018&usg=AOvVaw2FPv7X48hWdZp_dPycki7s) |
| 119 | [▪ Test-Time Attention Purification for Backdoored Large Vision Language Models (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12989&sa=D&source=editors&ust=1779048535980085&usg=AOvVaw2GPHexIoAJo1fH0EsUks9C) |
| 120 | [▪ RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection (Zhu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12582&sa=D&source=editors&ust=1779048535980153&usg=AOvVaw0LOhlKheJqHcAOSVqRW7Vw) |
| 121 | [▪ Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13028&sa=D&source=editors&ust=1779048535980221&usg=AOvVaw1QkyQaXyy0OwmrJmAx32PB) |
| 122 | [▪ Cascade: Composing Software-Hardware Attack Gadgets for Adversarial Threat Amplification in Compound AI Systems (Banerjee et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12023&sa=D&source=editors&ust=1779048535980292&usg=AOvVaw1H8YVViohEOnp8FXn7KT_G) |
| 123 | [▪ Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models (Ding et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11949&sa=D&source=editors&ust=1779048535980383&usg=AOvVaw3zV7S5PwJsW76j5ESjV3Ld) |
| 124 | [▪ Forging the Unforgeable: On the Feasibility of Counterfeit Watermarks in Backdoor-Based Dataset Ownership Verification (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.15450&sa=D&source=editors&ust=1779048535980464&usg=AOvVaw1iolhNuvrTV7n9Q5Q68H7d) |
| 125 | [▪ Backdoor Directions in Vision Transformers (Karayalcin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10806&sa=D&source=editors&ust=1779048535980533&usg=AOvVaw3Eytd4CyXDFnFembStGAcT) |
| 126 | [▪ Repurposing Backdoors for Good: Ephemeral Intrinsic Proofs for Verifiable Aggregation in Cross-silo Federated Learning (Qin, Yang, and Tang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10692&sa=D&source=editors&ust=1779048535980606&usg=AOvVaw3WkfrvXhdFMQ_TRNljGyku) |
| 127 | [▪ Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection (H{\\o}yheim et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10641&sa=D&source=editors&ust=1779048535980677&usg=AOvVaw2fUJTaLMLFe8WvZScUd7mu) |
| 128 | [▪ Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models (Raza et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10080&sa=D&source=editors&ust=1779048535980745&usg=AOvVaw0cE4yai8UiWol4DEl-ZvUW) |
| 129 | [▪ TASER: Task-Aware Spectral Energy Refine for Backdoor Suppression in UAV Swarms Decentralized Federated Learning (Huang and Yang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10075&sa=D&source=editors&ust=1779048535980817&usg=AOvVaw2FI9A4bASshTLn5DRej8ea) |
| 130 | [▪ Targeted Bit-Flip Attacks on LLM-Based Agents (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10042&sa=D&source=editors&ust=1779048535980881&usg=AOvVaw1qySGPqYcmz_rP8WGWt55m) |
| 131 | [▪ VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization Models (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.06782&sa=D&source=editors&ust=1779048535980948&usg=AOvVaw1dsD82ntRkuOEqD7buMz7l) |
| 132 | [▪ Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors (Abad et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09772&sa=D&source=editors&ust=1779048535981022&usg=AOvVaw2SB5MUuF6d477oUewkaIVk) |
| 133 | [▪ NetDiffuser: Deceiving DNN-Based Network Attack Detection Systems with Diffusion-Generated Adversarial Traffic (Kumar et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.08901&sa=D&source=editors&ust=1779048535981092&usg=AOvVaw1n2_hpqIpeitocxxrt3bgO) |
| 134 | [▪ DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models (Xu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10932&sa=D&source=editors&ust=1779048535981158&usg=AOvVaw1JYk9GWaIe-q9APOrM6NFq) |
| 135 | [▪ SFIBA: Spatial-based Full-target Invisible Backdoor Attacks (Yin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21052&sa=D&source=editors&ust=1779048535981225&usg=AOvVaw3hSgr-FFNVUcu4MXK2OJLU) |
| 136 | [▪ SlowBA: An efficiency backdoor attack towards VLM-based GUI agents (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.08316&sa=D&source=editors&ust=1779048535981291&usg=AOvVaw0k87LP8EkJcCJvuZUsZQoD) |
| 137 | [▪ Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07452&sa=D&source=editors&ust=1779048535981355&usg=AOvVaw0GC5hEcix-CzRvZBDhGg2V) |
| 138 | [▪ Structure-Aware Distributed Backdoor Attacks in Federated Learning (Jian et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03865&sa=D&source=editors&ust=1779048535981422&usg=AOvVaw1emgBGVYENPBAjNs907Iqc) |
| 139 | [▪ Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs (Pallakonda et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03371&sa=D&source=editors&ust=1779048535981488&usg=AOvVaw0RBReQCARSzWhMsDqkZQn-) |
| 140 | [▪ DSBA: Dynamic Stealthy Backdoor Attack with Collaborative Optimization in Self-Supervised Learning (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02849&sa=D&source=editors&ust=1779048535981579&usg=AOvVaw2eyJdsgMpY-urlH8dl0s7q) |
| 141 | [▪ TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models (Guo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02436&sa=D&source=editors&ust=1779048535981682&usg=AOvVaw0ZLrwmEdTD6YAglr4NxjAC) |
| 142 | [▪ Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection (D'Urso et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00217&sa=D&source=editors&ust=1779048535981760&usg=AOvVaw018w9DMiZFZFPWo2PXecii) |
| 143 | [▪ BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01019&sa=D&source=editors&ust=1779048535981831&usg=AOvVaw37YVzvOEtUranlIC5syVCf) |
| 144 | [▪ IU: Imperceptible Universal Backdoor Attack (Lin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00711&sa=D&source=editors&ust=1779048535981897&usg=AOvVaw39602jkg_MjwJ3qQopO0vb) |
| 145 | [▪ ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00516&sa=D&source=editors&ust=1779048535981966&usg=AOvVaw1M1ygKMrgQ9vp8kds_JzRL) |
| 146 | [▪ DropVLA: An Action-Level Backdoor Attack on Vision--Language--Action Models (Xu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10932&sa=D&source=editors&ust=1779048535982043&usg=AOvVaw12xvNIdwFPiF3vtTyxSTUh) |
| 147 | [▪ Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models (Wan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22246&sa=D&source=editors&ust=1779048535982111&usg=AOvVaw3dONYqUBrF-gJgZV432uo0) |
| 148 | [▪ Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity (Guo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.08747&sa=D&source=editors&ust=1779048535982177&usg=AOvVaw0NgNhWibxg3lMnkwj1SK_G) |
| 149 | [▪ Is the Trigger Essential? A Feature-Based Triggerless Backdoor Attack in Vertical Federated Learning (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20593&sa=D&source=editors&ust=1779048535982249&usg=AOvVaw00uBX5onUDLiy3J0B3z3L6) |
| 150 | [▪ When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks (Chen and Zhu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20193&sa=D&source=editors&ust=1779048535982328&usg=AOvVaw2euHdwvFuoRmzSx5XcKGtD) |
| 151 | [▪ Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation Systems (Shen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19539&sa=D&source=editors&ust=1779048535982429&usg=AOvVaw1N-Tcqn0KusRwFp9OBI2QA) |
| 152 | [▪ SafePickle: Robust and Generic ML Detection of Malicious Pickle-based ML Models (Ohayon, Gilkarov, and Dubin, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19818&sa=D&source=editors&ust=1779048535982504&usg=AOvVaw1nmP2vW8jasou2yacxJReA) |
| 153 | [▪ DCInject: Persistent Backdoor Attacks via Frequency Manipulation in Personal Federated Learning (Birhan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18489&sa=D&source=editors&ust=1779048535982577&usg=AOvVaw0uJ2XibfxmgrEpMEvI4ssq) |
| 154 | [▪ TrapFlow: Controllable Website Fingerprinting Defense via Dynamic Backdoor Learning (Liang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2412.11471&sa=D&source=editors&ust=1779048535982647&usg=AOvVaw0Wmms5sh0st6OCZMdOyj8J) |
| 155 | [▪ TFL: Targeted Bit-Flip Attack on Large Language Model (Guo, Chakrabarti, and Fan, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.17837&sa=D&source=editors&ust=1779048535982716&usg=AOvVaw2jSCpJYKj-dxVIA1lzQLq5) |
| 156 | [▪ Cert-SSBD: Certified Backdoor Defense with Sample-Specific Smoothing Noises (Qiao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21730&sa=D&source=editors&ust=1779048535982788&usg=AOvVaw3JB91NNa1aoamHEuu5-H9A) |
| 157 | [▪ Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective (Zhao, Hu, and Liu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15671&sa=D&source=editors&ust=1779048535982866&usg=AOvVaw0urOBtauasNN5M_QTbAbpc) |
| 158 | [▪ Weight space Detection of Backdoors in LoRA Adapters (Merenciano et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15195&sa=D&source=editors&ust=1779048535982938&usg=AOvVaw26ZJS1HPKDyRz0D_NvbRXV) |
| 159 | [▪ Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning (Foroughi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15161&sa=D&source=editors&ust=1779048535983011&usg=AOvVaw16-0ibhDc_D-HTDxyLETUh) |
| 160 | [▪ Backdooring Bias in Large Language Models (Das et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13427&sa=D&source=editors&ust=1779048535983081&usg=AOvVaw2zAM2_GhMY8qqQ1hBNsmvf) |
| 161 | [▪ Backdoor Attacks on Contrastive Continual Learning for IoT Systems (Tim and D, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13062&sa=D&source=editors&ust=1779048535983147&usg=AOvVaw0B2ZJpsBjJePGlwSwH1uC2) |
| 162 | [▪ PBP: Post-training Backdoor Purification for Malware Classifiers (Nguyen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2412.03441&sa=D&source=editors&ust=1779048535983214&usg=AOvVaw07HN1PmPWVK_r4J-S6I3Qv) |
| 163 | [▪ Transferable Backdoor Attacks for Code Models via Sharpness-Aware Adversarial Perturbation (Chang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11213&sa=D&source=editors&ust=1779048535983284&usg=AOvVaw0fPtK0lQH1tCBwznpUzcEl) |
| 164 | [▪ Kill it with FIRE: On Leveraging Latent Space Directions for Runtime Backdoor Mitigation in Deep Neural Networks (Ahlers et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10780&sa=D&source=editors&ust=1779048535983354&usg=AOvVaw3BfjMpY_Fa-SBLN64kM2MC) |
| 165 | [▪ Understanding and Enhancing Encoder-based Adversarial Transferability against Large Vision-Language Models (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09431&sa=D&source=editors&ust=1779048535983423&usg=AOvVaw3QPMMjUC8GL60u7b-TIKCR) |
| 166 | [▪ BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron (Miah, Vu, and Bi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07200&sa=D&source=editors&ust=1779048535983497&usg=AOvVaw1FrpL2nicim84omAMldwWn) |
| 167 | [▪ Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations (Miah and Bi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07197&sa=D&source=editors&ust=1779048535983569&usg=AOvVaw19ZZdyWXDJeo-0St5WPvnY) |
| 168 | [▪ Trojans in Artificial Intelligence (TrojAI) Final Report (Reese et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07152&sa=D&source=editors&ust=1779048535983634&usg=AOvVaw1hxJVCpwGwG6Kl9h6S4Ac2) |
| 169 | [▪ AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06771&sa=D&source=editors&ust=1779048535983706&usg=AOvVaw3spAy9gBtGpjr7GvsJZqSX) |
| 170 | [▪ Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06887&sa=D&source=editors&ust=1779048535983775&usg=AOvVaw3jI6JGjtITL3rI_izL07va) |
| 171 | [▪ Dependable Artificial Intelligence with Reliability and Security (DAIReS): A Unified Syndrome Decoding Approach for Hallucination and Backdoor Trigger Detection (Studies et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06532&sa=D&source=editors&ust=1779048535983849&usg=AOvVaw3IqSLYcugCRkUGvjUdhFFX) |
| 172 | [▪ BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05401&sa=D&source=editors&ust=1779048535983921&usg=AOvVaw3H7-SZbAiC8-m0fdwK7DAJ) |
| 173 | [▪ Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning (Rathbun et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05089&sa=D&source=editors&ust=1779048535984010&usg=AOvVaw2OiPbZBbCdF2Mg20aT9sKW) |
| 174 | [▪ Semantic-level Backdoor Attack against Text-to-Image Diffusion Models (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04898&sa=D&source=editors&ust=1779048535984087&usg=AOvVaw1sOJzhu3woHdJlmNWmtmDL) |
| 175 | [▪ Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation (Yang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04195&sa=D&source=editors&ust=1779048535984155&usg=AOvVaw2EW44E8vOzPS53EG4yrrPV) |
| 176 | [▪ Inference-Time Backdoors via Hidden Instructions in LLM Chat Templates (Fogel et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04653&sa=D&source=editors&ust=1779048535984221&usg=AOvVaw3mCvMCUwpX79mU0_egOepY) |
| 177 | [▪ Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks (Yu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03284&sa=D&source=editors&ust=1779048535984287&usg=AOvVaw10JXuMlV1KTiBwj9AsbofI) |
| 178 | [▪ The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers (Bullwinkel et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03085&sa=D&source=editors&ust=1779048535984363&usg=AOvVaw3s9Q3jFDBX22HsooIbOIbo) |
| 179 | [▪ DF-LoGiT: Data-Free Logic-Gated Backdoor Attacks in Vision Transformers (Shen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03040&sa=D&source=editors&ust=1779048535984450&usg=AOvVaw1jaIun07VmPJJG3Oe8eRu_) |
| 180 | [▪ Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection (Mukherjee et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00318&sa=D&source=editors&ust=1779048535984521&usg=AOvVaw3v-zE2Bo7Yi7wCgcQrBgLR) |
| 181 | [▪ HPE: Hallucinated Positive Entanglement for Backdoor Attacks in Federated Self-Supervised Learning (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02147&sa=D&source=editors&ust=1779048535984589&usg=AOvVaw2-BRGCaJwPwAoRBkkW5bXk) |
| 182 | [▪ Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01765&sa=D&source=editors&ust=1779048535984658&usg=AOvVaw2fzcac5wXPnW3bUgevGelj) |
| 183 | [▪ RPP: A Certified Poisoned-Sample Detection Framework for Backdoor Attacks under Dataset Imbalance (Lin et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00183&sa=D&source=editors&ust=1779048535984727&usg=AOvVaw2eWQe5dF-jf6Pn5KWYXP74) |
| 184 | [▪ Impact of Phonetics on Speaker Identity in Adversarial Voice Attack (Dar et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.15437&sa=D&source=editors&ust=1779048535984793&usg=AOvVaw0n-8ESVadvQkAljTTAAP3p) |
| 185 | [▪ Beauty and the Beast: Imperceptible Perturbations Against Diffusion-Based Face Swapping via Directional Attribute Editing (Huang and Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22744&sa=D&source=editors&ust=1779048535984862&usg=AOvVaw06fH4U0Sz4RbMCMfFws7GP) |
| 186 | [▪ An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems (Park and Kim, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22390&sa=D&source=editors&ust=1779048535984931&usg=AOvVaw0EJKpRhQgifDq0YDf4jFBn) |
| 187 | [▪ Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20002&sa=D&source=editors&ust=1779048535985021&usg=AOvVaw3dVNrOGlQuoASQDf_WUYK8) |
| 188 | [▪ Hardware-Triggered Backdoors (M\\"oller et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21902&sa=D&source=editors&ust=1779048535985097&usg=AOvVaw1idT224yNZIF1dc7h67dvG) |
| 189 | [▪ BadDet+: Robust Backdoor Attacks for Object Detection (Dunnett et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21066&sa=D&source=editors&ust=1779048535985164&usg=AOvVaw01T_6EGDFNT1lmBJ1ZiT0D) |
| 190 | [▪ SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning (Khan, Chandio, and Anwar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.22506&sa=D&source=editors&ust=1779048535985233&usg=AOvVaw3zebXb4ijHrrrGXQGXOvlA) |
| 191 | [▪ From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Online Test-Time Backdoor Defense (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19448&sa=D&source=editors&ust=1779048535985319&usg=AOvVaw3Qx5TgoGLSK8APZQkbGdKS) |
| 192 | [▪ TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion (Sreekumar et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17178&sa=D&source=editors&ust=1779048535985394&usg=AOvVaw2Ij42PwVAvAWCyiu3qsZbh) |
| 193 | [▪ Multi-Targeted Graph Backdoor Attack (Khan, Miah, and Bi, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15474&sa=D&source=editors&ust=1779048535985468&usg=AOvVaw1kP0RSn4-IwXRn8VgFEkJR) |
| 194 | [▪ SpooFL: Spoofing Federated Learning (Baglin, Zhu, and Hadfield, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15055&sa=D&source=editors&ust=1779048535985544&usg=AOvVaw2dH9yhqbNZlEPSeChdwbCW) |
| 195 | [▪ Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs (Lu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14340&sa=D&source=editors&ust=1779048535985619&usg=AOvVaw1JC5NTGsZV7YYCqtrGkMAe) |
| 196 | [▪ SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14323&sa=D&source=editors&ust=1779048535985695&usg=AOvVaw1PCe4uUPKlqfdWtSbTNbgC) |
| 197 | [▪ DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing (Hu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14302&sa=D&source=editors&ust=1779048535985772&usg=AOvVaw2l9g7zPVcG0fkXun8i5BxH) |
| 198 | [▪ SoK: On the Survivability of Backdoor Attacks on Unconstrained Face Recognition Systems (Roux et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.01607&sa=D&source=editors&ust=1779048535985847&usg=AOvVaw2YTzOlxQH98f_UiJJTRZM1) |
| 199 | [▪ SecureSplit: Mitigating Backdoor Attacks in Split Learning (Dou et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14054&sa=D&source=editors&ust=1779048535985919&usg=AOvVaw1JJwS9kXerAPCIkJc5ZuN6) |
| 200 | [▪ MirageNet:A Secure, Efficient, and Scalable On-Device Model Protection in Heterogeneous TEE and GPU System (Zheng, Cheng, and Ding, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13826&sa=D&source=editors&ust=1779048535986003&usg=AOvVaw0YnJBrxTr1otK2V9r0__vc) |
| 201 | [▪ DUAP: Dual-task Universal Adversarial Perturbations Against Voice Control Systems (Sun et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12786&sa=D&source=editors&ust=1779048535986090&usg=AOvVaw0oUu9gShEf97w8LnejOT5S) |
| 202 | [▪ Malware Classification using Diluted Convolutional Neural Network with Fast Gradient Sign Method (Anand et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09933&sa=D&source=editors&ust=1779048535986158&usg=AOvVaw0DspV5JFbbCYoGUQPRxNU9) |
| 203 | [▪ BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation (Wu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01595&sa=D&source=editors&ust=1779048535986228&usg=AOvVaw3VcjZJwgQ2xAbC9VCih3Jn) |
| 204 | [▪ STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio (Park et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08511&sa=D&source=editors&ust=1779048535986295&usg=AOvVaw0vTLvYKQflfqDOkSQN16X0) |
| 205 | [▪ Double Strike: Breaking Approximation-Based Side-Channel Countermeasures for DNNs (Casalino et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08698&sa=D&source=editors&ust=1779048535986362&usg=AOvVaw19GhQWwrWJGoTL02zXocXh) |
| 206 | [▪ ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08189&sa=D&source=editors&ust=1779048535986429&usg=AOvVaw1OnKdqK7wYhqd2BuAfIT1G) |
| 207 | [▪ BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts (Chan, Lintelo, and Picek, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.01811&sa=D&source=editors&ust=1779048535986495&usg=AOvVaw0zDBTOoAw-cFAJU7j1ye5I) |
| 208 | [▪ HogVul: Black-box Adversarial Code Generation Framework Against LM-based Vulnerability Detectors (Yang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05587&sa=D&source=editors&ust=1779048535986563&usg=AOvVaw3HyQZmKFe8Y-EKGWTFFL4j) |
| 209 | [▪ Unified Framework for Qualifying Security Boundary of PUFs Against Machine Learning Attacks (Fei et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04697&sa=D&source=editors&ust=1779048535986631&usg=AOvVaw0h03X_ziZNcHwW9Zhz8eea) |
| 210 | [▪ State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space (Guo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04266&sa=D&source=editors&ust=1779048535986700&usg=AOvVaw2ULRaxdpb7e42VZS0SaJSB) |
| 211 | [▪ Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models (Fu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04261&sa=D&source=editors&ust=1779048535986768&usg=AOvVaw1s6GTYjjLg7tM1nNHjHeoK) |
| 212 | [▪ Beyond Immediate Activation: Temporally Decoupled Backdoor Attacks on Time Series Forecasting (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04247&sa=D&source=editors&ust=1779048535986836&usg=AOvVaw06O4bVG8e3DVVO72GW7JnD) |
| 213 | [▪ Adversarial Contrastive Learning for LLM Quantization Attacks (Song et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02680&sa=D&source=editors&ust=1779048535986901&usg=AOvVaw3WK2BQC12rz6gYCqwwnI4H) |
| 214 | [▪ Non-omniscient backdoor injection with one poison sample: Proving the one-poison hypothesis for linear regression, linear classification, and 2-layer ReLU neural networks (Peinemann et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.05600&sa=D&source=editors&ust=1779048535986999&usg=AOvVaw1qTgWSVLdeaaRB01h6mSIl) |
| 215 | [▪ SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models (Xue, Zhang, and Xie, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14301&sa=D&source=editors&ust=1779048535987077&usg=AOvVaw2IowKMuJ1KgoEhCjl1a_JE) |
| 216 | [▪ SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs (Wu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06153&sa=D&source=editors&ust=1779048535987154&usg=AOvVaw1KA8tQDYLF0D3IWrfc3Paf) |
| 217 | [▪ Coward: Collision-based Watermark for Proactive Federated Backdoor Detection (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02115&sa=D&source=editors&ust=1779048535987229&usg=AOvVaw2GSPjI7OxZT9hnJDW7My0i) |
| 218 | [▪ From Chat Control to Robot Control: The Backdoors Left Open for the Sake of Safety (Akalin and Giaretta, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02205&sa=D&source=editors&ust=1779048535987304&usg=AOvVaw0jg6JVbn8XNdVKQ6xJzPUB) |
| 219 | [▪ IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection (Zhu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01239&sa=D&source=editors&ust=1779048535987380&usg=AOvVaw33xwOqJOXZ6uMzWOA8OJg3) |
| 220 | [▪ Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization (Guan, Jin, and Wang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01747&sa=D&source=editors&ust=1779048535987456&usg=AOvVaw1AIyzQOA16P3N1OUlsFER0) |
| 221 | [▪ NADD: Amplifying Noise for Effective Diffusion-based Adversarial Purification (Nguyen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01109&sa=D&source=editors&ust=1779048535987531&usg=AOvVaw3Z2FskhKOLU_i7z2msHHEg) |
| 222 | [▪ The Trojan in the Vocabulary: Stealthy Sabotage of LLM Composition (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00065&sa=D&source=editors&ust=1779048535987613&usg=AOvVaw211mtrGcfPHwHXaCytQgAd) |
| 223 | [▪ Rectifying Adversarial Examples Using Their Vulnerabilities (Morimoto, Morita, and Ono, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00270&sa=D&source=editors&ust=1779048535987679&usg=AOvVaw1Pm9Z2GQ1xJx9AZ6251NEE) |
| 224 | [▪ BadBlocks: Lightweight and Stealthy Backdoor Threat in Text-to-Image Diffusion Models (Pan et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.03221&sa=D&source=editors&ust=1779048535987753&usg=AOvVaw3hpObLdoeCUDpLVioLMrTO) |
| 225 | [▪ Breaking Audio Large Language Models by Attacking Only the Encoder: A Universal Targeted Latent-Space Audio Attack (Ziv, Lapid, and Sipper, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23881&sa=D&source=editors&ust=1779048535987844&usg=AOvVaw3vd28RryfI8ec7TXPoSFK6) |
| 226 | [▪ Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways (Frants and Agaian, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.24499&sa=D&source=editors&ust=1779048535987918&usg=AOvVaw1Uq8SUtR7yMRoygTLx41AH) |
| 227 | [▪ RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking (Liu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23307&sa=D&source=editors&ust=1779048535987993&usg=AOvVaw1Sv2SpOlxir1o2hDWAAAYv) |
| 228 | [▪ Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22046&sa=D&source=editors&ust=1779048535988061&usg=AOvVaw2eTHMmB_pExAzFVKtSY3Rr) |
| 229 | [▪ Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21681&sa=D&source=editors&ust=1779048535988130&usg=AOvVaw2bNiK2pwrdhCYbFHvi5_Iv) |
| 230 | [▪ LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors (Lan and Na\\"it-Abdesselam, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21404&sa=D&source=editors&ust=1779048535988205&usg=AOvVaw2eDxw0F3UAkP0xykrhuEfE) |
| 231 | [▪ WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.08602&sa=D&source=editors&ust=1779048535988277&usg=AOvVaw1nLMSNpA_t6Z2NFo3iwh9z) |
| 232 | [▪ CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21250&sa=D&source=editors&ust=1779048535988345&usg=AOvVaw0lhg598XJ3eqR7qICBXsRe) |
| 233 | [▪ Collision-based Watermark for Detecting Backdoor Manipulation in Federated Learning (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02115&sa=D&source=editors&ust=1779048535988412&usg=AOvVaw0V8i9rj7YPOYhc8oPzpA21) |
| 234 | [▪ ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19730&sa=D&source=editors&ust=1779048535988478&usg=AOvVaw3Zp8YYGW6zjU9UuBauY7TU) |
| 235 | [▪ One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models (Fang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05491&sa=D&source=editors&ust=1779048535988552&usg=AOvVaw1h6Yr8cyTBPFBMq_0k53Aj) |
| 236 | [▪ Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19297&sa=D&source=editors&ust=1779048535988619&usg=AOvVaw2_d0pD3Kmmb2pLzqxtm4Ws) |
| 237 | [▪ Adversarial Robustness of Vision in Open Foundation Models (Fox, Buchanan, and Papadopoulos, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17902&sa=D&source=editors&ust=1779048535988685&usg=AOvVaw1oor6BBVvfskRRHthwe-E6) |
| 238 | [▪ Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier? (Lu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15769&sa=D&source=editors&ust=1779048535988753&usg=AOvVaw13SSQ4VjZt7RyRGQ9i5Zji) |
| 239 | [▪ PHANTOM: Progressive High-fidelity Adversarial Network for Threat Object Modeling (Al-Karaki, Khan, and Athamneh, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15768&sa=D&source=editors&ust=1779048535988823&usg=AOvVaw0JQ75aSFkkiMt7F4a-ttXK) |
| 240 | [▪ Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs (Cui et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14741&sa=D&source=editors&ust=1779048535988888&usg=AOvVaw0LeuVE0aXIz87dWLb8R8wI) |
| 241 | [▪ CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World (Zhao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14158&sa=D&source=editors&ust=1779048535988958&usg=AOvVaw2_lNhSINenMcHolkdVMjWt) |
| 242 | [▪ Backdoors in DRL: Four Environments Focusing on In-distribution Triggers (Ashcraft et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17248&sa=D&source=editors&ust=1779048535989033&usg=AOvVaw29_SBpdlSwvRJTy64mXmju) |
| 243 | [▪ Trojan Cleansing with Neural Collapse (Gu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.12914&sa=D&source=editors&ust=1779048535989100&usg=AOvVaw3ii0XFBWXKCFVZXXSmxQiT) |
| 244 | [▪ We Can Always Catch You: Detecting Adversarial Patched Objects WITH or WITHOUT Signature (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2106.05261&sa=D&source=editors&ust=1779048535989186&usg=AOvVaw1aexT33qL5pETdVILO9uzV) |
| 245 | [▪ Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models (Chen, Wang, and Li, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13039&sa=D&source=editors&ust=1779048535989268&usg=AOvVaw1_jh49-6T-AeqeqHayaH3B) |
| 246 | [▪ Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks (Tang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13119&sa=D&source=editors&ust=1779048535989341&usg=AOvVaw0M7UjLPZEvskzqFURkhvQN) |
| 247 | [▪ Adversarial Attacks Against Deep Learning-Based Radio Frequency Fingerprint Identification (Ma et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12002&sa=D&source=editors&ust=1779048535989411&usg=AOvVaw3os2UqaFP61044A_Qva0OQ) |
| 248 | [▪ Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10600&sa=D&source=editors&ust=1779048535989497&usg=AOvVaw0cu2oWnbgzgWJ9T9UF-DpW) |
| 249 | [▪ Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs (Betley et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09742&sa=D&source=editors&ust=1779048535989567&usg=AOvVaw14KrYkdG2QnipjrWIxLIP9) |
| 250 | [▪ ByteShield: Adversarially Robust End-to-End Malware Detection through Byte Masking (Gibert and Many\\\`a, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09883&sa=D&source=editors&ust=1779048535989637&usg=AOvVaw2ozEqgEHbUUDw1ocbIvymd) |
| 251 | [▪ FBA$^2$D: Frequency-based Black-box Attack for AI-generated Image Detection (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09264&sa=D&source=editors&ust=1779048535989703&usg=AOvVaw2ky8PXroEyZtLTIRy-h9Bj) |
| 252 | [▪ AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing (Hong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07247&sa=D&source=editors&ust=1779048535989779&usg=AOvVaw1wz-4iI-jsKhj9Io_o508k) |
| 253 | [▪ Quantization Blindspots: How Model Compression Breaks Backdoor Defenses (Pandey and Ye, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06243&sa=D&source=editors&ust=1779048535989876&usg=AOvVaw1sxrNzL7rVRdNp29InwvLr) |
| 254 | [▪ Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models (Zhao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06899&sa=D&source=editors&ust=1779048535989958&usg=AOvVaw1g6Rc-cdGta7QiZwInSNv_) |
| 255 | [▪ Edge-Only Universal Adversarial Attacks in Distributed Learning (Rossolini et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.10500&sa=D&source=editors&ust=1779048535990033&usg=AOvVaw1N5PjfGfK_tVDeE0vRjWsg) |
| 256 | [▪ UltraClean: A Simple Framework to Train Robust Neural Networks against Backdoor Attacks (Zhao and Lao, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.10657&sa=D&source=editors&ust=1779048535990106&usg=AOvVaw3rP1wD6hsKla9o-JV-MsKJ) |
| 257 | [▪ SafeGenes: Evaluating the Adversarial Robustness of Genomic Foundation Models (Zhan, Barbour, and Moore, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00821&sa=D&source=editors&ust=1779048535990174&usg=AOvVaw0yQhgoFXqjDxYsKMvtDgo4) |
| 258 | [▪ Bones of Contention: Exploring Query-Efficient Attacks against Skeleton Recognition Systems (Cao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16843&sa=D&source=editors&ust=1779048535990248&usg=AOvVaw3_PsBIO7lDUn9IpEJmGovr) |
| 259 | [▪ Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas (Oe et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02062&sa=D&source=editors&ust=1779048535990332&usg=AOvVaw37CHYhPIu6n65ssE1bQCP3) |
| 260 | [▪ Concept-Guided Backdoor Attack on Vision Language Models (Shen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00713&sa=D&source=editors&ust=1779048535990402&usg=AOvVaw1bw2UhBjXB04Reh0DAHARR) |
| 261 | [▪ TrojanLoC: LLM-based Framework for RTL Trojan Localization (Xiao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00591&sa=D&source=editors&ust=1779048535990468&usg=AOvVaw00Y7aAxdwqn2zF9g5a8pSo) |
| 262 | [▪ CacheTrap: Injecting Trojans in LLMs without Leaving any Traces in Inputs or Weights (Nahian et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22681&sa=D&source=editors&ust=1779048535990537&usg=AOvVaw3_d8K3-jboGXZ7RmzG21Ja) |
| 263 | [▪ Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22415&sa=D&source=editors&ust=1779048535990614&usg=AOvVaw1kF0MjPI4iNbhyEfan_SLn) |
| 264 | [▪ On the Effectiveness of Adversarial Training on Malware Classifiers (Bostani et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.18218&sa=D&source=editors&ust=1779048535990683&usg=AOvVaw0Z2VOWF7KjCh7N0VNCgATO) |
| 265 | [▪ Special-Character Adversarial Attacks on Open-Source Language Model (Sarabamoun, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.14070&sa=D&source=editors&ust=1779048535990747&usg=AOvVaw0wriGcssfbBq2KQj0pPM0_) |
| 266 | [▪ Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI (Hoang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21291&sa=D&source=editors&ust=1779048535990815&usg=AOvVaw2E4bjny-N7UZjOY4veRp-P) |
| 267 | [▪ CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion (Xia et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21180&sa=D&source=editors&ust=1779048535990908&usg=AOvVaw1LU0XoA6WE_ZenDjHJByPR) |
| 268 | [▪ BackFed: An Efficient & Standardized Benchmark Suite for Backdoor Attacks in Federated Learning (Dao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.04903&sa=D&source=editors&ust=1779048535990994&usg=AOvVaw0PJ1_8NGAU-rvLix40704S) |
| 269 | [▪ Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains (Sanna, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19874&sa=D&source=editors&ust=1779048535991065&usg=AOvVaw0LxUVwxb0PQj0_re6C0-kQ) |
| 270 | [▪ IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.09456&sa=D&source=editors&ust=1779048535991131&usg=AOvVaw1idFA_TM6c2iFFpqS2_T18) |
| 271 | [▪ DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs (Guo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18617&sa=D&source=editors&ust=1779048535991196&usg=AOvVaw1FyNsjVw4ky0HbWBPk2kBx) |
| 272 | [▪ A Novel and Practical Universal Adversarial Perturbations against Deep Reinforcement Learning based Intrusion Detection Systems (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18223&sa=D&source=editors&ust=1779048535991266&usg=AOvVaw3Eb_nyg98U5VOsJ1coPkac) |
| 273 | [▪ Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models (Luo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17982&sa=D&source=editors&ust=1779048535991335&usg=AOvVaw1TNUbH2D7eti5soqjYieR_) |
| 274 | [▪ Evaluating Adversarial Vulnerabilities in Modern Large Language Models (Perel, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17666&sa=D&source=editors&ust=1779048535991401&usg=AOvVaw2ei7iyPyITBs2qkI3nQQeq) |
| 275 | [▪ Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning (Obioma, Sun, and Mustafa, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.10297&sa=D&source=editors&ust=1779048535991469&usg=AOvVaw2ziwOvmf1uOsYs2tee8X4K) |
| 276 | [▪ Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17194&sa=D&source=editors&ust=1779048535991542&usg=AOvVaw2tQTZPcPlDc4AahW3kTb_Y) |
| 277 | [▪ AutoBackdoor: Automating Backdoor Attacks via LLM Agents (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16709&sa=D&source=editors&ust=1779048535991610&usg=AOvVaw3RYMfzFY3mukRns4edbllV) |
| 278 | [▪ Transferable Dual-Domain Feature Importance Attack against AI-Generated Image Detector (Zhu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15571&sa=D&source=editors&ust=1779048535991679&usg=AOvVaw0WEv9H2m5WP3T-UX4EQyt9) |
| 279 | [▪ Attacking Autonomous Driving Agents with Adversarial Machine Learning: A Holistic Evaluation with the CARLA Leaderboard (Wong et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14876&sa=D&source=editors&ust=1779048535991748&usg=AOvVaw3gKHM_pgFvph_OOYd_Onhd) |
| 280 | [▪ High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine Attackers (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2307.13352&sa=D&source=editors&ust=1779048535991816&usg=AOvVaw1ZPgU6j8D2Cq2zBfAhqlkl) |
| 281 | [▪ BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting (Xiang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.04189&sa=D&source=editors&ust=1779048535991895&usg=AOvVaw2LBxnOGAURFdfBhOsuOQBy) |
| 282 | [▪ TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.09562&sa=D&source=editors&ust=1779048535991982&usg=AOvVaw3_XJpU3_YGShG0qrvUOQ2f) |
| 283 | [▪ Watch Out for the Lifespan: Evaluating Backdoor Attacks Against Federated Model Adaptation (Vuillod, Moellic, and Dutertre, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14406&sa=D&source=editors&ust=1779048535992071&usg=AOvVaw1XOdu-v7ZSaSzJNWxwrHsG) |
| 284 | [▪ Steganographic Backdoor Attacks in NLP: Ultra-Low Poisoning and Defense Evasion (Xue et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14301&sa=D&source=editors&ust=1779048535992150&usg=AOvVaw0Dy-pI8LUzRcG1nJ6MjbyO) |
| 285 | [▪ Dynamic Black-box Backdoor Attacks on IoT Sensory Data (Chathoth and Lee, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14074&sa=D&source=editors&ust=1779048535992220&usg=AOvVaw2bO520QXv_Ee8VO8aSMkHH) |
| 286 | [▪ Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks (Jin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13789&sa=D&source=editors&ust=1779048535992288&usg=AOvVaw3eWhuVVKvcoNKemlqMX4r2) |
| 287 | [▪ DiffProtect: Generate Adversarial Examples with Diffusion Models for Facial Privacy Protection (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2305.13625&sa=D&source=editors&ust=1779048535992357&usg=AOvVaw1mcnSsPEbdIWm_Ynf8HcHO) |
| 288 | [▪ NeuroStrike: Neuron-Level Attacks on Aligned LLMs (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11864&sa=D&source=editors&ust=1779048535992451&usg=AOvVaw0mQrrNnSKROh886hvoqNrt) |
| 289 | [▪ Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks (Sun et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.04094&sa=D&source=editors&ust=1779048535992523&usg=AOvVaw3Sg0ZTE1rsB1pi9KfCReoy) |
| 290 | [▪ Backdooring CLIP through Concept Confusion (Hu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.09095&sa=D&source=editors&ust=1779048535992590&usg=AOvVaw12VQlbfhcL7rhq1QH6XD1y) |
| 291 | [▪ The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models (Tan, Huang, and Li, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12414&sa=D&source=editors&ust=1779048535992664&usg=AOvVaw003XQgidN7adcH9xoyF4qK) |
| 292 | [▪ Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12265&sa=D&source=editors&ust=1779048535992734&usg=AOvVaw1kfoXZZbz-d4MZzyYwAV-m) |
| 293 | [▪ MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purification for Inference-Time Attacks (Karami et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11625&sa=D&source=editors&ust=1779048535992806&usg=AOvVaw3xbRvPclj2feCu-O1T2Onf) |
| 294 | [▪ Enhancing All-to-X Backdoor Attacks with Optimized Target Class Mapping (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13356&sa=D&source=editors&ust=1779048535992873&usg=AOvVaw3pm8lkXfohMYMaOA-ZPqDv) |
| 295 | [▪ SoK: The Last Line of Defense: On Backdoor Defense Evaluation (Abad et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13143&sa=D&source=editors&ust=1779048535992940&usg=AOvVaw0JlyKW3MvIwtRy22Ou2l8h) |
| 296 | [▪ Efficient Adversarial Malware Defense via Trust-Based Raw Override and Confidence-Adaptive Bit-Depth Reduction (Chaudhary and Doppalpudi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12827&sa=D&source=editors&ust=1779048535993037&usg=AOvVaw2Ie5Zqdg0Bdwg--hPFnC_Z) |
| 297 | [▪ AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12149&sa=D&source=editors&ust=1779048535993125&usg=AOvVaw2sevz-3pNiDx5XlnAD79kA) |
| 298 | [▪ BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning (Wang and Zhao, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12046&sa=D&source=editors&ust=1779048535993196&usg=AOvVaw1Chu4rZPfqJDkLmRS7Z0Fu) |
| 299 | [▪ Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10712&sa=D&source=editors&ust=1779048535993273&usg=AOvVaw2NHf3KOqztd_xhgMxG-z8w) |
| 300 | [▪ Transferable Hypergraph Attack via Injecting Nodes into Pivotal Hyperedges (He et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10698&sa=D&source=editors&ust=1779048535993344&usg=AOvVaw314EtIaN_qFAYij19bjMso) |
| 301 | [▪ Are Neural Networks Collision Resistant? (Benedetti et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.20262&sa=D&source=editors&ust=1779048535993427&usg=AOvVaw0_QXfuC_G3InwshVqheKQ7) |
| 302 | [▪ Trapped by Their Own Light: Deployable and Stealth Retroreflective Patch Attacks on Traffic Sign Recognition Systems (Tsuruoka et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10050&sa=D&source=editors&ust=1779048535993515&usg=AOvVaw16hzcfAUJVSmvfpHFV6t4A) |
| 303 | [▪ Improving Adversarial Transferability with Neighbourhood Gradient Information (Guo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.05745&sa=D&source=editors&ust=1779048535993592&usg=AOvVaw1Tgn6hDVOZuoNpuNZY7Qcp) |
| 304 | [▪ Robust Backdoor Removal by Reconstructing Trigger-Activated Changes in Latent Representation (Iwahana et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08944&sa=D&source=editors&ust=1779048535993670&usg=AOvVaw3K1vEN0lG62j8Z8zNnxFzO) |
| 305 | [▪ Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Backdoor Attacks in Federated Learning (Wang, Shen, and Lam, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09252&sa=D&source=editors&ust=1779048535993750&usg=AOvVaw27zGjj0ZgHUS-4jWRbh8kL) |
| 306 | [▪ Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07210&sa=D&source=editors&ust=1779048535993829&usg=AOvVaw1vKSDIHpvUoAJeP-pt58bp) |
| 307 | [▪ E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07099&sa=D&source=editors&ust=1779048535993906&usg=AOvVaw0GF6mXNo8sFkFMEdriscet) |
| 308 | [▪ From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge (Lu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07049&sa=D&source=editors&ust=1779048535993982&usg=AOvVaw2IHvxSTt1U9yB-pseCoLXW) |
| 309 | [▪ CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding (Tajalli, Koffas, and Picek, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06072&sa=D&source=editors&ust=1779048535994067&usg=AOvVaw1WKN82i-FNeZE0ROzgZS0Z) |
| 310 | [▪ Feature compression is the root cause of adversarial fragility in neural network classifiers (Gao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.16200&sa=D&source=editors&ust=1779048535994151&usg=AOvVaw1sxY6W2m_v79flOX7kW_3M) |
| 311 | [▪ Lorica: A Synergistic Fine-Tuning Framework for Advancing Personalized Adversarial Robustness (Qi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.05402&sa=D&source=editors&ust=1779048535994218&usg=AOvVaw3ZDnKr-VZIomCLOUA3m_Ll) |
| 312 | [▪ ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training (Yao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00446&sa=D&source=editors&ust=1779048535994283&usg=AOvVaw2q5N5PtRXAyuHHZ1wpNVOa) |
| 313 | [▪ ShadowLogic: Backdoors in Any Whitebox LLM (Schulz, Kawasaki, and Ring, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00664&sa=D&source=editors&ust=1779048535994348&usg=AOvVaw1V4c-1eJwFWADpgQwyQsyc) |
| 314 | [▪ Rethinking Robust Adversarial Concept Erasure in Diffusion Models (Yin, Tian, and Zhang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27285&sa=D&source=editors&ust=1779048535994417&usg=AOvVaw0DiXhnifpuh1l2vukg-beC) |
| 315 | [▪ GSE: Group-wise Sparse and Explainable Adversarial Attacks (Sadiku, Wagner, and Pokutta, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2311.17434&sa=D&source=editors&ust=1779048535994486&usg=AOvVaw2GkB_QVjIJshpN9A334ebl) |
| 316 | [▪ SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26420&sa=D&source=editors&ust=1779048535994555&usg=AOvVaw1wolXqSxYl4KkVwk2PbKhQ) |
| 317 | [▪ Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging (Latibari et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24976&sa=D&source=editors&ust=1779048535994624&usg=AOvVaw3TFefJif2N4Gu7Nhg8WcEp) |
| 318 | [▪ Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11173&sa=D&source=editors&ust=1779048535994696&usg=AOvVaw1yYPi1UwCRhEV8g3WXoV1j) |
| 319 | [▪ ME: Trigger Element Combination Backdoor Attack on Copyright Infringement (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10776&sa=D&source=editors&ust=1779048535994765&usg=AOvVaw3VSI04AwfvJW-HT-txr-K1) |
| 320 | [▪ Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents (Bazinska et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22620&sa=D&source=editors&ust=1779048535994833&usg=AOvVaw3U_Xf2CjJgDXVwwh0yChTc) |
| 321 | [▪ Cross-Paradigm Graph Backdoor Attacks with Promptable Subgraph Triggers (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22555&sa=D&source=editors&ust=1779048535994899&usg=AOvVaw2jWSQOfshzn_BVNVsQ715W) |
| 322 | [▪ Boosting Adversarial Transferability with Spatial Adversarial Alignment (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.01015&sa=D&source=editors&ust=1779048535994966&usg=AOvVaw3malh7LRsVZ6zSo3PFaj6O) |
| 323 | [▪ FPT-Noise: Dynamic Scene-Aware Counterattack for Test-Time Adversarial Defense in Vision-Language Models (Deng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20856&sa=D&source=editors&ust=1779048535995054&usg=AOvVaw3qr7ouhBKCCwUID6E1iAkV) |
| 324 | [▪ NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry (Gilkarov and Dubin, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20367&sa=D&source=editors&ust=1779048535995126&usg=AOvVaw21sStmRwFhQMHAQdbjQPE2) |
| 325 | [▪ Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection (Yi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19574&sa=D&source=editors&ust=1779048535995193&usg=AOvVaw0NrF7aqpaCYEZy6etTacxQ) |
| 326 | [▪ Exploring the Effect of DNN Depth on Adversarial Attacks in Network Intrusion Detection Systems (ElShehaby and Matrawy, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19761&sa=D&source=editors&ust=1779048535995262&usg=AOvVaw2EY4QVtuIR9mnm81XCD25y) |
| 327 | [▪ Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation (Muri et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18541&sa=D&source=editors&ust=1779048535995330&usg=AOvVaw3vogOutL9oog67CtzZbvAm) |
| 328 | [▪ Enhancing Adversarial Transferability with Adversarial Weight Tuning (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09469&sa=D&source=editors&ust=1779048535995397&usg=AOvVaw34bC17wAu9-XWT16rRbWy5) |
| 329 | [▪ Colliding with Adversaries at ECML-PKDD 2025 Adversarial Attack Competition 1st Prize Solution (Stefanopoulos and Voskou, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16440&sa=D&source=editors&ust=1779048535995467&usg=AOvVaw0b-VFuhdzHYYMWQ-nORVTl) |
| 330 | [▪ Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models (Hossain et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17098&sa=D&source=editors&ust=1779048535995536&usg=AOvVaw2R2N0Idx0JHma03txgBgFD) |
| 331 | [▪ UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks (Phute et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16923&sa=D&source=editors&ust=1779048535995609&usg=AOvVaw2etRZkG527r06FWxTJ3Lz-) |
| 332 | [▪ A Versatile Framework for Designing Group-Sparse Adversarial Attacks (Heshmati et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16637&sa=D&source=editors&ust=1779048535995678&usg=AOvVaw3wdPMO4-B8mQa78tJpXIeR) |
| 333 | [▪ Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16581&sa=D&source=editors&ust=1779048535995745&usg=AOvVaw1wQAFs1VL28ob4zE6XSobP) |
| 334 | [▪ PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models (Seddik et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15106&sa=D&source=editors&ust=1779048535995814&usg=AOvVaw2b1xqcjh9bR6g_l_ACurz6) |
| 335 | [▪ Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models (Xue et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14470&sa=D&source=editors&ust=1779048535995882&usg=AOvVaw02IVRr7ni1dbKu8Z5UQc3u) |
| 336 | [▪ An Information Asymmetry Game for Trigger-based DNN Model Watermarking (Huang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14218&sa=D&source=editors&ust=1779048535995950&usg=AOvVaw2kqBbTP44gs47pbkAP1sSI) |
| 337 | [▪ Improving Transferability of Adversarial Examples via Bayesian Attacks (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2307.11334&sa=D&source=editors&ust=1779048535996030&usg=AOvVaw2UIOUUJkBD-yBq4hm2DmMI) |
| 338 | [▪ Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13462&sa=D&source=editors&ust=1779048535996103&usg=AOvVaw3rs0ADJEOQUNU9adXrYSsK) |
| 339 | [▪ Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning (Song et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13322&sa=D&source=editors&ust=1779048535996172&usg=AOvVaw3opvr72yPzEwgGIkNE4nNn) |
| 340 | [▪ How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies (Kalra et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.03698&sa=D&source=editors&ust=1779048535996242&usg=AOvVaw29ofPbT5YzUcD0D-1-DYXf) |
| 341 | [▪ Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.16950&sa=D&source=editors&ust=1779048535996335&usg=AOvVaw33pqHoP9Sm6hOsbVuZCxEn) |
| 342 | [▪ DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent (Zhu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.12575&sa=D&source=editors&ust=1779048535996426&usg=AOvVaw1BOTpGuBzCv5TcIxrygFnq) |
| 343 | [▪ Adversarial Attacks on Downstream Weather Forecasting Models: Application to Tropical Cyclone Trajectory Prediction (Deng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10140&sa=D&source=editors&ust=1779048535996502&usg=AOvVaw3awQI0P9crBmkGGPTYMWx8) |
| 344 | [▪ Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems (Zhu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11246&sa=D&source=editors&ust=1779048535996572&usg=AOvVaw3Osx7k2o_Us7XGQlT3m13E) |
| 345 | [▪ TabVLA: Targeted Backdoor Attacks on Vision-Language-Action Models (Xu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10932&sa=D&source=editors&ust=1779048535996661&usg=AOvVaw09hEr79c-45zSTp-EvIFxH) |
| 346 | [▪ SASER: Stego attacks on open-source LLMs (Tan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10486&sa=D&source=editors&ust=1779048535996735&usg=AOvVaw2mu-wvFxVZ2d0QhZIY-o8J) |
| 347 | [▪ Rounding-Guided Backdoor Injection in Deep Learning Model Quantization (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09647&sa=D&source=editors&ust=1779048535996811&usg=AOvVaw08qDSxOmkV10h34ds3oCYK) |
| 348 | [▪ Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects (Zhou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09269&sa=D&source=editors&ust=1779048535996886&usg=AOvVaw1pFbha2JRDBxZWvWdvG6qR) |
| 349 | [▪ GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis (Dutta et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09260&sa=D&source=editors&ust=1779048535996961&usg=AOvVaw32iph8dOj-E7AwQBjaLf9F) |
| 350 | [▪ Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning (Gloaguen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16567&sa=D&source=editors&ust=1779048535997060&usg=AOvVaw2vM01oxIpBlZ8LHWxAOvvg) |
| 351 | [▪ Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses (Pawlak et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08016&sa=D&source=editors&ust=1779048535997136&usg=AOvVaw32ApO3rGUrNds44D1gDR8x) |
| 352 | [▪ Fewer Weights, More Problems: A Practical Attack on LLM Pruning (Egashira et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07985&sa=D&source=editors&ust=1779048535997209&usg=AOvVaw1OTf9KDzG5cyNkaeRubvhT) |
| 353 | [▪ Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs (Hu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07697&sa=D&source=editors&ust=1779048535997294&usg=AOvVaw0TGMSeXsXOtwm0TG0iqIqx) |
| 354 | [▪ Attacking the Spike: On the Transferability and Security of Spiking Neural Networks to Adversarial Examples (Xu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2209.03358&sa=D&source=editors&ust=1779048535997363&usg=AOvVaw1UCmm385oJSLV9MF76pR71) |
| 355 | [▪ Unsupervised Backdoor Detection and Mitigation for Spiking Neural Networks (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06629&sa=D&source=editors&ust=1779048535997441&usg=AOvVaw3coBhR4Ni1NpH0aG8zSa_P) |
| 356 | [▪ A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.19947&sa=D&source=editors&ust=1779048535997538&usg=AOvVaw1WetHrppt8est7xNq2gzSP) |
| 357 | [▪ From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs (Shen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05169&sa=D&source=editors&ust=1779048535997611&usg=AOvVaw1GKYhE5u2Naq4zrABQhZ1z) |
| 358 | [▪ Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies (Kalra et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.03698&sa=D&source=editors&ust=1779048535997679&usg=AOvVaw3YfRZaHL-sctc8lDElNr8U) |
| 359 | [▪ Cooperative Decentralized Backdoor Attacks on Vertical Federated Learning (Lee et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.09320&sa=D&source=editors&ust=1779048535997745&usg=AOvVaw3je4qFN_PqTJ9CX6Te-Xyb) |
| 360 | [▪ Backdoors in Code Summarizers: How Bad Is It? (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.01825&sa=D&source=editors&ust=1779048535997810&usg=AOvVaw1eD2CZ64Hh7b7_VXbVLY1x) |
| 361 | [▪ Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs (Rezkellah and Dakhmouche, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03567&sa=D&source=editors&ust=1779048535997877&usg=AOvVaw2AcngiecySiCQ_vTWJw4l5) |
| 362 | [▪ Adversarial training with restricted data manipulation (Benfield et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03254&sa=D&source=editors&ust=1779048535997941&usg=AOvVaw0d2zf2FzBRobRoJgbHq4aT) |
| 363 | [▪ NatGVD: Natural Adversarial Example Attack towards Graph-based Vulnerability Detection (Rath et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04987&sa=D&source=editors&ust=1779048535998015&usg=AOvVaw1pOUrnt-ACRcnGEQALZiin) |
| 364 | [▪ P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04503&sa=D&source=editors&ust=1779048535998081&usg=AOvVaw1NpnbPuq-mnIwyUB8p6thL) |
| 365 | [▪ Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03705&sa=D&source=editors&ust=1779048535998146&usg=AOvVaw17beysjkuvCJxBmwmgUxj0) |
| 366 | [▪ Attack logics, not outputs: Towards efficient robustification of deep neural networks by falsifying concept-based properties (Dankworth and Schwalbe, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03320&sa=D&source=editors&ust=1779048535998216&usg=AOvVaw0f9GDU5_jWdlDm2NnMVGv0) |
| 367 | [▪ Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders (Han et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.03508&sa=D&source=editors&ust=1779048535998284&usg=AOvVaw1UeJ39pS2cQUYRYnhauHCa) |
| 368 | [▪ A Statistical Method for Attack-Agnostic Adversarial Attack Detection with Compressive Sensing Comparison (Wimalasuriya and Tragoudas, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02707&sa=D&source=editors&ust=1779048535998353&usg=AOvVaw3kwMXRgJPCEJmo0oap6bYZ) |
| 369 | [▪ Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07670&sa=D&source=editors&ust=1779048535998419&usg=AOvVaw2X3kjjquOSHhxl9AbiwSBo) |
| 370 | [▪ Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems (Su et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02158&sa=D&source=editors&ust=1779048535998489&usg=AOvVaw1acdvP2oMzQ1B8OTH_lxIN) |
| 371 | [▪ Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01699&sa=D&source=editors&ust=1779048535998556&usg=AOvVaw0FSB99tNpfNNk46q-BX5YE) |
| 372 | [▪ Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks (Nasr et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01676&sa=D&source=editors&ust=1779048535998642&usg=AOvVaw3huQfDbUfncO_wSvMEA1Uj) |
| 373 | [▪ Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation (Chaudhari et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20485&sa=D&source=editors&ust=1779048535998711&usg=AOvVaw3WH4GURALv_uxZ-NT2n4f_) |
| 374 | [▪ A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions (Yang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.02344&sa=D&source=editors&ust=1779048535998785&usg=AOvVaw3pYVERErl5kEB1Y_4gFsjG) |
| 375 | [▪ Backdoor Attacks Against Speech Language Models (Fortier et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01157&sa=D&source=editors&ust=1779048535998852&usg=AOvVaw23AmAkuV6Wjb4hv5jqHydk) |
| 376 | [▪ Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness (Takahashi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00517&sa=D&source=editors&ust=1779048535998921&usg=AOvVaw2UShCQuT2UfT1zFKI_tFo_) |
| 377 | [▪ Has the Two-Decade-Old Prophecy Come True? Artificial Bad Intelligence Triggered by Merely a Single-Bit Flip in Large Language Models (Yan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00490&sa=D&source=editors&ust=1779048535998997&usg=AOvVaw3JFUVwfmcvdkHcSImt2Oc4) |
| 378 | [▪ Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution (Tong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.20476&sa=D&source=editors&ust=1779048535999065&usg=AOvVaw1h3egOqxfKCQ0vR0xl1ELm) |
| 379 | [▪ Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models (Yu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.21761&sa=D&source=editors&ust=1779048535999130&usg=AOvVaw0ZU9ZONgze6r-RVkPgpEJe) |
| 380 | [▪ Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.26032&sa=D&source=editors&ust=1779048535999198&usg=AOvVaw3ZN-BG0d6_Rs4RLfCdv6bc) |
| 381 | [▪ TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent (Meier et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.20118&sa=D&source=editors&ust=1779048535999266&usg=AOvVaw1SZKSYl10SvfQj1aehPzaJ) |
| 382 | [▪ Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack (Chen et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23871&sa=D&source=editors&ust=1779048535999334&usg=AOvVaw0tIvW9xlKayd6DXjr_EBnu) |
| 383 | [▪ GPM: The Gaussian Pancake Mechanism for Planting Undetectable Backdoors in Differential Privacy (Sun and He, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23834&sa=D&source=editors&ust=1779048535999402&usg=AOvVaw1peA94Ig0a2gGXaxqoRL9A) |
| 384 | [▪ DUP: Detection-guided Unlearning for Backdoor Purification in Language Models (Hu et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01647&sa=D&source=editors&ust=1779048535999498&usg=AOvVaw3YUqSZBTnFQKm2C5QNf2Yk) |
| 385 | [▪ Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models (Dai et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01605&sa=D&source=editors&ust=1779048535999598&usg=AOvVaw2YFChOurDlg2aWzV3uUCzv) |
| 386 | [▪ BeDKD: Backdoor Defense based on Dynamic Knowledge Distillation and Directional Mapping Modulator (Wu et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01595&sa=D&source=editors&ust=1779048535999678&usg=AOvVaw0bJuEA6tuNBV0d2aCJzJ4u) |
| 387 | [▪ Backdoor Attacks on Deep Learning Face Detection (Roux et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00620&sa=D&source=editors&ust=1779048535999790&usg=AOvVaw3Z_SPZ81111J-A5bvVLPC5) |
| 388 | [▪ SDBA: A Stealthy and Long-Lasting Durable Backdoor Attack in Federated Learning (Choe et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14805&sa=D&source=editors&ust=1779048535999869&usg=AOvVaw1JdvaoubRaLCpTeSEhmr3D) |
| 389 | [▪ FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning (Yan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21177&sa=D&source=editors&ust=1779048535999943&usg=AOvVaw0zMHsnSupPC41ovI5h5x2U) |
| 390 | [▪ Persistent Backdoor Attacks in Continual Learning (Guo, Kumar, and Tourani, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.13864&sa=D&source=editors&ust=1779048536000023&usg=AOvVaw1_Th0DVg5DR7kF9tbvBZGe) |
| 391 | [▪ ConSeg: Contextual Backdoor Attack Against Semantic Segmentation (Abbasi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19905&sa=D&source=editors&ust=1779048536000096&usg=AOvVaw1MMbb8_WuXgvUy4GYyNe6Z) |
| 392 | [▪ Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models (Pan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.20650&sa=D&source=editors&ust=1779048536000192&usg=AOvVaw38777cT9OZlcKra0_Pujen) |
| 393 | [▪ BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model (Lin et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.11798&sa=D&source=editors&ust=1779048536000266&usg=AOvVaw1MqMdlhH3-KiCxOiKfktTp) |
| 394 | [▪ VTarbel: Targeted Label Attack with Minimal Knowledge on Detector-enhanced Vertical Federated Learning (Tan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14625&sa=D&source=editors&ust=1779048536000352&usg=AOvVaw0fQJnjcp3CicGhXhrlPvM_) |
| 395 | [▪ BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems (Fang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.01593&sa=D&source=editors&ust=1779048536000447&usg=AOvVaw2tp9pN0tF-1WVNPVRujxs6) |
| 396 | [▪ Invisible Textual Backdoor Attacks based on Dual-Trigger (Hou et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.17531&sa=D&source=editors&ust=1779048536000561&usg=AOvVaw1YaiES1Sm9539PkQMCFpge) |
| 397 | [▪ An Adversarial-Driven Experimental Study on Deep Learning for RF Fingerprinting (Cao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14109&sa=D&source=editors&ust=1779048536000654&usg=AOvVaw2R-yp5N5iVM405ySSkeJcJ) |
| 398 | [▪ Architectural Backdoors in Deep Learning: A Survey of Vulnerabilities, Detection, and Defense (Childress, Collyer, and Knapp, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12919&sa=D&source=editors&ust=1779048536000725&usg=AOvVaw3M2_66XayDiLhRSbHW8k6y) |
| 399 | [▪ Non-Adaptive Adversarial Face Generation (Kim et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12107&sa=D&source=editors&ust=1779048536000789&usg=AOvVaw20E7yS3M2gu7eG8ndMHLAP) |
| 400 | [▪ Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks (Gosch et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.10867&sa=D&source=editors&ust=1779048536000857&usg=AOvVaw3dZdTu9pvJrkO_G56EkF0v) |
| 401 | [▪ Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs (Sivapiromrat et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11112&sa=D&source=editors&ust=1779048536000939&usg=AOvVaw3dP_OaXGVBMgixxOV366HQ) |
| 402 | [▪ 3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models (Yin et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10733&sa=D&source=editors&ust=1779048536001020&usg=AOvVaw2w05ghdZtPFPD5fEIR3B90) |
| 403 | [▪ No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms (Kazdan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.19537&sa=D&source=editors&ust=1779048536001096&usg=AOvVaw2Wf_qZDek8klSc1rdtPc79) |
| 404 | [▪ BURN: Backdoor Unlearning via Adversarial Boundary Analysis (Su et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10491&sa=D&source=editors&ust=1779048536001169&usg=AOvVaw2GPDnjq1g1NolIRJaWQYlB) |
| 405 | [▪ Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.17946&sa=D&source=editors&ust=1779048536001245&usg=AOvVaw2JqUhnraQl6t46ve6yEArY) |
| 406 | [▪ One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP (Xu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.19840&sa=D&source=editors&ust=1779048536001320&usg=AOvVaw1ogbGTdU9X6lOpvX-w4ca0) |
| 407 | [▪ Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems (Jiao et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20774&sa=D&source=editors&ust=1779048536001395&usg=AOvVaw1FigG8xWDBgBGgoyZiYGjy) |
| 408 | [▪ How to Backdoor Consistency Models (Wang and Kantarcioglu, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19785&sa=D&source=editors&ust=1779048536001466&usg=AOvVaw3ISfcq27Q_E7hV68HnmAaK) |
| 409 | [▪ Towards Backdoor Stealthiness in Model Parameter Space (Xu et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.05928&sa=D&source=editors&ust=1779048536001557&usg=AOvVaw04XTWOLsgv1fNvlZPS6TLv) |
| 410 | [▪ Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data (Zhang et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.05835&sa=D&source=editors&ust=1779048536001626&usg=AOvVaw1HPh5NL3ENGpYk7okXHPL0) |
| 411 | [▪ A Backdoor Attack Scheme with Invisible Triggers Based on Model Architecture Modification (Ma et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.16905&sa=D&source=editors&ust=1779048536001691&usg=AOvVaw2UPkOU3uKUWCHa3PDFQaER) |
| 412 | [▪ Double Landmines: Invisible Textual Backdoor Attacks based on Dual-Trigger (Hou et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.17531&sa=D&source=editors&ust=1779048536001756&usg=AOvVaw2h7Y526vFAgzaMdWqxAFCn) |
| 413 | [▪ Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models (Li et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05948&sa=D&source=editors&ust=1779048536001821&usg=AOvVaw2KRe_1dlW7f4Ja8z1s5Jri) |
| 414 | [▪ UOR: Universal Backdoor Attacks on Pre-trained Language Models (Du et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2305.09574%23&sa=D&source=editors&ust=1779048536001906&usg=AOvVaw1TLf1-j0aDm7PXgwdhX8m5) |
| 415 | [▪ Client-Side Patching against Backdoor Attacks in Federated Learning (Molina-Coronado, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.10605&sa=D&source=editors&ust=1779048536001978&usg=AOvVaw2SPqQ66LnlwJ1Er3zfZtNB) |
| 416 | [▪ Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models (Li et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.07667&sa=D&source=editors&ust=1779048536002050&usg=AOvVaw02Q13jy4MfOzPmoeMFf60s) |
| 417 | [▪ Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions (Cinà et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2106.07214&sa=D&source=editors&ust=1779048536002117&usg=AOvVaw3zqK92gSM48moCrXQI84_8) |
| 418 | [▪ Proactive Adversarial Defense: Harnessing Prompt Tuning in Vision-Language Models to Detect Unseen Backdoored Images (Stein et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.08755&sa=D&source=editors&ust=1779048536002194&usg=AOvVaw0oeRhLm-JJonAXu2L-iBw-) |
| 419 | [▪ Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP (Singh, Croce, and Hein, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.00727&sa=D&source=editors&ust=1779048536002278&usg=AOvVaw3036W500n4zZn0IIaHkQ0r) |
| 420 | [▪ RTL-Breaker: Assessing the Security of LLMs against Backdoor Attacks on HDL Code Generation (Mankali et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.17569&sa=D&source=editors&ust=1779048536002348&usg=AOvVaw3dQ-BSonPpZ5aGo5Ztnxlp) |
| 421 | [▪ Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining (Wu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.02454&sa=D&source=editors&ust=1779048536002440&usg=AOvVaw2nKbNLAgMipORkoMWSKErZ) |
| 422 | [▪ LoBAM: LoRA-Based Backdoor Attack on Model Merging (Yin et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.16746&sa=D&source=editors&ust=1779048536002526&usg=AOvVaw2atE1u6ndtz_zm3BH8-K6s) |
| 423 | [▪ Towards Clean-Label Backdoor Attacks in the Physical World (Dao et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.19203&sa=D&source=editors&ust=1779048536002598&usg=AOvVaw2xtNsXuM-st06MCCwXVDa0) |
| 424 | [▪ Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models (Shin and Park, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14449&sa=D&source=editors&ust=1779048536002673&usg=AOvVaw0-asGgVEjzkY0j7gb0xqML) |
| 425 | [▪ Memory Backdoor Attacks on Neural Networks (Luzon et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14516&sa=D&source=editors&ust=1779048536002745&usg=AOvVaw05i8IbpziwIQxBKgQfwx4E) |
| 426 | [▪ AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection (Lu et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14243&sa=D&source=editors&ust=1779048536002819&usg=AOvVaw0Yl1E-y3zD7T5Ll2M-1jBX) |
| 427 | [▪ When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations (Ge et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.12701&sa=D&source=editors&ust=1779048536002895&usg=AOvVaw3GUPMjSGiy1FuU9d_7IYcT) |
| 428 | [▪ Backdoor defense, learnability and obfuscation (Christiano et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03077&sa=D&source=editors&ust=1779048536002966&usg=AOvVaw1FFXn2tFq6bb9tNamwiNNz) |
| 429 | [▪ Combinational Backdoor Attack against Customized Text-to-Image Models (Jiang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.12389&sa=D&source=editors&ust=1779048536003047&usg=AOvVaw2eku5e_LD0er28bVrLE5Up) |
| 430 | [▪ Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion (Guo et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22678&sa=D&source=editors&ust=1779048536003132&usg=AOvVaw2juEBc1kXvBznuSpiOCyWT) |
| 431 | [▪ Backdoor Attacks against Image-to-Image Networks (Jiang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.10445&sa=D&source=editors&ust=1779048536003197&usg=AOvVaw0ViYOL99jbbN8BqVQv49PU) |
| 432 | [▪ Infighting in the Dark: Multi-Labels Backdoor Attack in Federated Learning (Li et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.19601&sa=D&source=editors&ust=1779048536003262&usg=AOvVaw0943KovVWiUD45D8AC-Yew) |
| 433 | [▪ Planting Undetectable Backdoors in Machine Learning Models (Goldwasser et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2204.06974&sa=D&source=editors&ust=1779048536003326&usg=AOvVaw0Ou6oSItHbMKXjXJoG5s2F) |
| 434 | [▪ On the Credibility of Backdoor Attacks Against Object Detectors in the Physical World (Gia Doan et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.12122&sa=D&source=editors&ust=1779048536003416&usg=AOvVaw0gsxAAhiSlFSFTxRNSHG_0) |
| 435 | ▪ Model X-ray:Detecting Backdoored Models via Decision Boundary (Su et al, Oct 2024) |
| 436 | [▪ Dullahan: Stealthy Backdoor Attack against Without-Label-Sharing Split Learning (Pu et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.12751&sa=D&source=editors&ust=1779048536003539&usg=AOvVaw1mGj2eMyWtFaTktzSkcivZ) |
| 437 | [▪ Mind Your Questions Towards Backdoor Attacks on Text-to-Visualization Models (Li et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.06782&sa=D&source=editors&ust=1779048536003606&usg=AOvVaw21UIUK87LRr2wTYQq3kCGt) |
| 438 | [▪ BadCM: Invisible Backdoor Attack Against Cross-Modal Learning (Zheng et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02182&sa=D&source=editors&ust=1779048536003671&usg=AOvVaw179NZMb06Gxk1BsQoQnlYq) |
| 439 | [▪ BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting (Lin et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02195&sa=D&source=editors&ust=1779048536003735&usg=AOvVaw3FI6rfCehNgQBA7ZxgdnYJ) |
| 440 | [▪ Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning (Zhao et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2401.05949&sa=D&source=editors&ust=1779048536003803&usg=AOvVaw3TLjZgXBPi7q6-Sy3bBxYb) |
| 441 | [▪ Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations (Bartolini, Stoyanov, and Giaretta, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.12553&sa=D&source=editors&ust=1779048536003871&usg=AOvVaw1i5cPay0IgED-QE_BVk_BJ) |
| 442 | [▪ EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody (Schoof et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.01178&sa=D&source=editors&ust=1779048536003936&usg=AOvVaw2W1_wxab-6AUly4KevGg0I) |
| 443 | [▪ Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers (Abad et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.04142&sa=D&source=editors&ust=1779048536004007&usg=AOvVaw307DTr36_N3LegGHCQi-BT) |
| 444 | [▪ Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning (Zhang, Gong, and Reiter, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.06206&sa=D&source=editors&ust=1779048536004076&usg=AOvVaw2rX6TyuYCGZMDc63_exQb8) |
| 445 | [▪ TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors (Mo et al , Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.05294&sa=D&source=editors&ust=1779048536004141&usg=AOvVaw1SWAiodgB9xm8JjG0dWudw) |
| 446 | [▪ INK: Inheritable Natural Backdoor Attack Against Model Distillation (Liu et al , Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2304.10985&sa=D&source=editors&ust=1779048536004207&usg=AOvVaw0lmdo4tx8s51umFuwz-gpg) |
| 447 | [▪ Exploiting the Vulnerability of Large Language Models via Defense-Aware Architectural Backdoor (Miah and Bi, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.01952&sa=D&source=editors&ust=1779048536004279&usg=AOvVaw1UBWlYpc8mbin8nKM3mANW) |
| 448 | [▪ Rethinking Backdoor Detection Evaluation for Language Models (Yan et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00399&sa=D&source=editors&ust=1779048536004348&usg=AOvVaw3uAXs5tt1yI8nA42mmCl2p) |
| 449 | [▪ Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks (Li et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2401.15295&sa=D&source=editors&ust=1779048536004414&usg=AOvVaw1-Dm_rBwEfmxUE7vxKEq3t) |
| 450 | [▪ Transferring Backdoors between Large Language Models by Knowledge Distillation (Cheng et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09878&sa=D&source=editors&ust=1779048536004481&usg=AOvVaw2i65jAbMnN2hAL0MEqcE35) |
| 451 | [▪ Revocable Backdoor for Deep Model Trading (Zu et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.00255&sa=D&source=editors&ust=1779048536004549&usg=AOvVaw12umtoUVrxes8QdDhPncrV) |
| 452 | [▪ BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning (Wu et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.19845&sa=D&source=editors&ust=1779048536004615&usg=AOvVaw0IQxNMbkjc6wR3rrwRUZPF) |
| 453 | [▪ Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models (Hao et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.21316&sa=D&source=editors&ust=1779048536004678&usg=AOvVaw20KFSSl0aDJNKbQdSh4jIL) |

|     |
| --- |
| Backdoor ML Model and Craft Adversarial Data |

**>**

**<**

‍

#### Supply Chain Vulnerabilities and Compromise

Covers:

- OWASP LLM 05/OWASP ML 06/MITRE ATLAS Initial Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Supply Chain Vulnerabilities and Compromise

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Exploiting LLM Agent Supply Chains via Payload-less Skills (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14460&sa=D&source=editors&ust=1779048535950723&usg=AOvVaw3Xql0_h4Tj3xhy8fJKWILe) |
| 3 | [▪ Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry (Saha, Faghih, and Feizi, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11418&sa=D&source=editors&ust=1779048535950885&usg=AOvVaw0_xLi5Njdn3ZNfxt6v4thy) |
| 4 | [▪ Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI Package Ecosystems (McCann, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03309&sa=D&source=editors&ust=1779048535950984&usg=AOvVaw3XQzZGsiB2y6fKmS65R931) |
| 5 | [▪ Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27426&sa=D&source=editors&ust=1779048535951091&usg=AOvVaw017yXzXFGdNd0FKWB4nUc3) |
| 6 | [▪ SOK: A Taxonomy of Attack Vectors and Defense Strategies for Agentic Supply Chain Runtime (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19555&sa=D&source=editors&ust=1779048535951180&usg=AOvVaw0m3ikCkuVxYW5bI1OG4Vkk) |
| 7 | [▪ Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06572&sa=D&source=editors&ust=1779048535951266&usg=AOvVaw2ejRZuKa4ha4YGInwJXUOI) |
| 8 | [▪ Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08407&sa=D&source=editors&ust=1779048535951348&usg=AOvVaw36Eb_qzD9oQoK3eXd3EUMi) |
| 9 | [▪ Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems (Qu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03081&sa=D&source=editors&ust=1779048535951450&usg=AOvVaw0KPNnKonwyqRX5bsN8KR7b) |
| 10 | [▪ How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks? (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2310.00710&sa=D&source=editors&ust=1779048535951535&usg=AOvVaw0XnXUYX7em6TDcF-Z0VyVh) |
| 11 | [▪ SynthChain: A Synthetic Benchmark and Forensic Analysis of Advanced and Stealthy Software Supply Chain Attacks (Tan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16694&sa=D&source=editors&ust=1779048535951624&usg=AOvVaw3ro4qsZC1puTkBcPRceL19) |
| 12 | [▪ On the (In)Security of Loading Machine Learning Models (Digregorio et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06703&sa=D&source=editors&ust=1779048535951706&usg=AOvVaw1Y9g1h6R5ILhhwnXpd-NFb) |
| 13 | [▪ Formal Analysis and Supply Chain Security for Agentic AI Skills (Bhardwaj, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00195&sa=D&source=editors&ust=1779048535951786&usg=AOvVaw3Zc2l0_pdVFIV3-2tZDb9F) |
| 14 | [▪ LLM Scalability Risk for Agentic-AI and Model Supply Chain Security (Ahi, Agrawal, and Valizadeh, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19021&sa=D&source=editors&ust=1779048535951869&usg=AOvVaw3KYHNculTQiE6UZEBs_Nn-) |
| 15 | [▪ Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines (Jamshidi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.23132&sa=D&source=editors&ust=1779048535951951&usg=AOvVaw1Od5jegeXN_oQhufAlXwNa) |
| 16 | [▪ AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training (Vandendriessche et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05703&sa=D&source=editors&ust=1779048535952037&usg=AOvVaw0Ev6ftlfVp14uLZlA9IcRZ) |
| 17 | [▪ Understanding Security Risks of AI Agents' Dependency Updates (Singla et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00205&sa=D&source=editors&ust=1779048535952116&usg=AOvVaw05k1xKdZegvLIpjhnYnW6C) |
| 18 | [▪ Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects? (Nguyen, Le, and Babar, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23385&sa=D&source=editors&ust=1779048535952206&usg=AOvVaw20DrqbvhD15vG-2M56ZjgS) |
| 19 | [▪ CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution (Zhu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04580&sa=D&source=editors&ust=1779048535952291&usg=AOvVaw08ZjaFWAuMGSzxZTX8QsWX) |
| 20 | [▪ A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution (Zhu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04580&sa=D&source=editors&ust=1779048535952376&usg=AOvVaw1ZM1xzVI35Qb9Bm0p5o2hZ) |
| 21 | [▪ Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications (Sheh and Geappen, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15763&sa=D&source=editors&ust=1779048535952460&usg=AOvVaw0R-CBTlhEN2ZBcxO8QMM0H) |
| 22 | [▪ Supply Chain Exploitation of Secure ROS 2 Systems: A Proof-of-Concept on Autonomous Platform Compromise via Keystore Exfiltration (Sakib et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00140&sa=D&source=editors&ust=1779048535952557&usg=AOvVaw0e4yL4plDAYMUETza_hlmK) |
| 23 | [▪ Lexo: Eliminating Stealthy Supply-Chain Attacks via LLM-Assisted Program Regeneration (Lamprou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14522&sa=D&source=editors&ust=1779048535952650&usg=AOvVaw3zqZsqggoiNxAJLUDmyqFt) |
| 24 | [▪ Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain (Boisvert et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05159&sa=D&source=editors&ust=1779048535952760&usg=AOvVaw2mgKWEJlLArmX5oShl1e7Y) |
| 25 | [▪ ExclaveFL: Providing Transparency to Federated Learning using Exclaves (Guo et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.10537&sa=D&source=editors&ust=1779048535952850&usg=AOvVaw2coEQqGuvGNVkVIg2InvC-) |
| 26 | [▪ VDGraph: A Graph-Theoretic Approach to Unlock Insights from SBOM and SCA Data (Xia et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20502&sa=D&source=editors&ust=1779048535952933&usg=AOvVaw3T9PnmtOj-X7kZhEtjt3ww) |
| 27 | [▪ Understanding the Supply Chain and Risks of Large Language Model Applications (Ma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18105&sa=D&source=editors&ust=1779048535953013&usg=AOvVaw2NMc6jfE6ouk0ZflZN9LDp) |
| 28 | [▪ PyPitfall: Dependency Chaos and Software Supply Chain Vulnerabilities in Python (Mahon, Hou, and Yao, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18075&sa=D&source=editors&ust=1779048535953095&usg=AOvVaw3CHGHr5l56eahDJ6kS1R6W) |
| 29 | [▪ Security Enclave Architecture for Heterogeneous Security Primitives for Supply-Chain Attacks (Raj et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10971&sa=D&source=editors&ust=1779048535953185&usg=AOvVaw2GddrscmETVPxxqhqVF0m6) |
| 30 | [▪ Exploiting Leaderboards for Large-Scale Distribution of Malicious Models (Suri et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08983&sa=D&source=editors&ust=1779048535953270&usg=AOvVaw2UpbdXMnpUw99RsrK78ySB) |
| 31 | [▪ Rugsafe: A multichain protocol for recovering from and defending against Rug Pulls (Pharr and Hussain, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06423&sa=D&source=editors&ust=1779048535953354&usg=AOvVaw1zcLcVmpT8JNADxMDTl3M2) |

|     |
| --- |
| Supply Chain Vulnerabilities and Compromise |

**>**

**<**

‍

#### Excessive Agency, Agentic Manipulation, Agentic Systems

We added 'agentic' manipulation to this subcategory.

Covers:

- OWASP LLM 08: Excessive Agency

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Excessive Agency, Agentic Manipulation, Agentic Systems

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Progent: Securing AI Agents with Privilege Control (Shi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.11703&sa=D&source=editors&ust=1779048536637145&usg=AOvVaw2k7E8iLhF6tnoidTYCPo7O) |
| 3 | [▪ Do Coding Agents Understand Least-Privilege Authorization? (Yan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14859&sa=D&source=editors&ust=1779048536637225&usg=AOvVaw0l9Z_4CZujHDvKyU_5jZAD) |
| 4 | [▪ Quantitative Certification of Agentic Tool Selection (Yeon, Chaudhary, and Singh, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03992&sa=D&source=editors&ust=1779048536637298&usg=AOvVaw1Q_31MCLY-5rSNDjcLb7UL) |
| 5 | [▪ Language-Based Agent Control (Zhou, D'Antoni, and Polikarpova, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12863&sa=D&source=editors&ust=1779048536637370&usg=AOvVaw2C-R0NwDKRisY3Koj39jqj) |
| 6 | [▪ Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety (Bilal et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12729&sa=D&source=editors&ust=1779048536637448&usg=AOvVaw1h7Q8Mu41gkS-bifsfOeiO) |
| 7 | [▪ No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13044&sa=D&source=editors&ust=1779048536637519&usg=AOvVaw1Ba6gVo_Xe-eFH2USB2LIY) |
| 8 | [▪ A Security Analysis of the OpenClaw AI Agent Framework (Suwansathit, Zhang, and Gu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27517&sa=D&source=editors&ust=1779048536637589&usg=AOvVaw1ggF6BXRKJp59fdWdIZI3G) |
| 9 | [▪ Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries (Laws, Oprea, and Nita-Rotaru, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12364&sa=D&source=editors&ust=1779048536637663&usg=AOvVaw2gbKpoo8LJqRZwOYrLu7jQ) |
| 10 | [▪ SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces (Jin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12015&sa=D&source=editors&ust=1779048536637759&usg=AOvVaw2zirIpJ6wZsnSEwItZpOtg) |
| 11 | [▪ Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems (Zhou, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11891&sa=D&source=editors&ust=1779048536637830&usg=AOvVaw2Fta8CVPr9n2CYsfL_gQvj) |
| 12 | [▪ Five Attacks on x402 Agentic Payment Protocol (Li, Wang, and Wang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11781&sa=D&source=editors&ust=1779048536637900&usg=AOvVaw0ZO6645jLurbclhW3iI19G) |
| 13 | [▪ Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11664&sa=D&source=editors&ust=1779048536637974&usg=AOvVaw0rkveLctN7XkmxMfAvPjEl) |
| 14 | [▪ FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11514&sa=D&source=editors&ust=1779048536638046&usg=AOvVaw1MWR7fM5rmdLWZZRg1kCUC) |
| 15 | [▪ Digital Identity for Agentic Systems: Toward a Portable Authorization Standard for Autonomous Agents (Madhira, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11487&sa=D&source=editors&ust=1779048536638117&usg=AOvVaw2J4GsgJo12y-eblgfVz7Fu) |
| 16 | [▪ Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution (Fendley et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11229&sa=D&source=editors&ust=1779048536638187&usg=AOvVaw1Gvb1WwJXgyGeGbZbV9I6F) |
| 17 | [▪ Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw (Yao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11047&sa=D&source=editors&ust=1779048536638282&usg=AOvVaw1eOIkW_MrZeVck4cMszqfm) |
| 18 | [▪ The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck (Fan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11039&sa=D&source=editors&ust=1779048536638357&usg=AOvVaw0-Mge1HO-VSxlo-su9tzC4) |
| 19 | [▪ Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents (Ravindran, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11032&sa=D&source=editors&ust=1779048536638433&usg=AOvVaw3IJZUvgX2sdVUOaUD2VLVx) |
| 20 | [▪ AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents (Rassul and Rashid, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11026&sa=D&source=editors&ust=1779048536638503&usg=AOvVaw1WzOkfP-B-axV0Hd9nYJNb) |
| 21 | [▪ The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents (Wu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11003&sa=D&source=editors&ust=1779048536638576&usg=AOvVaw27KPRQUAV-_YeZq2xkvRYI) |
| 22 | [▪ Formal Policy Enforcement for Real-World Agentic Systems (Palumbo et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16708&sa=D&source=editors&ust=1779048536638644&usg=AOvVaw2_LiMxLEeGanvq3xH9bUFP) |
| 23 | [▪ MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study (hamme et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10763&sa=D&source=editors&ust=1779048536638714&usg=AOvVaw3HLRQv6PV_Bas8M6WQOXFX) |
| 24 | [▪ Evaluating Tool Cloning in Agentic-AI Ecosystems (Kim et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09817&sa=D&source=editors&ust=1779048536638784&usg=AOvVaw1TsftbGe92JiWTmRTTXZl9) |
| 25 | [▪ Agentic Fuzzing: Opportunities and Challenges (Park and Yun, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10074&sa=D&source=editors&ust=1779048536638871&usg=AOvVaw0tNUpLzl9fONey95Qt1xn0) |
| 26 | [▪ Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents (Wang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09863&sa=D&source=editors&ust=1779048536638951&usg=AOvVaw26tCbz5IT7STsz_TNArLEr) |
| 27 | [▪ Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments (Goel, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09721&sa=D&source=editors&ust=1779048536639024&usg=AOvVaw3OptCt0leagBOmnVPSK-r5) |
| 28 | [▪ When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks (Cai, Zhang, and Hei, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08460&sa=D&source=editors&ust=1779048536639096&usg=AOvVaw3XJ9GkU6PQcvtm36PHdGAJ) |
| 29 | [▪ MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security (Avizeh et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06933&sa=D&source=editors&ust=1779048536639166&usg=AOvVaw1iH-ranYcA5GVBNm7o-GA_) |
| 30 | [▪ Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07135&sa=D&source=editors&ust=1779048536639248&usg=AOvVaw1fI22PLFGdDCGyifQeIAzN) |
| 31 | [▪ Language Models Can Autonomously Hack and Self-Replicate (Air et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06760&sa=D&source=editors&ust=1779048536639319&usg=AOvVaw1L6TsYVflYxZqWEMDak3U5) |
| 32 | [▪ Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand (Koch, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06713&sa=D&source=editors&ust=1779048536639394&usg=AOvVaw0DeCVywYrifIZlcZDjCK2c) |
| 33 | [▪ AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments? (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03117&sa=D&source=editors&ust=1779048536639472&usg=AOvVaw1L-gvJSruNsxSZeA6Rn3pH) |
| 34 | [▪ Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches (David and Gervais, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06601&sa=D&source=editors&ust=1779048536639544&usg=AOvVaw0R0RisLBlxhrr1mVVFqHyU) |
| 35 | [▪ Constraining Host-Level Abuse in Self-Hosted Computer-Use Agents via TEE-Backed Isolation (Lu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06393&sa=D&source=editors&ust=1779048536639614&usg=AOvVaw2qXxTzh3gsSUgrdTargpe6) |
| 36 | [▪ ClawGuard: Out-of-Band Detection of LLM Agent Workflow Hijacking via EM Side Channel (Gan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06205&sa=D&source=editors&ust=1779048536639683&usg=AOvVaw0-AeJnQWJcKcSK7k2vYE17) |
| 37 | [▪ SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills (Wu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05868&sa=D&source=editors&ust=1779048536639755&usg=AOvVaw0KmUoyBvWoNvRnhcXLG8tI) |
| 38 | [▪ WAAA! Web Adversaries Against Agentic Browsers (Datta et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05509&sa=D&source=editors&ust=1779048536639837&usg=AOvVaw1re_j2DaoIrfzIM8052w6k) |
| 39 | [▪ Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use (Arceo and Narsing, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05287&sa=D&source=editors&ust=1779048536639929&usg=AOvVaw0kcgLqOhE99QFxipLclsyT) |
| 40 | [▪ AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use (Yang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04785&sa=D&source=editors&ust=1779048536640002&usg=AOvVaw0rl6qZHqeWaZBCwGD0Aeo4) |
| 41 | [▪ Agentic Vulnerability Reasoning on Windows COM Binaries (Lee, Kim, and Zhang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05000&sa=D&source=editors&ust=1779048536640071&usg=AOvVaw1QtTeo9nWAhSjf2LBRzhce) |
| 42 | [▪ Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours (Dheekonda, Pearce, and Landers, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04019&sa=D&source=editors&ust=1779048536640141&usg=AOvVaw3SJXMbCQtRlKK1yRpL4D1b) |
| 43 | [▪ Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense (Prinos et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03034&sa=D&source=editors&ust=1779048536640212&usg=AOvVaw3XvXRN0qUMmuOWQzh923Qu) |
| 44 | [▪ When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI (Forough, Kogias, and Haddadi, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03213&sa=D&source=editors&ust=1779048536640282&usg=AOvVaw0dIQ6X9ghIRsg_F8WAXzmJ) |
| 45 | [▪ Towards Agentic Runtime Healing (Sun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2408.01055&sa=D&source=editors&ust=1779048536640348&usg=AOvVaw3VlFPCt-bxTmLgwbgafHz5) |
| 46 | [▪ Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit (Sun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07249&sa=D&source=editors&ust=1779048536640421&usg=AOvVaw0ZaBhpWD3EPryX6dRnJz0t) |
| 47 | [▪ When Alignment Isn't Enough: Response-Path Attacks on LLM Agents (Luo et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02187&sa=D&source=editors&ust=1779048536640493&usg=AOvVaw30Gw8E-EAoy1MjRhMbLzly) |
| 48 | [▪ Architectural Obsolescence of Unhardened Agentic-AI Runtimes (Metere, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01740&sa=D&source=editors&ust=1779048536640563&usg=AOvVaw0LDD_At5zBZ33IGLLN134w) |
| 49 | [▪ AgenticVM: Agentic AI for Adaptive Software Vulnerability Management (Arifin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01739&sa=D&source=editors&ust=1779048536640688&usg=AOvVaw1ET1e6_9s-j5kUlssWpE8O) |
| 50 | [▪ Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems (Jamshidi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00741&sa=D&source=editors&ust=1779048536640765&usg=AOvVaw0-JHoxx3MSD1DnJM6fdktH) |
| 51 | [▪ Alignment Contracts for Agentic Security Systems (David, Guarnieri, and Gervais, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00081&sa=D&source=editors&ust=1779048536640835&usg=AOvVaw1_mRVDK0sSTbSWJihLebw9) |
| 52 | [▪ Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure (Cuadros and Maiga, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00055&sa=D&source=editors&ust=1779048536640906&usg=AOvVaw1LtZajuVXRlNubJUddbeec) |
| 53 | [▪ Chronology of Multi-Agent Interactions for Provenance of Evolving Information (Chang and Echizen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.12612&sa=D&source=editors&ust=1779048536640980&usg=AOvVaw3jsuZz9SNLt9NV2GyO6-fY) |
| 54 | [▪ From surveillance to signalling: escalation channels as environmental controls for agentic AI (Gomez, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05192&sa=D&source=editors&ust=1779048536641049&usg=AOvVaw3zDtJvZsC4yrCaIo7J87BD) |
| 55 | [▪ From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems (Nagaraja, Bahsi, and Cunha, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27267&sa=D&source=editors&ust=1779048536641121&usg=AOvVaw1OYyep6PPo7L_xwK8b7LtB) |
| 56 | [▪ Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes (Mittal and Cruz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26997&sa=D&source=editors&ust=1779048536641202&usg=AOvVaw1rcUvSve55iOhe4RKJA-g1) |
| 57 | [▪ A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron? (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.10924&sa=D&source=editors&ust=1779048536641291&usg=AOvVaw0N6E4wKyU277hm-lUT2YJv) |
| 58 | [▪ Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents (Witt et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.02077&sa=D&source=editors&ust=1779048536641364&usg=AOvVaw20dnGQCRs0SWwJhvY1N0_j) |
| 59 | [▪ SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization (Meidan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26394&sa=D&source=editors&ust=1779048536641441&usg=AOvVaw3Og2Ha_GuR9w5yQfjcOXF1) |
| 60 | [▪ Towards Agentic Investigation of Security Alerts (Eilertsen, Mavroeidis, and Grov, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25846&sa=D&source=editors&ust=1779048536641510&usg=AOvVaw0RO2oRO5qXL_QDiner5wE1) |
| 61 | [▪ From CRUD to Autonomous Agents: Formal Validation and Zero-Trust Security for Semantic Gateways in AI-Native Enterprise Systems (Peyrano, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25555&sa=D&source=editors&ust=1779048536641608&usg=AOvVaw2CfZbURoSLdKTHQfeiQ8SL) |
| 62 | [▪ AgentDID: Trustless Identity Authentication for AI Agents (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25189&sa=D&source=editors&ust=1779048536641675&usg=AOvVaw06k2hh5ttfpEtycUU0Pn3r) |
| 63 | [▪ Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills (Lv et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25109&sa=D&source=editors&ust=1779048536641748&usg=AOvVaw0t0osqVLGeXlFtz0pJwfjq) |
| 64 | [▪ SUDP: Secret-Use Delegation Protocol for Agentic Systems (Yu, Geng, and Knottenbelt, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24920&sa=D&source=editors&ust=1779048536641826&usg=AOvVaw2XkN_CEfMICUK3JrO54gMx) |
| 65 | [▪ Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture (Xiang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23646&sa=D&source=editors&ust=1779048536641895&usg=AOvVaw1LCY6D0cLxl5deIBEvUe5E) |
| 66 | [▪ AI Identity: Standards, Gaps, and Research Directions for AI Agents (Otsuka, Toyoda, and Leung, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23280&sa=D&source=editors&ust=1779048536641966&usg=AOvVaw3h4XVlvudykaU1mouTY5eP) |
| 67 | [▪ AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24657&sa=D&source=editors&ust=1779048536642034&usg=AOvVaw1aEkpcVKwcOct8bZsz3v9-) |
| 68 | [▪ When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape (Mitchell, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23425&sa=D&source=editors&ust=1779048536642106&usg=AOvVaw0Q0DWX72EYO4BQM5Gkmkez) |
| 69 | [▪ Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents (Cai et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23374&sa=D&source=editors&ust=1779048536642177&usg=AOvVaw31PhTKvnhWhaCcPsqd8cEt) |
| 70 | [▪ From Stateless Queries to Autonomous Actions: A Layered Security Framework for Agentic AI Systems (Chu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23338&sa=D&source=editors&ust=1779048536642247&usg=AOvVaw2i2F0xSQIYr_C4hBINWSVH) |
| 71 | [▪ AgentBound: Securing Execution Boundaries of AI Agents (B\\"uhler et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21236&sa=D&source=editors&ust=1779048536642326&usg=AOvVaw1gqv_h-E0eDCpwsFvk6J4e) |
| 72 | [▪ Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation (Andreucci and Castiglione, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22427&sa=D&source=editors&ust=1779048536642401&usg=AOvVaw1rhvnDLQae0-87A8i6EOJQ) |
| 73 | [▪ Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems (He and Yu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22136&sa=D&source=editors&ust=1779048536642477&usg=AOvVaw0a35QoLYjwGY03VuThRi9Q) |
| 74 | [▪ Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms (Azarafrooz, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21131&sa=D&source=editors&ust=1779048536642548&usg=AOvVaw0ffNfFfZ7xzDydcczQXsJO) |
| 75 | [▪ Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models (Belkhiter et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20994&sa=D&source=editors&ust=1779048536642620&usg=AOvVaw3ZmqQjar8Mgx20Yy3FIjj9) |
| 76 | [▪ AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation (Roy and Singh, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20134&sa=D&source=editors&ust=1779048536642702&usg=AOvVaw1Ba9zCMM8IBHovpKZp3rtI) |
| 77 | [▪ Whispers in the Machine: Confidentiality in Agentic Systems (Evertz et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2402.06922&sa=D&source=editors&ust=1779048536642803&usg=AOvVaw27kcx3PFQeM08i_yXmIR_d) |
| 78 | [▪ ClawCoin: An Agentic AI-Native Cryptocurrency for Decentralized Agent Economies (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19026&sa=D&source=editors&ust=1779048536642878&usg=AOvVaw3SvK9wSfhtszHxXIAK_MDH) |
| 79 | [▪ HadAgent: Harness-Aware Decentralized Agentic AI Serving with Proof-of-Inference Blockchain Consensus (Jimenez et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18614&sa=D&source=editors&ust=1779048536642954&usg=AOvVaw1XfQW39HwUiXdyX0FDN2qp) |
| 80 | [▪ An AI Agent Execution Environment to Safeguard User Data (Stanley et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19657&sa=D&source=editors&ust=1779048536643024&usg=AOvVaw3KSzwAyzOJ8MeJFOo92ZMY) |
| 81 | [▪ Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps (Chona, Kozlov, and Kumar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19533&sa=D&source=editors&ust=1779048536643095&usg=AOvVaw3A2_hdHwIrubnAnCstMgZ3) |
| 82 | [▪ Towards Optimal Agentic Architectures for Offensive Security Tasks (David and Gervais, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18718&sa=D&source=editors&ust=1779048536643166&usg=AOvVaw3C_4BoQIZU41Vagy2rTfwR) |
| 83 | [▪ Owner-Harm: A Missing Threat Model for AI Agent Safety (Zhang and Jiang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18658&sa=D&source=editors&ust=1779048536643234&usg=AOvVaw3XL3-QtqAblOPt27NtEcZI) |
| 84 | [▪ From Craft to Kernel: A Governance-First Execution Architecture and Semantic ISA for Agentic Computers (Wen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18652&sa=D&source=editors&ust=1779048536643306&usg=AOvVaw1kT1zil6z-m3m1ZfBfZ5Iy) |
| 85 | [▪ Evaluating Privilege Usage of Agents with Real-World Tools (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28166&sa=D&source=editors&ust=1779048536643376&usg=AOvVaw0RQc3PiHeMtjRjbnDv50Ck) |
| 86 | [▪ From Admission to Invariants: Measuring Deviation in Delegated Agent Systems (Fernandez, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17517&sa=D&source=editors&ust=1779048536643462&usg=AOvVaw2MtSAieoZxSgSe1VI93pwn) |
| 87 | [▪ Atomic Decision Boundaries: A Structural Requirement for Guaranteeing Execution-Time Admissibility in Autonomous Systems (Fernandez, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17511&sa=D&source=editors&ust=1779048536643537&usg=AOvVaw1gwq1SuTOulTz8I9UEonDG) |
| 88 | [▪ AgenTEE: Confidential LLM Agent Execution on Edge Devices (Abdollahi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18231&sa=D&source=editors&ust=1779048536643614&usg=AOvVaw1sffoIHGfUUsxnCe5yORlN) |
| 89 | [▪ CapSeal: Capability-Sealed Secret Mediation for Secure Agent Execution (Jin, Guo, and Cheung, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16762&sa=D&source=editors&ust=1779048536643709&usg=AOvVaw293vY-CSLQLg5W4nL9qoil) |
| 90 | [▪ A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty (Lin, Li, and Chen, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16548&sa=D&source=editors&ust=1779048536643805&usg=AOvVaw32qhcEjCCsCyQiG37OrH8F) |
| 91 | [▪ CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations (Panjwani, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16521&sa=D&source=editors&ust=1779048536643902&usg=AOvVaw3sXlpYNojEkignY8rGjjOM) |
| 92 | [▪ HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents? (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15415&sa=D&source=editors&ust=1779048536643998&usg=AOvVaw3hn1Nhvffk8YOjkDuuWtn8) |
| 93 | [▪ An Agentic Workflow for Detecting Personally Identifiable Information in Crash Narratives (Ma et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15369&sa=D&source=editors&ust=1779048536644073&usg=AOvVaw0R_mis1SaDaaXr0gKLbv_B) |
| 94 | [▪ SoK: Security of Autonomous LLM Agents in Agentic Commerce (Mao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15367&sa=D&source=editors&ust=1779048536644142&usg=AOvVaw2VkdWkr5VB7V3F_rXymrPt) |
| 95 | [▪ CBCL: Safe Self-Extending Agent Communication (O'Connor, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14512&sa=D&source=editors&ust=1779048536644213&usg=AOvVaw0oJK2rc-ETHSgJ0ArKOblY) |
| 96 | [▪ Challenges and Future Directions in Agentic Reverse Engineering Systems (Radey, West, and Fawaz, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14317&sa=D&source=editors&ust=1779048536644290&usg=AOvVaw3IT0FJlECiaRXZ_RzBodNO) |
| 97 | [▪ Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection (Ghimire et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13298&sa=D&source=editors&ust=1779048536644360&usg=AOvVaw0mKUmdUMzfLSHOsn1RMI4V) |
| 98 | [▪ To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems (He et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.02546&sa=D&source=editors&ust=1779048536644445&usg=AOvVaw3XharZor_ms9GA_dcU25n6) |
| 99 | [▪ Policy-Invisible Violations in LLM-Based Agents (Wu and Gong, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12177&sa=D&source=editors&ust=1779048536644518&usg=AOvVaw03Pjf_yF1Igfa1useS1uCS) |
| 100 | [▪ Parallax: Why AI Agents That Think Must Never Act (Fokou, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12986&sa=D&source=editors&ust=1779048536644585&usg=AOvVaw2CdjCcsDM4ZK8vuU5SipdH) |
| 101 | [▪ Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents (Sidik and Rokach, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11839&sa=D&source=editors&ust=1779048536644655&usg=AOvVaw2_FTKdgPcgqa0i8-Qw0zM5) |
| 102 | [▪ Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval (Hamzic et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11419&sa=D&source=editors&ust=1779048536644727&usg=AOvVaw0a3557UhVCSAI3Z9_lCoBH) |
| 103 | [▪ Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering (Stantchev, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11430&sa=D&source=editors&ust=1779048536644798&usg=AOvVaw1cHZxgzEMUJgNVnYk_6rk8) |
| 104 | [▪ The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11309&sa=D&source=editors&ust=1779048536644891&usg=AOvVaw396nCruT9Rwj4xV42yHznV) |
| 105 | [▪ The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents (Ding et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10577&sa=D&source=editors&ust=1779048536644991&usg=AOvVaw30xoeDdUWOUpvHWbLysChb) |
| 106 | [▪ Self-Sovereign Agent (Qu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08551&sa=D&source=editors&ust=1779048536645063&usg=AOvVaw1yDJ5thDXppIweq_TBfi2d) |
| 107 | [▪ VCAO: Verifier-Centered Agentic Orchestration for Strategic OS Vulnerability Discovery (Mishra, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08291&sa=D&source=editors&ust=1779048536645135&usg=AOvVaw3lzhUmZhaJoqW2MBEzGnqT) |
| 108 | [▪ ACF: A Collaborative Framework for Agent Covert Communication under Cognitive Asymmetry (Wu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08276&sa=D&source=editors&ust=1779048536645206&usg=AOvVaw1kD0VxRBD9kYDKOHtlSyb6) |
| 109 | [▪ ARuleCon: Agentic Security Rule Conversion (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06762&sa=D&source=editors&ust=1779048536645280&usg=AOvVaw13IJOxLTFc-CVeZ_waYnUt) |
| 110 | [▪ SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills (Hou and Yang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06550&sa=D&source=editors&ust=1779048536645357&usg=AOvVaw1qfCV23vdq-SMCIam4pkMk) |
| 111 | [▪ ClawLess: A Security Model of AI Agents (Lu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06284&sa=D&source=editors&ust=1779048536645441&usg=AOvVaw3M_by-2ZpTvDunPb7WcxqU) |
| 112 | [▪ ZitPit: Consumer-Side Admission Control for Agentic Software Intake (Taylor et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06241&sa=D&source=editors&ust=1779048536645515&usg=AOvVaw3AZD-ZHfPBm2wx-JKfKr91) |
| 113 | [▪ Foundations for Agentic AI Investigations from the Forensic Analysis of OpenClaw (Gruber and Hilgert, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05589&sa=D&source=editors&ust=1779048536645593&usg=AOvVaw27AeemvMxJNAvPjA4sX9WM) |
| 114 | [▪ LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations (Abdennebi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05440&sa=D&source=editors&ust=1779048536645664&usg=AOvVaw0yruj7ppPgQRI7XurNc9ma) |
| 115 | [▪ Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04465&sa=D&source=editors&ust=1779048536645739&usg=AOvVaw0rILVJkMSbbcI0E43Dn96d) |
| 116 | [▪ Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04759&sa=D&source=editors&ust=1779048536645821&usg=AOvVaw18EA4SnT-OMKzCeu_IuuMQ) |
| 117 | [▪ Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange (Vaikuntanathan and Zamir, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04757&sa=D&source=editors&ust=1779048536645894&usg=AOvVaw0T8WFd4a4uMPu1czWDRpYd) |
| 118 | [▪ HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems (Dalugoda, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04522&sa=D&source=editors&ust=1779048536645979&usg=AOvVaw25PPD4b0TeaY3WRiSx1pa-) |
| 119 | [▪ Governance-Constrained Agentic AI: Blockchain-Enforced Human Oversight for Safety-Critical Wildfire Monitoring (Akarma et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04265&sa=D&source=editors&ust=1779048536646053&usg=AOvVaw2GKcIWuurcZVB9bMauu_dT) |
| 120 | [▪ AutoVerifier: An Agentic Automated Verification Framework Using Large Language Models (Du et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02617&sa=D&source=editors&ust=1779048536646125&usg=AOvVaw12_4qrkKF0V3kjdM75F2Kb) |
| 121 | [▪ Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02837&sa=D&source=editors&ust=1779048536646193&usg=AOvVaw1gp78MYtGu2Xc0CXQj1Y4u) |
| 122 | [▪ SentinelAgent: Intent-Verified Delegation Chains for Securing Federal Multi-Agent AI Systems (Patil, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02767&sa=D&source=editors&ust=1779048536646264&usg=AOvVaw2pvh2R8kuSLjp3YWQ1sC_a) |
| 123 | [▪ Type-Checked Compliance: Deterministic Guardrails for Agentic Financial Systems Using Lean 4 Theorem Proving (Rashie and Rashi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01483&sa=D&source=editors&ust=1779048536646334&usg=AOvVaw1QrrQCxj9BRydWhRuhEWGD) |
| 124 | [▪ APEX: Agent Payment Execution with Policy for Autonomous Agent API Access (Uddin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02023&sa=D&source=editors&ust=1779048536646404&usg=AOvVaw25OQbvUpn58FOydU2tORIY) |
| 125 | [▪ Multi-Agent LLM Governance for Safe Two-Timescale Reinforcement Learning in SDN-IoT Defense (Jamshidi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01127&sa=D&source=editors&ust=1779048536646479&usg=AOvVaw1_RlwwboMjeWPTUrGNIS2X) |
| 126 | [▪ AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01014&sa=D&source=editors&ust=1779048536646549&usg=AOvVaw1MW2dIWL-JemX8LdC-J6C0) |
| 127 | [▪ Do Phone-Use Agents Respect Your Privacy? (Tang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00986&sa=D&source=editors&ust=1779048536646618&usg=AOvVaw0T6yOyxznIQOKCFSCrk7td) |
| 128 | [▪ SafeClaw-R: Towards Safe and Secure Multi-Agent Personal Assistants (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28807&sa=D&source=editors&ust=1779048536646688&usg=AOvVaw0VSJJXuMZZISGJ5XBXUro0) |
| 129 | [▪ Multi-Agent Actor-Critics in Autonomous Cyber Defense (Wang and Dechene, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.09134&sa=D&source=editors&ust=1779048536646757&usg=AOvVaw1iZcVNzqEBaYC6HSmxAYTl) |
| 130 | [▪ "What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents (Peng, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28551&sa=D&source=editors&ust=1779048536646829&usg=AOvVaw33Z71w901jmRufAK2ZdFyi) |
| 131 | [▪ Evaluating Privilege Usage of Agents on Real-World Tools (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28166&sa=D&source=editors&ust=1779048536646897&usg=AOvVaw3yavnxNW_YJA9f_Kj7YkEA) |
| 132 | [▪ Red-MIRROR: Agentic LLM-based Autonomous Penetration Testing with Reflective Verification and Knowledge-augmented Interaction (Khang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27127&sa=D&source=editors&ust=1779048536646973&usg=AOvVaw3KqKhtq2J_YY8Q2zftmCQe) |
| 133 | [▪ Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization (Kong et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26270&sa=D&source=editors&ust=1779048536647044&usg=AOvVaw3kgtnxNY1InmQIdOH3SkQh) |
| 134 | [▪ Clawed and Dangerous: Can We Trust Open Agentic Systems? (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26221&sa=D&source=editors&ust=1779048536647112&usg=AOvVaw0SDsrxXeOC3RNJrZK0zC3t) |
| 135 | [▪ AIP: Agent Identity Protocol for Verifiable Delegation Across MCP and A2A (Prakash, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24775&sa=D&source=editors&ust=1779048536647181&usg=AOvVaw0Fztajye5uq0EpxyeGI_t1) |
| 136 | [▪ Infrastructure for Valuable, Tradable, and Verifiable Agent Memory (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24564&sa=D&source=editors&ust=1779048536647250&usg=AOvVaw2H8UMGXftcb1W_gtq4ptya) |
| 137 | [▪ AgentRFC: Security Design Principles and Conformance Testing for Agent Protocols (Zheng and Zhang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23801&sa=D&source=editors&ust=1779048536647320&usg=AOvVaw0yr9CPrMXoTU1Kfv2XHalj) |
| 138 | [▪ SoK: The Attack Surface of Agentic AI -- Tools, and Autonomy (Dehghantanha and Homayoun, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22928&sa=D&source=editors&ust=1779048536647391&usg=AOvVaw3CqbD91r1yLRqmTudIx0qw) |
| 139 | [▪ Agent-Sentry: Bounding LLM Agents via Execution Provenance (Sequeira et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22868&sa=D&source=editors&ust=1779048536647465&usg=AOvVaw1XCVZsKu5EJAIGYP2RTK80) |
| 140 | [▪ STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving (Hugglestone et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22577&sa=D&source=editors&ust=1779048536647537&usg=AOvVaw1tY0lL9_oPXD54kaUHYtUz) |
| 141 | [▪ When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents (Lu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21231&sa=D&source=editors&ust=1779048536647610&usg=AOvVaw3Pa-YZT6MFUYMDt5VMv6P1) |
| 142 | [▪ Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions (Xiang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21194&sa=D&source=editors&ust=1779048536647680&usg=AOvVaw20t2lwg2J0pNEv7prmyqdd) |
| 143 | [▪ Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents (Uchibeke, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20953&sa=D&source=editors&ust=1779048536647749&usg=AOvVaw03OZKLZK1MIApEbSNknK8c) |
| 144 | [▪ AC4A: Access Control for Agents (Sharma and Grossman, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20933&sa=D&source=editors&ust=1779048536647819&usg=AOvVaw02Qi2c7ZoSLC2XUe0YSsZC) |
| 145 | [▪ MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20351&sa=D&source=editors&ust=1779048536647891&usg=AOvVaw19uXbyXOI8JnB7P39Ekn_S) |
| 146 | [▪ Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20198&sa=D&source=editors&ust=1779048536647969&usg=AOvVaw3IA9HgVkZZuKeEb1hVtu3V) |
| 147 | [▪ An Agentic Multi-Agent Architecture for Cybersecurity Risk Management (Gupta et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20131&sa=D&source=editors&ust=1779048536648051&usg=AOvVaw0k6-Nq_yuDUkdKWbRNqn2n) |
| 148 | [▪ The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents (Sah et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19328&sa=D&source=editors&ust=1779048536648126&usg=AOvVaw1FPb5QfJKSJ707DJSeOpcr) |
| 149 | [▪ A402: Binding Cryptocurrency Payments to Service Execution for Agentic Commerce (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01179&sa=D&source=editors&ust=1779048536648197&usg=AOvVaw2cc-0yDgJne1zLFA37mk7K) |
| 150 | [▪ Access Controlled Website Interaction for Agentic AI with Delegated Critical Tasks (Kim and Kim, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18197&sa=D&source=editors&ust=1779048536648268&usg=AOvVaw3r8_LN4BzVp_pgKwDb8t2_) |
| 151 | [▪ Security, privacy, and agentic AI in a regulatory view: From definitions and distinctions to provisions and reflections (Zhang and Maharjan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18914&sa=D&source=editors&ust=1779048536648341&usg=AOvVaw1cEoC6p3bIK-dCnL7YwhVE) |
| 152 | [▪ Agent Control Protocol: Admission Control for Agent Actions (Fernandez, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18829&sa=D&source=editors&ust=1779048536648416&usg=AOvVaw0WbnVOrRKXG9_9hobB_LwQ) |
| 153 | [▪ Capability-Priced Micro-Markets: A Micro-Economic Framework for the Agentic Web over HTTP 402 (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16899&sa=D&source=editors&ust=1779048536648488&usg=AOvVaw1t93BUCYAsOSES3yEnZXn6) |
| 154 | [▪ Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare (Maiti, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17419&sa=D&source=editors&ust=1779048536648559&usg=AOvVaw0AyEnrOjfHryJC3cqWT6hT) |
| 155 | [▪ LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems (Atta et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17239&sa=D&source=editors&ust=1779048536648630&usg=AOvVaw0MTsYFAj139FTWXbZeFlOH) |
| 156 | [▪ PAuth - Precise Task-Scoped Authorization For Agents (Sharma et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17170&sa=D&source=editors&ust=1779048536648698&usg=AOvVaw0NJwA080E3wkX5VFe5bGfF) |
| 157 | [▪ Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback (Jiralerspong, Kondrup, and Bengio, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16928&sa=D&source=editors&ust=1779048536648769&usg=AOvVaw1LX4MsBjOX1nNXvmUJoq7P) |
| 158 | [▪ ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15727&sa=D&source=editors&ust=1779048536648837&usg=AOvVaw102sE_u2AHvQjoZ28noAqf) |
| 159 | [▪ ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems (Chitan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13247&sa=D&source=editors&ust=1779048536648905&usg=AOvVaw1gaZ-nq3rtahdCXwjTDqws) |
| 160 | [▪ From Storage to Steering: Memory Control Flow Attacks on LLM Agents (Xu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15125&sa=D&source=editors&ust=1779048536648976&usg=AOvVaw2tEQmbnH0JP0rXWvJfF5Hb) |
| 161 | [▪ Towards Agentic Honeynet Configuration (Mirra et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14122&sa=D&source=editors&ust=1779048536649043&usg=AOvVaw0ukSAQ7hAioXxNN8qwAEAU) |
| 162 | [▪ Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline (Yuan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14011&sa=D&source=editors&ust=1779048536649116&usg=AOvVaw0q83DY-yP1EUT6YRZ-m8-a) |
| 163 | [▪ Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents (Li, Li, and Li, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13151&sa=D&source=editors&ust=1779048536649184&usg=AOvVaw05Nt0u6RE-T6BZ47qhXsPf) |
| 164 | [▪ Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats (Deng et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11619&sa=D&source=editors&ust=1779048536649254&usg=AOvVaw1-Jylp-kHSskfh_pOwCJpC) |
| 165 | [▪ The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey (Kim et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11088&sa=D&source=editors&ust=1779048536649333&usg=AOvVaw0Bm7llxmHq0cFkqMoYFs9x) |
| 166 | [▪ Re-Evaluating EVMBench: Are AI Agents Ready for Smart Contract Security? (Peng, Wu, and Zhou, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10795&sa=D&source=editors&ust=1779048536649423&usg=AOvVaw170bJkwvh6kf4Uq2Kw8RkH) |
| 167 | [▪ Execution Is the New Attack Surface: Survivability-Aware Agentic Crypto Trading with OpenClaw-Style Local Executors (Borjigin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10092&sa=D&source=editors&ust=1779048536649513&usg=AOvVaw3830AEedZ7noHaH9Aa5OjS) |
| 168 | [▪ SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation (Radanliev et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10057&sa=D&source=editors&ust=1779048536649589&usg=AOvVaw3oPEvITUHb3IXpZjE7eRlF) |
| 169 | [▪ ProvAgent: Threat Detection Based on Identity-Behavior Binding and Multi-Agent Collaborative Attack Investigation (Yan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09358&sa=D&source=editors&ust=1779048536649671&usg=AOvVaw3A18ul41sZEJ_iGj8ItHDw) |
| 170 | [▪ AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations (Mitra et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09134&sa=D&source=editors&ust=1779048536649744&usg=AOvVaw3LDutHQ0XU0aXwZ9-R0i9S) |
| 171 | [▪ Security Considerations for Multi-agent Systems (Nguyen, Ndebugre, and Arremsetty, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09002&sa=D&source=editors&ust=1779048536649815&usg=AOvVaw0OqaeDRTAJnGi8IyrVQRiU) |
| 172 | [▪ SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions (Mishra et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07379&sa=D&source=editors&ust=1779048536649887&usg=AOvVaw0y8_ZFCWDnN-lnbxOpgPTS) |
| 173 | [▪ From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07496&sa=D&source=editors&ust=1779048536649985&usg=AOvVaw0S5Z6ZOsobsyqQZZx5N1VG) |
| 174 | [▪ Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice (Ge, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07191&sa=D&source=editors&ust=1779048536650061&usg=AOvVaw3BZ3oaRNNtnG4XFT-sVg4r) |
| 175 | [▪ Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems (Asif and Amiri, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05520&sa=D&source=editors&ust=1779048536650133&usg=AOvVaw2CB6qhC59k4E63S0YGbFTv) |
| 176 | [▪ Evolving Deception: When Agents Evolve, Deception Wins (Ying et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05872&sa=D&source=editors&ust=1779048536650203&usg=AOvVaw27qlF0nOr_-rSCYj4vCwzY) |
| 177 | [▪ CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics (Fumero et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20643&sa=D&source=editors&ust=1779048536650272&usg=AOvVaw1NOzuHnyrW-Di6fIv3tNsN) |
| 178 | [▪ AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows (Ngong et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04902&sa=D&source=editors&ust=1779048536650341&usg=AOvVaw0PtTdmcoxfx9XPc2thI_-I) |
| 179 | [▪ Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization (Mumcu and Yilmaz, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04378&sa=D&source=editors&ust=1779048536650418&usg=AOvVaw0zSNepID6ICPAQOQtWTNbL) |
| 180 | [▪ On the Suitability of LLM-Driven Agents for Dark Pattern Audits (Sun, Vekaria, and Nithyanand, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03881&sa=D&source=editors&ust=1779048536650489&usg=AOvVaw0ejONrfcSEw_FQOFrs9iym) |
| 181 | [▪ From Secure Agentic AI to Secure Agentic Web: Challenges, Threats, and Future Directions (Deng, Gui, and Zhang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01564&sa=D&source=editors&ust=1779048536650560&usg=AOvVaw1rfyVXpz6nplbkm2byZnNg) |
| 182 | [▪ Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents (Jiang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00476&sa=D&source=editors&ust=1779048536650630&usg=AOvVaw0bBnN68HS64lh9YI-xPcNF) |
| 183 | [▪ LiaisonAgent: An Multi-Agent Framework for Autonomous Risk Investigation and Governance (Tang, Qing, and Chen, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00200&sa=D&source=editors&ust=1779048536650700&usg=AOvVaw2BOf5q0kBZCz46rEbrUR1v) |
| 184 | [▪ Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG (Singh et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21447&sa=D&source=editors&ust=1779048536650770&usg=AOvVaw18H8eQ-juRImKWMSEScaPr) |
| 185 | [▪ "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21127&sa=D&source=editors&ust=1779048536650842&usg=AOvVaw1x6ydJNb7zjczf8P4v6exp) |
| 186 | [▪ SoK: Agentic Skills -- Beyond Tool Use in LLM Agents (Jiang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20867&sa=D&source=editors&ust=1779048536650912&usg=AOvVaw04CPqAdDLCOjaxQgyTOmht) |
| 187 | [▪ The LLMbda Calculus: AI Agents, Conversations, and Information Flow (Garby, Gordon, and Sands, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20064&sa=D&source=editors&ust=1779048536651002&usg=AOvVaw1MzcwC0tM_00RcpEN4KEeb) |
| 188 | [▪ Agentic AI as a Cybersecurity Attack Surface: Threats, Exploits, and Defenses in Runtime Supply Chains (Jiang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19555&sa=D&source=editors&ust=1779048536651083&usg=AOvVaw3eR-nmMYbWowbpF9nDNE0w) |
| 189 | [▪ Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study (Mehta, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19514&sa=D&source=editors&ust=1779048536651153&usg=AOvVaw0-N1Eu0kkkIdgIHLssgX20) |
| 190 | [▪ OpenSage: Self-programming Agent Generation Engine (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16891&sa=D&source=editors&ust=1779048536651222&usg=AOvVaw31ZjhUrV4nm83Uxp_z-x4p) |
| 191 | [▪ Policy Compiler for Secure Agentic Systems (Palumbo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16708&sa=D&source=editors&ust=1779048536651292&usg=AOvVaw0eUXxsHpGQfCZI49AgpTF2) |
| 192 | [▪ Intellicise Wireless Networks Meet Agentic AI: A Security and Privacy Perspective (Meng et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15290&sa=D&source=editors&ust=1779048536651362&usg=AOvVaw1RFepVYAV5bghQnambhsd-) |
| 193 | [▪ Overthinking Loops in Agents: A Structural Risk via MCP Tools (Lee et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14798&sa=D&source=editors&ust=1779048536651437&usg=AOvVaw16Wcb_F6DM0vozFRnwVK8k) |
| 194 | [▪ AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports (Sajadi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14345&sa=D&source=editors&ust=1779048536651509&usg=AOvVaw0FqtrpO_5xQ7pZ4d8ifdCv) |
| 195 | [▪ In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach (Gao, Hammar, and Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13156&sa=D&source=editors&ust=1779048536651606&usg=AOvVaw0lCiRJ_tiXTmy2BB19-Wo2) |
| 196 | [▪ Agentic AI for Cybersecurity: A Meta-Cognitive Architecture for Governable Autonomy (Kojukhov and Bovshover, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11897&sa=D&source=editors&ust=1779048536651677&usg=AOvVaw0KzHb2U9fRWnrQUXUZBttX) |
| 197 | [▪ Optimizing Agent Planning for Security and Autonomy (Kolluri et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11416&sa=D&source=editors&ust=1779048536651752&usg=AOvVaw1ngzSVx7TySOoySZ9HRVcS) |
| 198 | [▪ Agentic Knowledge Distillation: Autonomous Training of Small Language Models for SMS Threat Detection (ElZemity et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10869&sa=D&source=editors&ust=1779048536651830&usg=AOvVaw1Ko5-DdXxzr9RBNP2qeIYH) |
| 199 | [▪ Authenticated Workflows: A Systems Approach to Protecting Agentic AI (Rajagopalan and Rao, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10465&sa=D&source=editors&ust=1779048536651900&usg=AOvVaw2poLK1K0Lb8JcM6k4-ThDV) |
| 200 | [▪ Trustworthy Agentic AI Requires Deterministic Architectural Boundaries (Bhattarai and Vu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09947&sa=D&source=editors&ust=1779048536651977&usg=AOvVaw3aEMmBeXhkS35UMZDYPy8t) |
| 201 | [▪ Focus Session: LLM4PQC -- An Agentic Framework for Accurate and Efficient Synthesis of PQC Cores (Perera et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09919&sa=D&source=editors&ust=1779048536652048&usg=AOvVaw0M7mPuki24tNZjt4Eionyg) |
| 202 | [▪ Autonomous Action Runtime Management(AARM):A System Specification for Securing AI-Driven Actions at Runtime (Errico, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09433&sa=D&source=editors&ust=1779048536652118&usg=AOvVaw1HCWZ1ujWC2ig43hUmLIur) |
| 203 | [▪ Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents (Puppala et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07652&sa=D&source=editors&ust=1779048536652188&usg=AOvVaw3m47s8sjKAJX7sY4-YydEj) |
| 204 | [▪ AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management (Wen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07398&sa=D&source=editors&ust=1779048536652257&usg=AOvVaw0PeiRTLwXy_TGyVw6_oKyc) |
| 205 | [▪ Aegis: Towards Governance, Integrity, and Security of AI Voice Agents (Li, Chen, and Wei, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07379&sa=D&source=editors&ust=1779048536652327&usg=AOvVaw00Vm4qGf321-nuDe09yeOu) |
| 206 | [▪ Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction (Pu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07287&sa=D&source=editors&ust=1779048536652397&usg=AOvVaw1SmBx6Um4VOMsoO4NhdvJh) |
| 207 | [▪ Malicious Agent Skills in the Wild: A Large-Scale Security Empirical Study (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06547&sa=D&source=editors&ust=1779048536652474&usg=AOvVaw0ih5nnjBMsOejamlTBuadp) |
| 208 | [▪ Zero-Trust Runtime Verification for Agentic Payment Protocols: Mitigating Replay and Context-Binding Failures in AP2 (Lan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06345&sa=D&source=editors&ust=1779048536652547&usg=AOvVaw2IpKFcEbYrZQ_PqjzSU9Q7) |
| 209 | [▪ CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability (Luo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03012&sa=D&source=editors&ust=1779048536652616&usg=AOvVaw1y96fPhmaG_dgIGgqHpeFV) |
| 210 | [▪ Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework (Abuadbba et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01942&sa=D&source=editors&ust=1779048536652683&usg=AOvVaw1ITaCRJia657QoZSc71kVu) |
| 211 | [▪ Protocol Agent: What If Agents Could Use Cryptography In Everyday Life? (Rossi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01304&sa=D&source=editors&ust=1779048536652751&usg=AOvVaw2PparIsFyfmj7fWWCMrJ9S) |
| 212 | [▪ Semantic-Aware Advanced Persistent Threat Detection Using Autoencoders on LLM-Encoded System Logs (Mohammed et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00204&sa=D&source=editors&ust=1779048536652821&usg=AOvVaw3ArXZJome3dVDU2pEC9QJi) |
| 213 | [▪ StepShield: When, Not Whether to Intervene on Rogue Agents (Felicia et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22136&sa=D&source=editors&ust=1779048536652892&usg=AOvVaw09DsbPw6bs-Q9skX5alt6V) |
| 214 | [▪ AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection (Charoenwet et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19138&sa=D&source=editors&ust=1779048536652966&usg=AOvVaw3kpVODKvRwdwgDMHdbXcZZ) |
| 215 | [▪ Multi-Agent Collaborative Intrusion Detection for Low-Altitude Economy IoT: An LLM-Enhanced Agentic AI Framework (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17817&sa=D&source=editors&ust=1779048536653036&usg=AOvVaw3XGcUyHt7Bfnz5NayojZhs) |
| 216 | [▪ Secure Intellicise Wireless Network: Agentic AI for Coverless Semantic Steganography Communication (Meng et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16472&sa=D&source=editors&ust=1779048536653106&usg=AOvVaw1cSynBZtogOFDweEeF5bdq) |
| 217 | [▪ An LLM Agent-based Framework for Whaling Countermeasures (Miyamoto, Iimura, and Michishita, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14606&sa=D&source=editors&ust=1779048536653173&usg=AOvVaw2aXBiWxUykMNCufw9IX6rj) |
| 218 | [▪ AgenTRIM: Tool Risk Mitigation for Agentic AI (Betser et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12449&sa=D&source=editors&ust=1779048536653239&usg=AOvVaw3dmSJh6dE7y1DvbpJ83llJ) |
| 219 | [▪ Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents? (Qian et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12349&sa=D&source=editors&ust=1779048536653309&usg=AOvVaw2nvNRSKdQeyOw_f9MBCuFi) |
| 220 | [▪ Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework (Ji et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11893&sa=D&source=editors&ust=1779048536653380&usg=AOvVaw05rotZrduxjnJONZGxOG8U) |
| 221 | [▪ Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents (Zhou et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10955&sa=D&source=editors&ust=1779048536653475&usg=AOvVaw1DNBE-DUUfLOvzH3hXehq-) |
| 222 | [▪ Too Helpful to Be Safe: User-Mediated Attacks on Planning and Web-Use Agents (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10758&sa=D&source=editors&ust=1779048536653546&usg=AOvVaw3p23PWCBWVkCvo1RWQ7mcW) |
| 223 | [▪ AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior (Abaev et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10440&sa=D&source=editors&ust=1779048536653617&usg=AOvVaw0sKDihLnMkO7jPyDzfzUCr) |
| 224 | [▪ Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10338&sa=D&source=editors&ust=1779048536653693&usg=AOvVaw1OOnjf_b9OSDt0X5JIwkis) |
| 225 | [▪ Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility (Engelberg et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07880&sa=D&source=editors&ust=1779048536653765&usg=AOvVaw0okGCEB8A-uxuAmN5K18R8) |
| 226 | [▪ FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments (Yang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07853&sa=D&source=editors&ust=1779048536653833&usg=AOvVaw35nZ4lMZsXgHPziHD8_6UD) |
| 227 | [▪ Agentic AI Microservice Framework for Deepfake and Document Fraud Detection in KYC Pipelines (Kubam, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06241&sa=D&source=editors&ust=1779048536653903&usg=AOvVaw0aAII8tuPuRipHRDc6CGHM) |
| 228 | [▪ A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes (Lazer et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05293&sa=D&source=editors&ust=1779048536653979&usg=AOvVaw2hTrMKMS5YyWAbDGZt1Bpx) |
| 229 | [▪ Integrating Multi-Agent Simulation, Behavioral Forensics, and Trust-Aware Machine Learning for Adaptive Insider Threat Detection (Kausar et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04243&sa=D&source=editors&ust=1779048536654050&usg=AOvVaw3TszDTDd3Ahfqq42Os2mru) |
| 230 | [▪ Web Fraud Attacks Against LLM-Driven Multi-Agent Systems (Kong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.01211&sa=D&source=editors&ust=1779048536654118&usg=AOvVaw3PW50n3YLa2lrt6Dr2-HGV) |
| 231 | [▪ SastBench: A Benchmark for Testing Agentic SAST Triage (Feiglin and Dar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02941&sa=D&source=editors&ust=1779048536654186&usg=AOvVaw2ECcItXRbI29PuZAloc7sP) |
| 232 | [▪ MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI (Xing et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.10991&sa=D&source=editors&ust=1779048536654267&usg=AOvVaw0Za0RHKWm2IMeThTH66mf8) |
| 233 | [▪ Device-Native Autonomous Agents for Privacy-Preserving Negotiations (Roy, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00911&sa=D&source=editors&ust=1779048536654337&usg=AOvVaw2x8_-UDyU5KdkNltryQtXa) |
| 234 | [▪ The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models (Canale and Thimmaraju, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00867&sa=D&source=editors&ust=1779048536654420&usg=AOvVaw0mhF6jNbKdPOyouKIwfpdt) |
| 235 | [▪ Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub (Siddiq et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00477&sa=D&source=editors&ust=1779048536654496&usg=AOvVaw2GIQxeAI9177Cy8FLemlVJ) |
| 236 | [▪ Zero-Trust Agentic Federated Learning for Secure IIoT Defense Systems (Singh, Roy, and So, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23809&sa=D&source=editors&ust=1779048536654566&usg=AOvVaw3Lc9OexKrr8fibBZw0Yfg_) |
| 237 | [▪ Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory (Huang and Huang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23760&sa=D&source=editors&ust=1779048536654647&usg=AOvVaw2efJwv1k2uPAfX7Trvolrn) |
| 238 | [▪ Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation (Syed et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23480&sa=D&source=editors&ust=1779048536654731&usg=AOvVaw1_RMhTX_vdIldpnDbVwikv) |
| 239 | [▪ Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations (Li and Zhu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22883&sa=D&source=editors&ust=1779048536654803&usg=AOvVaw1P-erqa2fQqVFRgc3ItqGA) |
| 240 | [▪ Securing Agentic AI Systems -- A Multilayer Security Framework (Arora and Hastings, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18043&sa=D&source=editors&ust=1779048536654879&usg=AOvVaw1Q4J3Bg15RHvuAzFTdFT6v) |
| 241 | [▪ Binding Agent ID: Unleashing the Power of AI Agents with accountability and credibility (Lin et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17538&sa=D&source=editors&ust=1779048536654968&usg=AOvVaw2IB8eX4PuuypWp9Cpy5Aut) |
| 242 | [▪ Biosecurity-Aware AI: Agentic Risk Auditing of Soft Prompt Attacks on ESM-Based Variant Predictors (Zhan, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17146&sa=D&source=editors&ust=1779048536655049&usg=AOvVaw0MBgUmsgKMgPznJ63tvHpS) |
| 243 | [▪ BashArena: A Control Setting for Highly Privileged AI Agents (Kaufman et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15688&sa=D&source=editors&ust=1779048536655129&usg=AOvVaw0FoGxeypzvkPlGudQziAqK) |
| 244 | [▪ Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks (Nguyen and Husain, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14860&sa=D&source=editors&ust=1779048536655220&usg=AOvVaw2BsJHTsDfvL8wpas0O6Tk5) |
| 245 | [▪ Factor(U,T): Controlling Untrusted AI by Monitoring their Plans (Lip et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14745&sa=D&source=editors&ust=1779048536655297&usg=AOvVaw0YxIjfdMXQGLMkiEwyVRJZ) |
| 246 | [▪ ceLLMate: Sandboxing Browser AI Agents (Meng et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12594&sa=D&source=editors&ust=1779048536655371&usg=AOvVaw2-YgdIk6X_off-1NUI-xvg) |
| 247 | [▪ MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents (Zhu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11147&sa=D&source=editors&ust=1779048536655454&usg=AOvVaw35g7HlSszdfeZTiZuwHa3j) |
| 248 | [▪ Locus: Agentic Predicate Synthesis for Directed Fuzzing (Zhu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21302&sa=D&source=editors&ust=1779048536655522&usg=AOvVaw0kwOa-Yh4eFAUzlzyZdpv6) |
| 249 | [▪ AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration (Karthikeyan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08104&sa=D&source=editors&ust=1779048536655590&usg=AOvVaw2KKVj4sjjoDWfJF5HO5ynp) |
| 250 | [▪ Agentic Artificial Intelligence for Ethical Cybersecurity in Uganda: A Reinforcement Learning Framework for Threat Detection in Resource-Constrained Environments (Adabara et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07909&sa=D&source=editors&ust=1779048536655674&usg=AOvVaw3O9AFEGgXH1idsqffgVShn) |
| 251 | [▪ SoK: Trust-Authorization Mismatch in LLM Agent Interactions (Shi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06914&sa=D&source=editors&ust=1779048536655746&usg=AOvVaw1zJmlF2ZG8LCJinLbRz8fE) |
| 252 | [▪ The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines (Vinay, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06659&sa=D&source=editors&ust=1779048536655817&usg=AOvVaw1rIwm3ilkn8Nr6EcHiniSs) |
| 253 | [▪ AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity (Roy, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06396&sa=D&source=editors&ust=1779048536655887&usg=AOvVaw2bIcpJgObBw_4j2sgNNmAr) |
| 254 | [▪ Please Don't Kill My Vibe: Empowering Agents with Data Flow Control (Summers, Mohammed, and Wu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.05374&sa=D&source=editors&ust=1779048536655964&usg=AOvVaw3Glq7lRVPnx46_LWJEIeRH) |
| 255 | [▪ ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications (Bandara et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04785&sa=D&source=editors&ust=1779048536656033&usg=AOvVaw2JUS1-6gmfZrk4ZHY43tuU) |
| 256 | [▪ PBFuzz: Agentic Directed Fuzzing for PoV Generation (Zeng et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04611&sa=D&source=editors&ust=1779048536656129&usg=AOvVaw2C5moVGpJ8YgYzcYsBQibd) |
| 257 | [▪ Password-Activated Shutdown Protocols for Misaligned Frontier Agents (Williams, Subramani, and Ward, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03089&sa=D&source=editors&ust=1779048536656258&usg=AOvVaw3rYOnrPRXf43Kqo7SKFJEN) |
| 258 | [▪ LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02321&sa=D&source=editors&ust=1779048536656372&usg=AOvVaw130bX2GshSx06Q-SuRI1Ie) |
| 259 | [▪ Systems Security Foundations for Agentic Computing (Christodorescu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01295&sa=D&source=editors&ust=1779048536656483&usg=AOvVaw1SaVDvkWPWkFmqpDq8OOmE) |
| 260 | [▪ AgentShield: Make MAS more secure and efficient (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22924&sa=D&source=editors&ust=1779048536656577&usg=AOvVaw1t-GV3xBfiz5F4_vwmrkQY) |
| 261 | [▪ A Safety and Security Framework for Real-World Agentic Systems (Ghosh et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21990&sa=D&source=editors&ust=1779048536656674&usg=AOvVaw3WYG6o67PF_x3ADY0sptSZ) |
| 262 | [▪ FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18653&sa=D&source=editors&ust=1779048536656765&usg=AOvVaw2c4VOWQYpfaqaHLoOMqjk1) |
| 263 | [▪ LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18438&sa=D&source=editors&ust=1779048536656837&usg=AOvVaw1j2j7uFAQGs92LGqp9XEP7) |
| 264 | [▪ ASTRA: Agentic Steerability and Risk Assessment Framework (Hazan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18114&sa=D&source=editors&ust=1779048536656906&usg=AOvVaw0-JQd4eydWl-BILDEOlVL4) |
| 265 | [▪ Towards Automating Data Access Permissions in AI Agents (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17959&sa=D&source=editors&ust=1779048536656977&usg=AOvVaw0wnhFuKg-MqqWbl3-1wILL) |
| 266 | [▪ LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures (Hassouna, Chaari, and Belhaj, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11393&sa=D&source=editors&ust=1779048536657050&usg=AOvVaw3O2ODc1GdSvtjjB7ngKOQG) |
| 267 | [▪ Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming (Janjuesvic, Garcia, and Kazerounian, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15998&sa=D&source=editors&ust=1779048536657120&usg=AOvVaw1-_lsJbvnaLNiakjDIjmn2) |
| 268 | [▪ Secure Autonomous Agent Payments: Verifying Authenticity and Intent in a Trustless Environment (Acharya, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15712&sa=D&source=editors&ust=1779048536657187&usg=AOvVaw2idpxxDh5X1QlcnGl3FQ9M) |
| 269 | [▪ MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm (Narajala et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15097&sa=D&source=editors&ust=1779048536657257&usg=AOvVaw1RlOdNWJkFE3mNrpQYYtqd) |
| 270 | [▪ Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation (Zhao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11693&sa=D&source=editors&ust=1779048536657325&usg=AOvVaw364Ruve1FyI9SmRZFvpCZa) |
| 271 | [▪ Towards a Generalisable Cyber Defence Agent for Real-World Computer Networks (Dudman and Bull, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09114&sa=D&source=editors&ust=1779048536657394&usg=AOvVaw2tE6geTUtyVpU_Tuodga-b) |
| 272 | [▪ 3D Guard-Layer: An Integrated Agentic AI Safety System for Edge Artificial Intelligence (Kurshan, Xie, and Franzon, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08842&sa=D&source=editors&ust=1779048536657480&usg=AOvVaw3ZNCUarEczjTOsUv4dygdv) |
| 273 | [▪ From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection (Han et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08060&sa=D&source=editors&ust=1779048536657565&usg=AOvVaw2k2MSBzU3CHLNJOK_yn48A) |
| 274 | [▪ PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts (Andersson et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02780&sa=D&source=editors&ust=1779048536657636&usg=AOvVaw1k2NtXNEY4P1onZynx2kiw) |
| 275 | [▪ Security Analysis of Agentic AI Communication Protocols: A Comparative Evaluation (Louck, Stulman, and Dvir, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03841&sa=D&source=editors&ust=1779048536657705&usg=AOvVaw2l63vffp8LxEQ4dv0PFdkM) |
| 276 | [▪ 1 PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts (Andersson et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02780&sa=D&source=editors&ust=1779048536657773&usg=AOvVaw0NNclrqW32bMny_la2oUEz) |
| 277 | [▪ LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources (Ji et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18477&sa=D&source=editors&ust=1779048536657842&usg=AOvVaw2Rwz8SRO2kLec9aNtlbIIE) |
| 278 | [▪ AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI (Huang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25863&sa=D&source=editors&ust=1779048536657909&usg=AOvVaw3xQDR3GUIFqGGU_rb5mho2) |
| 279 | [▪ Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world (South et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25819&sa=D&source=editors&ust=1779048536657984&usg=AOvVaw2fzn8S0MzaOliseC-q5EM8) |
| 280 | [▪ AgentCyTE: Leveraging Agentic AI to Generate Cybersecurity Training & Experimentation Scenarios (Rodriguez et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25189&sa=D&source=editors&ust=1779048536658056&usg=AOvVaw1aFtmWb3vWc62RtCFVBw6W) |
| 281 | [▪ Securing AI Agent Execution (B\\"uhler et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21236&sa=D&source=editors&ust=1779048536658124&usg=AOvVaw0nG96y63U2N4nqxnVCq6Ha) |
| 282 | [▪ AI Agentic Vulnerability Injection And Transformation with Optimized Reasoning (Lbath et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20866&sa=D&source=editors&ust=1779048536658192&usg=AOvVaw0d5Y1gVOvszGxT6tiZWsBZ) |
| 283 | [▪ CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection (Trad and Chehab, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18585&sa=D&source=editors&ust=1779048536658259&usg=AOvVaw1pMXriXeXKE7ZicBxIFm9t) |
| 284 | [▪ The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability (Xu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18563&sa=D&source=editors&ust=1779048536658327&usg=AOvVaw35L2TtLcQBJ043t2yaVRlJ) |
| 285 | [▪ Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs (Balassone et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17521&sa=D&source=editors&ust=1779048536658395&usg=AOvVaw0Hzm6PGdOBKunTelZaBZTv) |
| 286 | [▪ Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies (Nakamura et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14312&sa=D&source=editors&ust=1779048536658467&usg=AOvVaw2cFblUE2qIikUaIAd12TJz) |
| 287 | [▪ Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems (Allegrini, Shreekumar, and Celik, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14133&sa=D&source=editors&ust=1779048536658535&usg=AOvVaw0UgrmPA49X_HFhDOXTBIez) |
| 288 | [▪ A2AS: Agentic AI Runtime Security and Self-Defense (Neelou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13825&sa=D&source=editors&ust=1779048536658603&usg=AOvVaw1l4zLwVhSzd7oJ000UykZe) |
| 289 | [▪ AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement (Rosser and Foerster, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.00757&sa=D&source=editors&ust=1779048536658672&usg=AOvVaw0QkR9NBHhruaze-_9q0IWi) |
| 290 | [▪ A Vision for Access Control in LLM-based Agent Systems (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11108&sa=D&source=editors&ust=1779048536658741&usg=AOvVaw3T5yGvZsLYvtHMN_QnTlrw) |
| 291 | [▪ Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model (Fleming, Kundu, and Kompella, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11414&sa=D&source=editors&ust=1779048536658813&usg=AOvVaw0rEAVP04j3eCL5XMF2bCnU) |
| 292 | [▪ Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols (Terekhov et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09462&sa=D&source=editors&ust=1779048536658882&usg=AOvVaw2A9nTbZL9nKZPA7n-g9_mi) |
| 293 | [▪ A Survey on Agentic Security: Applications, Threats and Defenses (Shahriar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06445&sa=D&source=editors&ust=1779048536658954&usg=AOvVaw1-8Sz32izxcU7Lw6LU4Ukc) |
| 294 | [▪ Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent (Luo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06607&sa=D&source=editors&ust=1779048536659024&usg=AOvVaw1Gjxv7C3PXnYj2hdJbvvbk) |
| 295 | [▪ AutoPentester: An LLM Agent-based Framework for Automated Pentesting (Ginige et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05605&sa=D&source=editors&ust=1779048536659093&usg=AOvVaw0V5pjtf-sNiIRdnHKJv3ks) |
| 296 | [▪ Adapting Insider Risk mitigations for Agentic Misalignment: an empirical study (Gomez, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05192&sa=D&source=editors&ust=1779048536659177&usg=AOvVaw1vAzSzNXOFIF2O1pUmOArL) |
| 297 | [▪ Agentic Misalignment: How LLMs Could Be Insider Threats (Lynch et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05179&sa=D&source=editors&ust=1779048536659249&usg=AOvVaw3h0N58X1ql2ImS8sMuK4dd) |
| 298 | [▪ Autonomy Matters: A Study on Personalization-Privacy Dilemma in LLM Agents (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04465&sa=D&source=editors&ust=1779048536659318&usg=AOvVaw0ceY0t7TfpA2q_gBwQOkEi) |
| 299 | [▪ Quantifying Distributional Robustness of Agentic Tool-Selection (Yeon, Chaudhary, and Singh, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03992&sa=D&source=editors&ust=1779048536659396&usg=AOvVaw1v3-QRswXcL3ywD4zc3LU0) |
| 300 | [▪ PentestMCP: A Toolkit for Agentic Penetration Testing (Ezetta and Feng, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03610&sa=D&source=editors&ust=1779048536659471&usg=AOvVaw3agprzdc59MbueYLzsKNRd) |
| 301 | [▪ MobiLLM: An Agentic AI Framework for Closed-Loop Threat Mitigation in 6G Open RANs (Sharma et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.21634&sa=D&source=editors&ust=1779048536659542&usg=AOvVaw22NNVcJDBBHu1ZbqphLPkT) |
| 302 | [▪ ToolTweak: An Attack on Tool Selection in LLM-based Agents (Sneh et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02554&sa=D&source=editors&ust=1779048536659610&usg=AOvVaw2CQKoeV-Ke4ehsu6W6HQRn) |
| 303 | [▪ Agentic-AI Healthcare: Multilingual, Privacy-First Framework with MCP Agents (Shehab, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02325&sa=D&source=editors&ust=1779048536659677&usg=AOvVaw0kkzCLWvh6Hv2REXdtdSxq) |
| 304 | [▪ FalseCrashReducer: Mitigating False Positive Crashes in OSS-Fuzz-Gen Using Agentic AI (Amusuo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02185&sa=D&source=editors&ust=1779048536659772&usg=AOvVaw1ny4CYw1C1AvyMqrRqjSVc) |
| 305 | [▪ Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness (Shayegani et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01670&sa=D&source=editors&ust=1779048536659897&usg=AOvVaw19KYsGOYSRfvOqFqJQCf-M) |
| 306 | [▪ Better Privilege Separation for Agents by Restricting Data Types (Jacob et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25926&sa=D&source=editors&ust=1779048536659996&usg=AOvVaw3NhO-grsgVdDZaczi_xMQL) |
| 307 | [▪ Agentic Specification Generator for Move Programs (Fu, Xu, and Kim, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24515&sa=D&source=editors&ust=1779048536660072&usg=AOvVaw3qatrDcGSKLFaAUan8p1vF) |
| 308 | [▪ LISA Technical Report: An Agentic Framework for Smart Contract Auditing (Sun, Tan, and Deng, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24698&sa=D&source=editors&ust=1779048536660165&usg=AOvVaw1oCUiv7Z3UHEk_FT6ILrI9) |
| 309 | [▪ DIRF: A Framework for Digital Identity Protection and Clone Governance in Agentic AI Systems (Atta et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01997&sa=D&source=editors&ust=1779048536660255&usg=AOvVaw0QCK27gKtY5XlO9NRUuLuy) |
| 310 | [▪ Measuring Harmfulness of Computer-Using Agents (Tian et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00935&sa=D&source=editors&ust=1779048536660330&usg=AOvVaw0d6AwTFpZ14tvUclv5Cgpk) |
| 311 | [▪ Beyond DNS: Unlocking the Internet of AI Agents via the NANDA Index and Verified AgentFacts (Raskar et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14263&sa=D&source=editors&ust=1779048536660429&usg=AOvVaw0wOS6wV5aCIyO9rCfveahk) |
| 312 | [▪ From Semantic Web and MAS to Agentic AI: A Unified Narrative of the Web of Agents (SnT et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10644&sa=D&source=editors&ust=1779048536660565&usg=AOvVaw2e0NxZHvD2rnRNRhxgqvPe) |
| 313 | [▪ Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats (Zhu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10621&sa=D&source=editors&ust=1779048536660664&usg=AOvVaw2gyptVKe8YryGuTm4Q4BEN) |
| 314 | [▪ Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm (Marino and Juels, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08249&sa=D&source=editors&ust=1779048536660765&usg=AOvVaw2QFONsf7zvMdmBMxWYk-2h) |
| 315 | [▪ The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover (Lupinacci et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06850&sa=D&source=editors&ust=1779048536660842&usg=AOvVaw3Tx1WEK1EzMPQR-JNsBg-W) |
| 316 | [▪ The Trust Fabric: Decentralized Interoperability and Economic Coordination for the Agentic Web (Balija et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07901&sa=D&source=editors&ust=1779048536660917&usg=AOvVaw2m6l8QwFJWvk7ena99Idid) |
| 317 | [▪ The Dark Side of LLMs Agent-based Attacks for Complete Computer Takeover (Lupinacci et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06850&sa=D&source=editors&ust=1779048536661000&usg=AOvVaw0RvGa9ok-IEcxhZgmLzVcO) |
| 318 | [▪ A Systematization of Security Vulnerabilities in Computer Use Agents (Jones et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05445&sa=D&source=editors&ust=1779048536661076&usg=AOvVaw28vjhFvMy7J_7AgXzhfvQM) |
| 319 | [▪ CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution\<br>\<br> (Shao et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17107&sa=D&source=editors&ust=1779048536661189&usg=AOvVaw0O-s1Vd1JGRcnKNysP6CzK) |
| 320 | [▪ The Odyssey of the Fittest: Can Agents Survive and Still Be Good?\<br>\<br> (Waldner and Miikkulainen, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05442&sa=D&source=editors&ust=1779048536661368&usg=AOvVaw2sNyy-YTHGR66cRTvNT6Z1) |
| 321 | [▪ Agency Is Frame-Dependent (Abel et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.04403&sa=D&source=editors&ust=1779048536661465&usg=AOvVaw2J2Yqxy07Nfuyva7NFEMvo) |
| 322 | [▪ Security of AI Agents (He et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.08689&sa=D&source=editors&ust=1779048536661577&usg=AOvVaw2LYcIp1xbtB2BhG-LSC8n2) |
| 323 | [▪ Dissecting Adversarial Robustness of Multimodal LM Agents (Wu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.12814&sa=D&source=editors&ust=1779048536661710&usg=AOvVaw0hxlMewK2QbUYnXlpIUG7c) |
| 324 | [▪ RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors (Bai et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.10713&sa=D&source=editors&ust=1779048536661848&usg=AOvVaw2Y9Cju59Rgm1wNuLGUxjzT) |
| 325 | [▪ AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (Debenedetti et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.13352&sa=D&source=editors&ust=1779048536661955&usg=AOvVaw0u3xyCCMqP3bVEtwv12E2N) |
| 326 | [▪ AdvWeb: Controllable Black-box Attacks on VLM-powered Web Agents (Xu et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.17401&sa=D&source=editors&ust=1779048536662087&usg=AOvVaw03c2GQYVe1-yPmt1rYzC-j) |
| 327 | [▪ Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation (Kwartler et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.14262&sa=D&source=editors&ust=1779048536662180&usg=AOvVaw0OgHICl6Kl-JmeJjNsqn5T) |
| 328 | [▪ Bayes-Nash Generative Privacy Protection Against Membership Inference Attacks (Zhang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07414&sa=D&source=editors&ust=1779048536662264&usg=AOvVaw2wB3wgPS76zQnh8gkVaA53) |
| 329 | [▪ Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems (Lee and Tiwari, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07283&sa=D&source=editors&ust=1779048536662345&usg=AOvVaw1ykD8i5cff5ZGqIMpm2mNl) |
| 330 | [▪ Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents (Zhang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02644&sa=D&source=editors&ust=1779048536662444&usg=AOvVaw04j45mOlJ72cr9Kkl_psde) |
| 331 | [▪ BreachSeek: A Multi-Agent Automated Penetration Tester (Alshehri et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03789&sa=D&source=editors&ust=1779048536662535&usg=AOvVaw0KmPqhYbK7VNJOlpEtNC61) |
| 332 | [▪ Safeguarding AI Agents: Developing and Analyzing Safety Architectures (Domkunwar and N S, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.03793&sa=D&source=editors&ust=1779048536662620&usg=AOvVaw0o46241JhBu5XnTILVTuMT) |
| 333 | [▪ Secret Collusion among Generative AI Agents (Motwani et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.07510&sa=D&source=editors&ust=1779048536662698&usg=AOvVaw2NUfqY0cH3aZgDFZbCAnqF) |
| 334 | [▪ Large Language Model Sentinel: LLM Agent for Adversarial Purification (Lin and Zhao, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20770&sa=D&source=editors&ust=1779048536662776&usg=AOvVaw0drdFO6TEZI_CriKnokmZ9) |
| 335 | [▪ Compromising Embodied Agents with Contextual Backdoor Attacks (Liu et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02882&sa=D&source=editors&ust=1779048536662853&usg=AOvVaw25n4qPIyr-x-WRJN1W21Ws) |
| 336 | [▪ Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification (Zhang et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20859&sa=D&source=editors&ust=1779048536662937&usg=AOvVaw0Cgp4bP1OhE9SS50h0ogbb) |

|     |
| --- |
| Excessive Agency, Agentic Manipulation, Agentic Systems |

**>**

**<**

‍

#### Copyright Infringement

Covers:

- MITRE ATLAS Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Copyright Infringement

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ From Compression to Accountability: Harmless Copyright Protection for Dataset Distillation (Liang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12942&sa=D&source=editors&ust=1779048536747575&usg=AOvVaw2TcGi8IUonPBSDDMBiK0S2) |
| 3 | [▪ Position: No Retroactive Cure for Infringement during Training (Utsunomiya et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18649&sa=D&source=editors&ust=1779048536747795&usg=AOvVaw2Kr5PFPJn8E4H6Gs5JULAf) |
| 4 | [▪ MATRIX: Multi-Layer Code Watermarking via Dual-Channel Constrained Parity-Check Encoding (Nie et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16001&sa=D&source=editors&ust=1779048536747907&usg=AOvVaw2ASMKvFcRCy-NxCcruScEH) |
| 5 | [▪ Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11344&sa=D&source=editors&ust=1779048536748052&usg=AOvVaw2rCpqZq3BOMZopdkRpb4S4) |
| 6 | [▪ Towards trustworthy management of AIGC copyright: blockchain-enabled full lifecycle recording and multi-party auditing approach (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2406.14966&sa=D&source=editors&ust=1779048536748184&usg=AOvVaw1n_42_GfMyez63WWLY2kH9) |
| 7 | [▪ Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.11548&sa=D&source=editors&ust=1779048536748329&usg=AOvVaw2e63UtDziECVCp2ui865c5) |
| 8 | [▪ A PUF-Based Approach for Copy Protection of Intellectual Property in Neural Network Models (Dorfmeister et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10753&sa=D&source=editors&ust=1779048536748510&usg=AOvVaw3enLS-rhY3xSaPnK2g6-yj) |
| 9 | [▪ Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment (Chen, Lou, and Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10030&sa=D&source=editors&ust=1779048536748728&usg=AOvVaw3QgtXWmffOTuCYsH93CezL) |
| 10 | [▪ DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing (Ren et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13541&sa=D&source=editors&ust=1779048536748853&usg=AOvVaw0yan3jdpc7CFMzf_iMpRTM) |
| 11 | [▪ Traceable Black-box Watermarks for Federated Learning (Xu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.13651&sa=D&source=editors&ust=1779048536748965&usg=AOvVaw2yB-ID9JGrWlhm6GTqza4i) |
| 12 | [▪ SEW: Strengthening Robustness of Black-box DNN Watermarking via Specificity Enhancement (Qiu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03377&sa=D&source=editors&ust=1779048536749073&usg=AOvVaw1dvAB6oR_i1BnQW-m0PzDq) |
| 13 | [▪ Lossless Copyright Protection via Intrinsic Model Fingerprinting (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21252&sa=D&source=editors&ust=1779048536749190&usg=AOvVaw3DrESmmtB3VRPVzjWn94zV) |
| 14 | [▪ On the Evidentiary Limits of Membership Inference for Copyright Auditing (Ertan et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12937&sa=D&source=editors&ust=1779048536749270&usg=AOvVaw17I1nYy0GGRkZSVjWVAegD) |
| 15 | [▪ DNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08223&sa=D&source=editors&ust=1779048536749374&usg=AOvVaw132FtaNXVXBRzIX3OcPSAk) |
| 16 | [▪ Multi-Agent Framework for Controllable and Protected Generative Content Creation: Addressing Copyright and Provenance in AI-Generated Media (Khan, Asif, and Asif, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06232&sa=D&source=editors&ust=1779048536749480&usg=AOvVaw2Pszgy6J2ftEl0iiKbuSMj) |
| 17 | [▪ SRAF: Stealthy and Robust Adversarial Fingerprint for Copyright Verification of Large Language Models (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.06304&sa=D&source=editors&ust=1779048536749596&usg=AOvVaw3hKLWcLzFP6Be295qq6Cf4) |
| 18 | [▪ SWaRL: Safeguard Code Watermarking via Reinforcement Learning (Javidnia et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02602&sa=D&source=editors&ust=1779048536749689&usg=AOvVaw0ToGqIuALt8bQk-WNW_4QD) |
| 19 | [▪ Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content? (Xu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21871&sa=D&source=editors&ust=1779048536749776&usg=AOvVaw1eABs4Di7XoQ7tQzmvaFUO) |
| 20 | [▪ Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models (Gao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.02824&sa=D&source=editors&ust=1779048536749871&usg=AOvVaw0TTwYqL5qc3V8ZvO2agD9n) |
| 21 | [▪ From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16439&sa=D&source=editors&ust=1779048536750005&usg=AOvVaw1ku2wr_nYjI0HT9MDW0Afs) |
| 22 | [▪ VICTOR: Dataset Copyright Auditing in Video Recognition Systems (Yuan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14439&sa=D&source=editors&ust=1779048536750143&usg=AOvVaw1vPq_eZAXUKGXe-Yj-5lPD) |
| 23 | [▪ Blameless Users in a Clean Room: Defining Copyright Protection for Generative Models (Cohen, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.19881&sa=D&source=editors&ust=1779048536750237&usg=AOvVaw1ib4UK4WMGS6oaoNSrkyQh) |
| 24 | [▪ Copyright in AI Pre-Training Data Filtering: Regulatory Landscape and Mitigation Strategies (Kyrychenko, Mudryi, and Chaklosh, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02047&sa=D&source=editors&ust=1779048536750324&usg=AOvVaw2IUGpPAFnsheKAXvW-X8gf) |
| 25 | [▪ EnTruth: Enhancing the Traceability of Unauthorized Dataset Usage in Text-to-image Diffusion Models with Minimal and Robust Alterations (Ren et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.13933&sa=D&source=editors&ust=1779048536750420&usg=AOvVaw2uIX6QRDyWNmnZw4S6sczQ) |
| 26 | [▪ AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters (Shi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21216&sa=D&source=editors&ust=1779048536750521&usg=AOvVaw2e_tk6B_mW2QykF6YMsTzQ) |
| 27 | [▪ PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.03117&sa=D&source=editors&ust=1779048536750638&usg=AOvVaw3Rbygas-7uVh9Bj7x5fjKz) |
| 28 | [▪ SoK: Large Language Model Copyright Auditing via Fingerprinting (Shao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.19843&sa=D&source=editors&ust=1779048536750750&usg=AOvVaw0medG36R671_-oxvm7JdP4) |
| 29 | [▪ RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13329&sa=D&source=editors&ust=1779048536750839&usg=AOvVaw0g8VRqjOBJfJ8mlMRClMdo) |
| 30 | [▪ SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04711&sa=D&source=editors&ust=1779048536750941&usg=AOvVaw2jm3u6QL9CjK7QNIzwQGaM) |
| 31 | [▪ SLIP: Securing LLMs IP Using Weights Decomposition (Refael et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.10886&sa=D&source=editors&ust=1779048536751055&usg=AOvVaw3P46HxUsW9hcnkvKU8Y_ZK) |
| 32 | [▪ Large Language Models Are Effective Code Watermarkers (Xu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11251&sa=D&source=editors&ust=1779048536751172&usg=AOvVaw3TcGoUsp99Yoo1RRRfYO5j) |
| 33 | [▪ Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.15740&sa=D&source=editors&ust=1779048536751267&usg=AOvVaw0rlscXlAoSOSEGbGLUJs_F) |
| 34 | [▪ Content ARCs: Decentralized Content Rights in the Age of Generative AI (Balan, Gilbert, Collomosse, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.14519&sa=D&source=editors&ust=1779048536751355&usg=AOvVaw1AP_IE3Kv2yUYO86qWjlI0) |
| 35 | [▪ SoK: Dataset Copyright Auditing in Machine Learning Systems (Du et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.16618&sa=D&source=editors&ust=1779048536751462&usg=AOvVaw1nwb2LwsGvQyA-dvECWgvI) |
| 36 | [▪ Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Watermarking (Kong et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10570&sa=D&source=editors&ust=1779048536751593&usg=AOvVaw0q2ZP2Q8VAVXr-5sLSayQQ) |
| 37 | [▪ Strong Copyright Protection for Language Models via Adaptive Model Fusion (Wang et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20105&sa=D&source=editors&ust=1779048536751717&usg=AOvVaw0axcsq9YP1hYJJEANN-0OM) |

|     |
| --- |
| Copyright Infringement |

**>**

**<**

‍

#### Defenses

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Defenses

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Toward Securing AI Agents Like Operating Systems (Pirch et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14932&sa=D&source=editors&ust=1779048536910284&usg=AOvVaw3g-wOG9e8GeL-Ai1Bdvj2D) |
| 3 | [▪ One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries (Zloczower et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14605&sa=D&source=editors&ust=1779048536910481&usg=AOvVaw3YggZRJHKglGAqMC3lFQx9) |
| 4 | [▪ Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models (Meng et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14514&sa=D&source=editors&ust=1779048536910581&usg=AOvVaw2-vr4xl-elrAfiOqZXPWat) |
| 5 | [▪ MemLineage: Lineage-Guided Enforcement for LLM Agent Memory (Ouyang and Hou, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14421&sa=D&source=editors&ust=1779048536910669&usg=AOvVaw3ryyyqx1P9ECLo0Cap5_Po) |
| 6 | [▪ ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents (Lee and Brumley, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14153&sa=D&source=editors&ust=1779048536910783&usg=AOvVaw0ZY1EPJUAy3uvd--0WvTXA) |
| 7 | [▪ Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13411&sa=D&source=editors&ust=1779048536910887&usg=AOvVaw2UyqsGW26ddjTc44S9MDGn) |
| 8 | [▪ RACC: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02280&sa=D&source=editors&ust=1779048536910974&usg=AOvVaw3i6I3WkN-pv4WhwYraaekI) |
| 9 | [▪ QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild (Safdar et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04056&sa=D&source=editors&ust=1779048536911057&usg=AOvVaw2urvxFTBEgh9v2uZXs4Hcz) |
| 10 | [▪ Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11202&sa=D&source=editors&ust=1779048536911143&usg=AOvVaw2LMkvJO8aGn8akJ47jetjT) |
| 11 | [▪ Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions (Susan, Arusoaie, and Lucanu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11163&sa=D&source=editors&ust=1779048536911231&usg=AOvVaw2iW06HdSQ_U_j4ZfT2PTSY) |
| 12 | [▪ What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10180&sa=D&source=editors&ust=1779048536911327&usg=AOvVaw06p1Yw86D4v2SpIvrzmDTS) |
| 13 | [▪ Containment Verification: AI Safety Guarantees Independent of Alignment (Moon and Varshney, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09045&sa=D&source=editors&ust=1779048536911416&usg=AOvVaw05OpGMIIaY3OhCL6rVDmet) |
| 14 | [▪ Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims (Konrad, Tanyel, and Ayvaz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10575&sa=D&source=editors&ust=1779048536911500&usg=AOvVaw04-h34wBfIa0Bhim36C9-W) |
| 15 | [▪ Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning (Huff et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09792&sa=D&source=editors&ust=1779048536911585&usg=AOvVaw2g2LcL5GtC52nnXcwok_Yb) |
| 16 | [▪ MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring (Jotautait\\.e et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09684&sa=D&source=editors&ust=1779048536911668&usg=AOvVaw0ocAIkpeYk3ZuhKZMluLsL) |
| 17 | [▪ SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08382&sa=D&source=editors&ust=1779048536911755&usg=AOvVaw0Cway1N5EsWoYKmVxPKC46) |
| 18 | [▪ Seed Hijacking of LLM Sampling and Quantum Random Number Defense (You et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08313&sa=D&source=editors&ust=1779048536911840&usg=AOvVaw0tu5ODF5naerqM8ip9m59Q) |
| 19 | [▪ Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks (Hu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08257&sa=D&source=editors&ust=1779048536911950&usg=AOvVaw3PUmDeYC6NPgMNDjhCltmp) |
| 20 | [▪ GLiGuard: Schema-Conditioned Classification for LLM Safeguard (Zaratiana et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07982&sa=D&source=editors&ust=1779048536912034&usg=AOvVaw1H8FzTYEVa2vKpI38aln9m) |
| 21 | [▪ FIT to Forget: Robust Continual Unlearning for Large Language Models (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21682&sa=D&source=editors&ust=1779048536912115&usg=AOvVaw1HswHx3Dc_7GpTUx-OGmYi) |
| 22 | [▪ SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering (Lin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17633&sa=D&source=editors&ust=1779048536912197&usg=AOvVaw1_zLa3zY-meJm1FjvLAZWP) |
| 23 | [▪ Information Theoretic Adversarial Training of Large Language Models (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05415&sa=D&source=editors&ust=1779048536912278&usg=AOvVaw3njZAi-F1QaPcXXYKQaIyY) |
| 24 | [▪ Autonomous Adversary: Red-Teaming in the age of LLM (Mamun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06486&sa=D&source=editors&ust=1779048536912369&usg=AOvVaw1R0n8oCmSpgqS1Ocs10kjI) |
| 25 | [▪ Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks (Lu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05995&sa=D&source=editors&ust=1779048536912451&usg=AOvVaw1QEAGYyaoREVly2NLRbrMu) |
| 26 | [▪ SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05704&sa=D&source=editors&ust=1779048536912532&usg=AOvVaw0ECnKr13zuuLPR_hGzGFIA) |
| 27 | [▪ GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy (Minko, Sadiekh, and Kokuykin, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05277&sa=D&source=editors&ust=1779048536912617&usg=AOvVaw3_gKYt1VhcceOdFibCrYke) |
| 28 | [▪ A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework (Chu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23338&sa=D&source=editors&ust=1779048536912700&usg=AOvVaw29Dani93IbMzYj0Q1mPUo9) |
| 29 | [▪ On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04901&sa=D&source=editors&ust=1779048536912786&usg=AOvVaw2rZYnV865IS326QQ244ODq) |
| 30 | [▪ Large Language Model assisted Hybrid Fuzzing (Meng, Duck, and Roychoudhury, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2412.15931&sa=D&source=editors&ust=1779048536912878&usg=AOvVaw2I2yHZDsaWaQK53zZ0KgRf) |
| 31 | [▪ MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents (Steinberg and Gal, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03952&sa=D&source=editors&ust=1779048536912960&usg=AOvVaw2knwVrzR9x8amIpUfj2jKT) |
| 32 | [▪ ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection (Weng et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03378&sa=D&source=editors&ust=1779048536913041&usg=AOvVaw1qW3_OuelCJ9XrX-z289G8) |
| 33 | [▪ MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03228&sa=D&source=editors&ust=1779048536913124&usg=AOvVaw3sJfV8MnD7K-oMqhd5A4J6) |
| 34 | [▪ Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02900&sa=D&source=editors&ust=1779048536913209&usg=AOvVaw2lFp5d1rILSe3PwYgI-0Fk) |
| 35 | [▪ RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs (Asif and Amiri, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01913&sa=D&source=editors&ust=1779048536913297&usg=AOvVaw0Go1PACKnoZNZNGWOSKKPw) |
| 36 | [▪ Toward a Principled Framework for Agent Safety Measurement (Lin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01644&sa=D&source=editors&ust=1779048536913381&usg=AOvVaw2I-s5bxjRoFniAXT2X7-yK) |
| 37 | [▪ When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems (Zhang, Zheng, and Chen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01133&sa=D&source=editors&ust=1779048536913464&usg=AOvVaw0Q8tjyr2FwIBK-TFAkwYsD) |
| 38 | [▪ Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts (Hasan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22628&sa=D&source=editors&ust=1779048536913546&usg=AOvVaw1vLIIyQ5sfYLEm4ASm1kyf) |
| 39 | [▪ ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00689&sa=D&source=editors&ust=1779048536913630&usg=AOvVaw2AuC_V4OyrpoBuz-WOYIVl) |
| 40 | [▪ GAVEL: Towards Rule-Based Safety Through Activation Monitoring (Rozenfeld et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19768&sa=D&source=editors&ust=1779048536913711&usg=AOvVaw1tyXBb5N1G4E1Dmf5A7DV_) |
| 41 | [▪ Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures (Mahi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27804&sa=D&source=editors&ust=1779048536913797&usg=AOvVaw2WLYlMVtssEH15Q9Izidcy) |
| 42 | [▪ AdaBFL: Multi-Layer Defensive Adaptive Aggregation for Bzantine-Robust Federated Learning (Tang, Liu, and Huang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27434&sa=D&source=editors&ust=1779048536913882&usg=AOvVaw0g80RbYdMJ59I8PGcHwD4E) |
| 43 | [▪ Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches (Wienczkowski, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27000&sa=D&source=editors&ust=1779048536913969&usg=AOvVaw1q5xC3B-C8XN4BO7NMbpp2) |
| 44 | [▪ Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study (Xu and Chen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27464&sa=D&source=editors&ust=1779048536914054&usg=AOvVaw0bQFxT3KkuXEgP1aZsWGwC) |
| 45 | [▪ Toward Autonomous SOC Operations: End-to-End LLM Framework for Threat Detection, Query Generation, and Resolution in Security Operations (Saju and Azim, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27321&sa=D&source=editors&ust=1779048536914142&usg=AOvVaw3YIYeM8v4HacU09XTeTM4m) |
| 46 | [▪ An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code (Elsayed, Fulton, and Yang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27001&sa=D&source=editors&ust=1779048536914227&usg=AOvVaw000fIsm5xS_bXOCQz-AB-X) |
| 47 | [▪ Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents (Dang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26274&sa=D&source=editors&ust=1779048536914316&usg=AOvVaw0zCMfy4XHhFY5A5dzH39FU) |
| 48 | [▪ Large Language Models as Explainable Cyberattack Detectors for Energy Industrial Control Systems (Kong et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26079&sa=D&source=editors&ust=1779048536914400&usg=AOvVaw3nvJSd2daMJkU4hQqJnL6L) |
| 49 | [▪ A Comparative Evaluation of AI Agent Security Guardrails (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24826&sa=D&source=editors&ust=1779048536914479&usg=AOvVaw0BDjSS1quTmwIiI6DHQNEj) |
| 50 | [▪ Mitigating Error Amplification in Fast Adversarial Training (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24332&sa=D&source=editors&ust=1779048536914561&usg=AOvVaw2cKLHxDYCwGJkqGFsI3gIf) |
| 51 | [▪ AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization (Ying et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24118&sa=D&source=editors&ust=1779048536914643&usg=AOvVaw10g3GsEayr7egfi9eVQNlP) |
| 52 | [▪ Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24020&sa=D&source=editors&ust=1779048536914729&usg=AOvVaw1ezoYArpJeVMevLeIsFo1v) |
| 53 | [▪ Evaluation of Prompt Injection Defenses in Large Language Models (Deep et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23887&sa=D&source=editors&ust=1779048536914812&usg=AOvVaw3jZjDpxzM2WMH6_FkAOQDB) |
| 54 | [▪ UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks (Yu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23141&sa=D&source=editors&ust=1779048536914894&usg=AOvVaw2sM2P4g2gmeAQOG6esB2YX) |
| 55 | [▪ Secure eFPGA-Enabled Edge LLM Inference: Architectural and Hardware Countermeasures (Das et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22935&sa=D&source=editors&ust=1779048536914974&usg=AOvVaw2hf8Hr5qvRKhZ7LgHWNl9t) |
| 56 | [▪ AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models (Gautam, Bahramali, and Atluri, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22871&sa=D&source=editors&ust=1779048536915056&usg=AOvVaw0_TXzpKvpEEBGevjD5rImU) |
| 57 | [▪ SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22097&sa=D&source=editors&ust=1779048536915140&usg=AOvVaw3FhTYaGS6rynq94SAZBnwL) |
| 58 | [▪ Secure LLM Fine-Tuning via Safety-Aware Probing (Wu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16737&sa=D&source=editors&ust=1779048536915220&usg=AOvVaw3F5l8n5fAQcu_r9cD4SvgI) |
| 59 | [▪ Adaptive Instruction Composition for Automated LLM Red-Teaming (Zymet et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21159&sa=D&source=editors&ust=1779048536915306&usg=AOvVaw3LK-3678b34ziNLgMt8-Mg) |
| 60 | [▪ Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks (Pallerla et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20932&sa=D&source=editors&ust=1779048536915400&usg=AOvVaw3qAt5e4slgruKfkYbnFucA) |
| 61 | [▪ Verification of Machine Unlearning is Fragile (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2408.00929&sa=D&source=editors&ust=1779048536915484&usg=AOvVaw1z3R6sYa0rwqDLBjkIl523) |
| 62 | [▪ AVISE: Framework for Evaluating the Security of AI Systems (Lempinen, Kemppainen, and Raesalmi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20833&sa=D&source=editors&ust=1779048536915569&usg=AOvVaw3r_Wb_Cip_M4K5epWCV0Cs) |
| 63 | [▪ Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing (Talluri, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20704&sa=D&source=editors&ust=1779048536915650&usg=AOvVaw2exvEnkdVH-03r-hnUktjG) |
| 64 | [▪ Towards Certified Malware Detection: Provable Guarantees Against Evasion Attacks (Giri et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20495&sa=D&source=editors&ust=1779048536915734&usg=AOvVaw0f6hMxGDgZ7FgaBkdUcZ8B) |
| 65 | [▪ Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning (Ni, Christodorescu, and Jia, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20179&sa=D&source=editors&ust=1779048536915818&usg=AOvVaw2bT2XK84ftAAdYsNRIuOt-) |
| 66 | [▪ Benchmarking Misuse Mitigation Against Covert Adversaries (Brown et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06414&sa=D&source=editors&ust=1779048536915898&usg=AOvVaw0tZF9v0Ytyo08KMQYE-NMi) |
| 67 | [▪ ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System (Liang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18789&sa=D&source=editors&ust=1779048536915978&usg=AOvVaw1sjexupJWQq_pS3pVZKWM3) |
| 68 | [▪ Malicious ML Model Detection by Learning Dynamic Behaviors (Nambiar, Pradhan, and Soremekun, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19438&sa=D&source=editors&ust=1779048536916058&usg=AOvVaw3DazBEfCYFgp41jD9CghBN) |
| 69 | [▪ SAGE: Signal-Amplified Guided Embeddings for LLM-based Vulnerability Detection (Shan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19031&sa=D&source=editors&ust=1779048536916139&usg=AOvVaw25EXtDujQH9j5b7ODhGKEE) |
| 70 | [▪ Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents (Xu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18860&sa=D&source=editors&ust=1779048536916222&usg=AOvVaw1GY13RKkamO3LE7-nZbnHN) |
| 71 | [▪ ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction (Wei, Wu, and Sun, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.01770&sa=D&source=editors&ust=1779048536916314&usg=AOvVaw3SoSW2j7zhad3x_VSPy-Hm) |
| 72 | [▪ SDLLMFuzz: Dynamic-static LLM-assisted greybox fuzzing for structured input programs (Zou et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17750&sa=D&source=editors&ust=1779048536916397&usg=AOvVaw3CuJhy_L0-KrBREdAgayFt) |
| 73 | [▪ Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives (Son, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16870&sa=D&source=editors&ust=1779048536916478&usg=AOvVaw3ybeBJaX5G-ziz5oTp2eIS) |
| 74 | [▪ enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways (Metere, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16838&sa=D&source=editors&ust=1779048536916560&usg=AOvVaw0oChQNq6EbfJfD2-FSiF3d) |
| 75 | [▪ SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models (Mohammad and Bayaz{\\i}t, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16606&sa=D&source=editors&ust=1779048536916741&usg=AOvVaw0UQhNWatqmKY9kEDeqhuZV) |
| 76 | [▪ TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts (Chu, Wang, and Huang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16542&sa=D&source=editors&ust=1779048536916842&usg=AOvVaw0jJi2CU4t2gl-eaIDjkPkz) |
| 77 | [▪ Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility (Hong et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15579&sa=D&source=editors&ust=1779048536916931&usg=AOvVaw04Ak931GtvmU_qqzNHPLY3) |
| 78 | [▪ Privacy-Preserving LLMs Routing (Wu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15728&sa=D&source=editors&ust=1779048536917012&usg=AOvVaw3qKD39a_KLiFdmFtPmmpL5) |
| 79 | [▪ SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment (Lin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13630&sa=D&source=editors&ust=1779048536917101&usg=AOvVaw2cvBaV7YSXPoDaECGpNywG) |
| 80 | [▪ Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory (Fu and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12817&sa=D&source=editors&ust=1779048536917188&usg=AOvVaw0JgsJvZEZTSk2tVNavPtVC) |
| 81 | [▪ SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents (Begimher et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12040&sa=D&source=editors&ust=1779048536917272&usg=AOvVaw2YVcQZmjtYOj9YCwJBwf-R) |
| 82 | [▪ Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16057&sa=D&source=editors&ust=1779048536917362&usg=AOvVaw1IJ0nz-gF3s1VnBs4wLDbk) |
| 83 | [▪ Towards Automated Pentesting with Large Language Models (Bessa et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11772&sa=D&source=editors&ust=1779048536917443&usg=AOvVaw08294SgBF4b2kzo4KUTsUm) |
| 84 | [▪ QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits (Azimi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10933&sa=D&source=editors&ust=1779048536917526&usg=AOvVaw2cLJ-lxYBvldItbzJ1l-Tm) |
| 85 | [▪ PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification (Gong and Deng, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10134&sa=D&source=editors&ust=1779048536917611&usg=AOvVaw2t_RXXHQYqQGlmKdLXvLfv) |
| 86 | [▪ Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark (Shao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.05674&sa=D&source=editors&ust=1779048536917694&usg=AOvVaw267IvkmBYgAsqlQ1CDhSAD) |
| 87 | [▪ Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.20472&sa=D&source=editors&ust=1779048536917778&usg=AOvVaw1cloTEJlnKpoRNzP90Mtxk) |
| 88 | [▪ Vulnerability Detection with Interprocedural Context in Multiple Languages: Assessing Effectiveness and Cost of Modern LLMs (Lira et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08417&sa=D&source=editors&ust=1779048536917864&usg=AOvVaw0Dx3Wl39f-ipsN_AarT_ab) |
| 89 | [▪ Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08304&sa=D&source=editors&ust=1779048536917945&usg=AOvVaw3cCkcx41H5nSfeJIssdFVh) |
| 90 | [▪ Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models (Qi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08297&sa=D&source=editors&ust=1779048536918026&usg=AOvVaw3q-_bImgT6qLGLJ6R3XzGP) |
| 91 | [▪ MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security (Rostamzadeh et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07551&sa=D&source=editors&ust=1779048536918109&usg=AOvVaw1XuKJ8hudMsw3CPcfdwTmU) |
| 92 | [▪ Can Drift-Adaptive Malware Detectors Be Made Robust? Attacks and Defenses Under White-Box and Black-Box Threats (Li, Akil, and Bertino, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06599&sa=D&source=editors&ust=1779048536918193&usg=AOvVaw0aRQdgJHqJVZwQefl32YFx) |
| 93 | [▪ The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? (Bhatt et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06436&sa=D&source=editors&ust=1779048536918274&usg=AOvVaw1_EZFyWDHWEabOYS_NHdn-) |
| 94 | [▪ Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization (Maljkovic et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06285&sa=D&source=editors&ust=1779048536918375&usg=AOvVaw2Ijh_gPv2c9Mvm-33ftAa5) |
| 95 | [▪ SALLIE: Safeguarding Against Latent Language & Image Exploits (Azov, Rivlin, and Shtar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06247&sa=D&source=editors&ust=1779048536918457&usg=AOvVaw26eXJq8nJpREIWgFqF3Sma) |
| 96 | [▪ CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments (Keppler, Gst\\"ur, and Hagenmeyer, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06019&sa=D&source=editors&ust=1779048536918545&usg=AOvVaw34VJX1c0UXTrwWsVXbjgol) |
| 97 | [▪ A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms (Acharya and Gupta, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05969&sa=D&source=editors&ust=1779048536918637&usg=AOvVaw1lS_hUC6WbaDf2J1GwmMll) |
| 98 | [▪ Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts (Uenal, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05872&sa=D&source=editors&ust=1779048536918724&usg=AOvVaw38j62kBMvnrQ4KXvTo5aFw) |
| 99 | [▪ BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents (Ma, Wu, and Yan, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05793&sa=D&source=editors&ust=1779048536918813&usg=AOvVaw2UyPOlcdPIedLxL2sg7RY-) |
| 100 | [▪ Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code (Blain and Noiseux, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05292&sa=D&source=editors&ust=1779048536918897&usg=AOvVaw290Aj2YgHgC4AKk4TtGbkL) |
| 101 | [▪ SoSBench: Benchmarking Safety Alignment on Six Scientific Domains (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.21605&sa=D&source=editors&ust=1779048536918976&usg=AOvVaw2gipKZC8sHR9MtBFByqzhS) |
| 102 | [▪ Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework (Zhou et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04852&sa=D&source=editors&ust=1779048536919058&usg=AOvVaw1kPo_LShgL0maPAgUJliDF) |
| 103 | [▪ Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning (Zhang, Goel, and Ahmad, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04442&sa=D&source=editors&ust=1779048536919142&usg=AOvVaw2niyrBhuitc-2bu6CKGEnu) |
| 104 | [▪ CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04060&sa=D&source=editors&ust=1779048536919223&usg=AOvVaw3X5UOMwY5PpOKYQ62Gwm69) |
| 105 | [▪ Automating Cloud Security and Forensics Through a Secure-by-Design Generative AI Framework (Alharthi and Garcia, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03912&sa=D&source=editors&ust=1779048536919310&usg=AOvVaw0gAgdZcwa9CE7DmgLyGVC5) |
| 106 | [▪ SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03587&sa=D&source=editors&ust=1779048536919393&usg=AOvVaw2CcDkeLDMPVKwQbjGPXFSC) |
| 107 | [▪ Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI (Cheruvu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05660&sa=D&source=editors&ust=1779048536919476&usg=AOvVaw0gMaIIrrg6ipnucgW344iD) |
| 108 | [▪ Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection (Halder et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01637&sa=D&source=editors&ust=1779048536919559&usg=AOvVaw30SJgd17LobA-5hAjLWZsz) |
| 109 | [▪ Assertain: Automated Security Assertion Generation Using Large Language Models (Tarek et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01583&sa=D&source=editors&ust=1779048536919645&usg=AOvVaw0RZU_oapPqRnZ2z1GTr3nl) |
| 110 | [▪ Certifiably Robust RAG against Retrieval Corruption (Xiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2405.15556&sa=D&source=editors&ust=1779048536919748&usg=AOvVaw1oXO66v7J4XgH8CGptHfy5) |
| 111 | [▪ Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents (Ye et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00430&sa=D&source=editors&ust=1779048536919844&usg=AOvVaw23cmRjiMgT6jNQvidpghNA) |
| 112 | [▪ VibeGuard: A Security Gate Framework for AI-Generated Code (Xie, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01052&sa=D&source=editors&ust=1779048536919933&usg=AOvVaw291MLtIilshKHVNZJorK_T) |
| 113 | [▪ Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks (Sahu, Samanta, and Soosahabi, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01039&sa=D&source=editors&ust=1779048536920027&usg=AOvVaw0CjQyPWyX4RUKtsDm2uU1F) |
| 114 | [▪ Quantum-Safe Code Auditing: LLM-Assisted Static Analysis and Quantum-Aware Risk Scoring for Post-Quantum Cryptography Migration (Shaw, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00560&sa=D&source=editors&ust=1779048536920120&usg=AOvVaw1z4h-VZ6hBTubC9Kb_IqBK) |
| 115 | [▪ SecureVibeBench: Evaluating Secure Coding Capabilities of Code Agents with Realistic Vulnerability Scenarios (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22097&sa=D&source=editors&ust=1779048536920213&usg=AOvVaw3FIpbUpr_9yCCFEgG04l7P) |
| 116 | [▪ Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks (Xiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.30016&sa=D&source=editors&ust=1779048536920329&usg=AOvVaw0MJvjqdeboKwiE1r_bZIes) |
| 117 | [▪ CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks (Patil, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29062&sa=D&source=editors&ust=1779048536920418&usg=AOvVaw3692DwLw0pnrTLqCVc7fTL) |
| 118 | [▪ Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems (Cai et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28998&sa=D&source=editors&ust=1779048536920502&usg=AOvVaw3R21NnqDkCuQtVyFh64PIx) |
| 119 | [▪ Attesting LLM Pipelines: Enforcing Verifiable Training and Release Claims (Tan, Singer, and Anagnostopoulos, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28988&sa=D&source=editors&ust=1779048536920584&usg=AOvVaw2it0dmDYURwx1njO60rQfs) |
| 120 | [▪ GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models (Mia et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28817&sa=D&source=editors&ust=1779048536920667&usg=AOvVaw1igfI0qSlP5qQtYEFxXPfx) |
| 121 | [▪ SkillTester: Benchmarking Utility and Security of Agent Skills (Wang, Wang, and Xu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28815&sa=D&source=editors&ust=1779048536920747&usg=AOvVaw0SBW_hdHejBAUkS01IIS9l) |
| 122 | [▪ VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications (Zhu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04014&sa=D&source=editors&ust=1779048536920862&usg=AOvVaw2Sy8GdJMw2EWScU8zlvY6H) |
| 123 | [▪ Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries (Raz et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28655&sa=D&source=editors&ust=1779048536920946&usg=AOvVaw3X44vuArDfGhp5U3baJQMl) |
| 124 | [▪ A Systematic Taxonomy of Security Vulnerabilities in the OpenClaw AI Agent Framework (Suwansathit, Zhang, and Gu, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27517&sa=D&source=editors&ust=1779048536921031&usg=AOvVaw3NuE5ebm-u2t6UxtEtZXTw) |
| 125 | [▪ Protecting User Prompts Via Character-Level Differential Privacy (Arachchige et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26032&sa=D&source=editors&ust=1779048536921112&usg=AOvVaw18_U3MRL9FZoE4Zwotgtuo) |
| 126 | [▪ Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25412&sa=D&source=editors&ust=1779048536921194&usg=AOvVaw1LexImSBOp1WBBdpP-BXBa) |
| 127 | [▪ Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage (Sahay et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23966&sa=D&source=editors&ust=1779048536921275&usg=AOvVaw1JYrQ0owelig6y_RJiQTZP) |
| 128 | [▪ Agent Audit: A Security Analysis System for LLM Agent Applications (Zhang, Nian, and Zhao, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22853&sa=D&source=editors&ust=1779048536921362&usg=AOvVaw23VwpY3NJG65fv3dDL83BK) |
| 129 | [▪ Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval (Sabir et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22717&sa=D&source=editors&ust=1779048536921445&usg=AOvVaw0UBYYVM25veTKkSo4HuTHl) |
| 130 | [▪ BioShield: A Context-Aware Firewall for Securing Bio-LLMs (Das et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22612&sa=D&source=editors&ust=1779048536921525&usg=AOvVaw1yRXGRoxCUsUGcL5lqQEWV) |
| 131 | [▪ T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search (Lee et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22341&sa=D&source=editors&ust=1779048536921605&usg=AOvVaw3x_6ZETQY8rwV3zphnvBmv) |
| 132 | [▪ Agentproof: Static Verification of Agent Workflow Graphs (Xavier et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20356&sa=D&source=editors&ust=1779048536921683&usg=AOvVaw27nue2TnL9-Q9lF_q4EbLH) |
| 133 | [▪ SecureBreak -- A dataset towards safe and secure models (Arazzi, Kembu, and Nocera, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21975&sa=D&source=editors&ust=1779048536921762&usg=AOvVaw17m6-Nfq3BAHDbobIu003u) |
| 134 | [▪ Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks (Mu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21654&sa=D&source=editors&ust=1779048536921850&usg=AOvVaw3gNyIgtusYDXwaGNMGG0mz) |
| 135 | [▪ DeepXplain: XAI-Guided Autonomous Defense Against Multi-Stage APT Campaigns (Phan and Bauschert, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21296&sa=D&source=editors&ust=1779048536921953&usg=AOvVaw0TIpIezXM6AcQW3p4w8UKs) |
| 136 | [▪ LISAA: A Framework for Large Language Model Information Security Awareness Assessment (Cohen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.13207&sa=D&source=editors&ust=1779048536922038&usg=AOvVaw12dEWNV_vXRx-kWUl5_mxl) |
| 137 | [▪ A Framework for Formalizing LLM Agent Security (Siu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19469&sa=D&source=editors&ust=1779048536922117&usg=AOvVaw3LyelL4Zr05vsBOWjmonLc) |
| 138 | [▪ The Autonomy Tax: Defense Training Breaks LLM Agents (Li and Zhao, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19423&sa=D&source=editors&ust=1779048536922196&usg=AOvVaw3dyWbrEyomXXie6KA8LVpJ) |
| 139 | [▪ Network and Device Level Cyber Deception for Contested Environments Using RL and LLMs (Sahu, Paul, and Macwan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17272&sa=D&source=editors&ust=1779048536922279&usg=AOvVaw0mEgeWFiV18Ghm7JeqfvUQ) |
| 140 | [▪ Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18245&sa=D&source=editors&ust=1779048536922371&usg=AOvVaw0DlzLtMlNqgUsOxlFYPivf) |
| 141 | [▪ Security awareness in LLM agents: the NDAI zone case (Bottazzi and Park, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19011&sa=D&source=editors&ust=1779048536922451&usg=AOvVaw1H5GOOj28Kz_9G3f9nTiYB) |
| 142 | [▪ Toward Reliable, Safe, and Secure LLMs for Scientific Applications (Chaturvedi, Bergerson, and Mallick, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18235&sa=D&source=editors&ust=1779048536922533&usg=AOvVaw0wW0vGWOAvh-BLzCvEFS7j) |
| 143 | [▪ Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows (Sigdel and Baral, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18059&sa=D&source=editors&ust=1779048536922614&usg=AOvVaw2_asnJpov9zTAJ5YS9wShl) |
| 144 | [▪ Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs (Yang and Zhu, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17902&sa=D&source=editors&ust=1779048536922695&usg=AOvVaw2W4P4dtLLNTqohol1QlNbV) |
| 145 | [▪ Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework (Onitiju and Vakilinia, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17123&sa=D&source=editors&ust=1779048536922783&usg=AOvVaw1U19gVPTo_OQRlXceo3PHz) |
| 146 | [▪ DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns (Phan, Nguyen, and Bauschert, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16969&sa=D&source=editors&ust=1779048536922865&usg=AOvVaw0YQrHKy9_HAIEwqmbEmFtF) |
| 147 | [▪ Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15800&sa=D&source=editors&ust=1779048536922947&usg=AOvVaw2WkJFv3ks0MVcu2D07E2H4) |
| 148 | [▪ Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models (Liu and Chen, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16382&sa=D&source=editors&ust=1779048536923030&usg=AOvVaw12LdfcmwXb0GwWr9GI7XTT) |
| 149 | [▪ TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15408&sa=D&source=editors&ust=1779048536923110&usg=AOvVaw0fZGOS13mjU1OmQRuMCllR) |
| 150 | [▪ SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration (Pan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15397&sa=D&source=editors&ust=1779048536923198&usg=AOvVaw1CrJvzzpG37Hlv13oiWfcy) |
| 151 | [▪ Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats (Zhang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14860&sa=D&source=editors&ust=1779048536923280&usg=AOvVaw1YWR-j2iXsfKg84Mc44cNi) |
| 152 | [▪ Governing Dynamic Capabilities: Cryptographic Binding and Reproducibility Verification for AI Agent Tool Use (Zhou, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14332&sa=D&source=editors&ust=1779048536923367&usg=AOvVaw3tWBYPqQ5r7RBFR2eh8_4f) |
| 153 | [▪ SecDTD: Dynamic Token Drop for Secure Transformers Inference (Cai et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13670&sa=D&source=editors&ust=1779048536923446&usg=AOvVaw2hcC_NFzlK20dF4COqFyas) |
| 154 | [▪ CTI-REALM: Benchmark to Evaluate Agent Performance on Security Detection Rule Generation Capabilities (Chakraborty et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13517&sa=D&source=editors&ust=1779048536923529&usg=AOvVaw0t19jewK3kNdMDDcsYBuKC) |
| 155 | [▪ Why Neural Structural Obfuscation Can't Kill White-Box Watermarks for Good! (Jiang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12679&sa=D&source=editors&ust=1779048536923611&usg=AOvVaw3y28qTWJyCLpcjJTtBfgqw) |
| 156 | [▪ Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw (Ying et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12644&sa=D&source=editors&ust=1779048536923696&usg=AOvVaw057RROwizCyeRWyZh4dRXR) |
| 157 | [▪ AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents (Yuan, Su, and Zhao, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12621&sa=D&source=editors&ust=1779048536923780&usg=AOvVaw1FQPjEFSeocsvJWXCAO8to) |
| 158 | [▪ Security Considerations for Artificial Intelligence Agents (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12230&sa=D&source=editors&ust=1779048536923859&usg=AOvVaw0_DHRgm6FaOyj1tVxKoFOp) |
| 159 | [▪ OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents (Li, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11853&sa=D&source=editors&ust=1779048536923944&usg=AOvVaw0-HNraOq-0Bmb0Ccr43kIo) |
| 160 | [▪ Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms (Wendlinger et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11212&sa=D&source=editors&ust=1779048536924031&usg=AOvVaw21QjFttuJbucZ0A1J3JzkS) |
| 161 | [▪ Differential Privacy in Machine Learning: A Survey from Symbolic AI to LLMs (Aguilera-Mart\\'inez and Berzal, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.11687&sa=D&source=editors&ust=1779048536924113&usg=AOvVaw3z2WHXZaj_SCKY2WgcaTtB) |
| 162 | [▪ TOSSS: a CVE-based Software Security Benchmark for Large Language Models (Damie et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10969&sa=D&source=editors&ust=1779048536924194&usg=AOvVaw0w_JJZODNjA689c8Gale7G) |
| 163 | [▪ Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial Attacks (Soltani et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10413&sa=D&source=editors&ust=1779048536924276&usg=AOvVaw2FvWqVd4meaFYLeKQe6Aex) |
| 164 | [▪ AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07427&sa=D&source=editors&ust=1779048536924392&usg=AOvVaw1ObFNGKW-JhBwFjIOhed18) |
| 165 | [▪ SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.08520&sa=D&source=editors&ust=1779048536924479&usg=AOvVaw0L9ecqV_OQlr7B_KfKrfVB) |
| 166 | [▪ DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation (Jiang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07835&sa=D&source=editors&ust=1779048536924561&usg=AOvVaw1ihpelSeJiN2T51_K-eoiy) |
| 167 | [▪ Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI (Jin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07466&sa=D&source=editors&ust=1779048536924643&usg=AOvVaw0AnhjnwWEEcbpUNBTxAGqS) |
| 168 | [▪ Where Do LLM-based Systems Break? A System-Level Security Framework for Risk Assessment and Treatment (Nagaraja and Bahsi, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07460&sa=D&source=editors&ust=1779048536924725&usg=AOvVaw12h27V74TyuXEzcqYX_Wjw) |
| 169 | [▪ ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code (Filho, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.06365&sa=D&source=editors&ust=1779048536924809&usg=AOvVaw3arPURY8VOcWKl9Zi_UegU) |
| 170 | [▪ Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05773&sa=D&source=editors&ust=1779048536924891&usg=AOvVaw229mbFwlR9-A04N4CmxlUx) |
| 171 | [▪ Secure human oversight of AI: Threat modeling in a socio-technical context (Ditz et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.12290&sa=D&source=editors&ust=1779048536924970&usg=AOvVaw3EQt8b5t9wAp0i17b8LEHG) |
| 172 | [▪ EVMbench: Evaluating AI Agents on Smart Contract Security (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04915&sa=D&source=editors&ust=1779048536925050&usg=AOvVaw3LsrcDy32hcuYmWGV1Et0I) |
| 173 | [▪ Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks (Chu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04459&sa=D&source=editors&ust=1779048536925132&usg=AOvVaw3WBcWUjm8bjvinNsgwZQ06) |
| 174 | [▪ Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study (Nagaraja and Bahsi, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03633&sa=D&source=editors&ust=1779048536925214&usg=AOvVaw3QQb86eP5NNwnvOyYzKMgP) |
| 175 | [▪ WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols (Maheri et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00272&sa=D&source=editors&ust=1779048536925299&usg=AOvVaw348SBJEUoL4SIvoijc0Pst) |
| 176 | [▪ LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities (Nakka et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.05619&sa=D&source=editors&ust=1779048536925386&usg=AOvVaw36VT0lmwGoGIAUuVWHRW5a) |
| 177 | [▪ Contextualized Privacy Defense for LLM Agents (Wen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02983&sa=D&source=editors&ust=1779048536925465&usg=AOvVaw0J7N-LoIjL-HieVjbQ5EOp) |
| 178 | [▪ SOSecure: Safer Code Generation with RAG and StackOverflow Discussions (Mukherjee and Hellendoorn, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.13654&sa=D&source=editors&ust=1779048536925546&usg=AOvVaw0Gr0B9otVHkct3syF5Vne-) |
| 179 | [▪ ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline (Singh et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19149&sa=D&source=editors&ust=1779048536925627&usg=AOvVaw3gHWX1cSi6jTRWEwfdzreX) |
| 180 | [▪ Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision (Mukherjee and Hellendoorn, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01494&sa=D&source=editors&ust=1779048536925708&usg=AOvVaw2nopo4Q7_MPjmPowqY-gwc) |
| 181 | [▪ Token-level Data Selection for Safe LLM Fine-tuning (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01185&sa=D&source=editors&ust=1779048536925786&usg=AOvVaw3vU9HggcIbLXmLxqWnorSJ) |
| 182 | [▪ MPU: Towards Secure and Privacy-Preserving Knowledge Unlearning for Large Language Models (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23798&sa=D&source=editors&ust=1779048536925866&usg=AOvVaw31Vq1Tj79TRyTLa3K9Hsu7) |
| 183 | [▪ Learning to Generate Secure Code via Token-Level Rewards (Quan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23407&sa=D&source=editors&ust=1779048536925946&usg=AOvVaw1phjsBr6Bxf3IuiH-TPEoE) |
| 184 | [▪ Automated Vulnerability Detection in Source Code Using Deep Representation Learning (Seas et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23121&sa=D&source=editors&ust=1779048536926026&usg=AOvVaw2JDhX6n4m1_My5aEex70hd) |
| 185 | [▪ A Lightweight Defense Mechanism against Next Generation of Phishing Emails using Distilled Attention-Augmented BiLSTM (Eskandarian et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22250&sa=D&source=editors&ust=1779048536926110&usg=AOvVaw3p_2MQJuhA7spXOVkWpUOa) |
| 186 | [▪ Secure Semantic Communications via AI Defenses: Fundamentals, Solutions, and Future Directions (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22134&sa=D&source=editors&ust=1779048536926192&usg=AOvVaw06Zl2X3blnfm3vgwtdN8E0) |
| 187 | [▪ Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning (Garc\\'ia-M\\'arquez et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21841&sa=D&source=editors&ust=1779048536926277&usg=AOvVaw0WrfRSW0IPPFPIcuV1p6rL) |
| 188 | [▪ A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications (Srivastava, Janardhan, and Jauhari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21267&sa=D&source=editors&ust=1779048536926371&usg=AOvVaw3shlKkSHKeY7ikO-cWmatn) |
| 189 | [▪ AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06326&sa=D&source=editors&ust=1779048536926454&usg=AOvVaw19MjL2t7FmX-7_Dc4AnWAt) |
| 190 | [▪ Dynamic Probabilistic Noise Injection for Membership Inference Defense (Forough and Haddadi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.13362&sa=D&source=editors&ust=1779048536926536&usg=AOvVaw3yWESTu281bNagKunm72mh) |
| 191 | [▪ LLM-enabled Applications Require System-Level Threat Monitoring (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19844&sa=D&source=editors&ust=1779048536926615&usg=AOvVaw1WmWd7ZwKfum42LzzmKdJ7) |
| 192 | [▪ CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents (Ba, Li, and Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19547&sa=D&source=editors&ust=1779048536926696&usg=AOvVaw3A0rxBkp5z7OmCKZU-7WCc) |
| 193 | [▪ KUDA: Knowledge Unlearning by Deviating Representation for Large Language Models (Fang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19275&sa=D&source=editors&ust=1779048536926777&usg=AOvVaw2ZGOgci68IhdFDPqd5mSCp) |
| 194 | [▪ MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs (Kan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18782&sa=D&source=editors&ust=1779048536926861&usg=AOvVaw3CAutOoCCWOBbwuH79Q4-N) |
| 195 | [▪ DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing (Brinkhege and Menon, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.17413&sa=D&source=editors&ust=1779048536926942&usg=AOvVaw3XNWavDLTNUm33iQS0zFH6) |
| 196 | [▪ NeST: Neuron Selective Tuning for LLM Safety (Behrouzi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16835&sa=D&source=editors&ust=1779048536927019&usg=AOvVaw1_Nc5MfiaXG_Vqr5HkctgB) |
| 197 | [▪ NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist (Bertram and Geiping, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16756&sa=D&source=editors&ust=1779048536927100&usg=AOvVaw0FQvSCxXcggNggKOWtCiYi) |
| 198 | [▪ Secure Coding with AI -- From Detection to Repair (Belozerov, Barclay, and Sami, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.20814&sa=D&source=editors&ust=1779048536927182&usg=AOvVaw0Y3Z7qsavkvcvGCmVM6f7J) |
| 199 | [▪ PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models (Yuan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.03544&sa=D&source=editors&ust=1779048536927263&usg=AOvVaw3wTrBjsbKlM6_lC0ZcvmuZ) |
| 200 | [▪ Large Language Models for Secure Code Assessment: A Multi-Language Empirical Study (Dozono, Gasiba, and Stocco, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2408.06428&sa=D&source=editors&ust=1779048536927351&usg=AOvVaw39DFGnpqOOiIyQeI-4Zsyo) |
| 201 | [▪ ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI (Tong et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14135&sa=D&source=editors&ust=1779048536927433&usg=AOvVaw3B2QFuENLmuLbqkt3xQIUo) |
| 202 | [▪ MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents (Zhou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14281&sa=D&source=editors&ust=1779048536927515&usg=AOvVaw0cSAFIK3DwMgAk4Rcjx0zl) |
| 203 | [▪ From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection (Li, Yu, and Wang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14012&sa=D&source=editors&ust=1779048536927597&usg=AOvVaw38XboK4uB4rXhxk3iw02pE) |
| 204 | [▪ Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13379&sa=D&source=editors&ust=1779048536927680&usg=AOvVaw2LHMa3FsyOK79J7vA5QZ3I) |
| 205 | [▪ Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks (Zafar et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12943&sa=D&source=editors&ust=1779048536927767&usg=AOvVaw1_vlXNLmOMJkv2r0QnioMP) |
| 206 | [▪ TensorCommitments: A Lightweight Verifiable Inference for Language Models (Baser et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12630&sa=D&source=editors&ust=1779048536927848&usg=AOvVaw071114MlJfeCQJXxXCyyaq) |
| 207 | [▪ Sparse Autoencoders are Capable LLM Jailbreak Mitigators (Assogba et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12418&sa=D&source=editors&ust=1779048536927927&usg=AOvVaw0qpCqvjiu8pmyAuVgteUeb) |
| 208 | [▪ Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset (Kang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.19054&sa=D&source=editors&ust=1779048536928008&usg=AOvVaw39uXARL9JtGAow7SloxJQG) |
| 209 | [▪ DeepSight: An All-in-One LM Safety Toolkit (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12092&sa=D&source=editors&ust=1779048536928086&usg=AOvVaw2d57h9YhIYjc0AsVy9lWuM) |
| 210 | [▪ The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates (Willis, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11301&sa=D&source=editors&ust=1779048536928169&usg=AOvVaw3D5n4mLkE-Hngf75_A6ZGP) |
| 211 | [▪ SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models (Thornton, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18542&sa=D&source=editors&ust=1779048536928251&usg=AOvVaw1O3JxF9tCYtG98hsLKLzsC) |
| 212 | [▪ GoodVibe: Security-by-Vibe for LLM-Based Code Generation (Thang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10778&sa=D&source=editors&ust=1779048536928346&usg=AOvVaw16S-mOfIF8Mn6ps2Ocw23s) |
| 213 | [▪ MAPS: A Multilingual Benchmark for Agent Performance and Security (Hofman et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15935&sa=D&source=editors&ust=1779048536928432&usg=AOvVaw2w08vd79-VnhB1Ml40EBbw) |
| 214 | [▪ Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference (Chu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.08438&sa=D&source=editors&ust=1779048536928513&usg=AOvVaw0zMKXHlYoksw1kr_Ncppxf) |
| 215 | [▪ Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks (Dhabhi and Thimmaraju, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09629&sa=D&source=editors&ust=1779048536928598&usg=AOvVaw1tIrAhFRkY5Vb6HrOZgEZV) |
| 216 | [▪ Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks (Fu, Guo, and Luo, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08679&sa=D&source=editors&ust=1779048536928679&usg=AOvVaw0Mp37M78gVfD4RNOWtmoG9) |
| 217 | [▪ MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07517&sa=D&source=editors&ust=1779048536928760&usg=AOvVaw1ragvuEyn_620ZYzkipyP-) |
| 218 | [▪ Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model (Wu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07422&sa=D&source=editors&ust=1779048536928841&usg=AOvVaw0ynrrGXn5SzJp_z2tMQl-Q) |
| 219 | [▪ Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention (Singh et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06623&sa=D&source=editors&ust=1779048536928924&usg=AOvVaw1UQrbxqcax93CE4k9vME9g) |
| 220 | [▪ TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (Hossain et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06911&sa=D&source=editors&ust=1779048536929006&usg=AOvVaw1g3dasVTDDXCagE352cnw_) |
| 221 | [▪ Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models (Lu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06687&sa=D&source=editors&ust=1779048536929086&usg=AOvVaw1S-Nd2_q8k4uOunrqZXxol) |
| 222 | [▪ How Catastrophic is Your LLM? Certifying Risk in Conversation (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03969&sa=D&source=editors&ust=1779048536929166&usg=AOvVaw0V4VOfIi3iGRBAklqvoNQa) |
| 223 | [▪ Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection (Firouzi and Ghafari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05868&sa=D&source=editors&ust=1779048536929249&usg=AOvVaw0wvxv2cPG9jcSg-SUjV1Go) |
| 224 | [▪ Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening (Yu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05386&sa=D&source=editors&ust=1779048536929340&usg=AOvVaw1CsCtKStQA-D39rlAmc1_B) |
| 225 | [▪ Comparative Insights on Adversarial Machine Learning from Industry and Academia: A User-Study Approach (Kakkad et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04753&sa=D&source=editors&ust=1779048536929425&usg=AOvVaw2yTXv0U5cbsRF_9juEx1GA) |
| 226 | [▪ Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models (Ding et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18533&sa=D&source=editors&ust=1779048536929507&usg=AOvVaw36Es0Wq9gpgO65EIyKS4eA) |
| 227 | [▪ Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation (Marri, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02584&sa=D&source=editors&ust=1779048536929589&usg=AOvVaw2-ciT4G3MFdzKaK-3rLo0_) |
| 228 | [▪ GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video (Zhu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03328&sa=D&source=editors&ust=1779048536929669&usg=AOvVaw0RihYpLeamFhj2UXaE7mY1) |
| 229 | [▪ TinyGuard:A lightweight Byzantine Defense for Resource-Constrained Federated Learning via Statistical Update Fingerprints (Mahdavi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02615&sa=D&source=editors&ust=1779048536929756&usg=AOvVaw2D6cPHVWSZRc__zSJLyTCn) |
| 230 | [▪ To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack (Zhuo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02595&sa=D&source=editors&ust=1779048536929835&usg=AOvVaw3ce-0vja_x3dsXzbKDLMA2) |
| 231 | [▪ MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22924&sa=D&source=editors&ust=1779048536929915&usg=AOvVaw1gTqIrLa_gG3BpVWvDT372) |
| 232 | [▪ When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models (Ou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09689&sa=D&source=editors&ust=1779048536929995&usg=AOvVaw20n43YVafoBT9DXHcLSgBQ) |
| 233 | [▪ RACA: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02280&sa=D&source=editors&ust=1779048536930078&usg=AOvVaw1o3xFB6yx81sd05bdOmqBv) |
| 234 | [▪ Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents (He et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02164&sa=D&source=editors&ust=1779048536930158&usg=AOvVaw3a6QYzt81Y1d-PjKETS1Ew) |
| 235 | [▪ SMCP: Secure Model Context Protocol (Hou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01129&sa=D&source=editors&ust=1779048536930236&usg=AOvVaw1ZrkA27v6kSlPXisdfchl1) |
| 236 | [▪ Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation (Jamshidi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00219&sa=D&source=editors&ust=1779048536930325&usg=AOvVaw12BzRLPoJnlCFhUTw4E_yI) |
| 237 | [▪ Detecting Instruction Fine-tuning Attacks using Influence Function (Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.09026&sa=D&source=editors&ust=1779048536930409&usg=AOvVaw3CeGi-Lhyyza7ikGNgWeDq) |
| 238 | [▪ No More, No Less: Least-Privilege Language Models (Rauba et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.23157&sa=D&source=editors&ust=1779048536930495&usg=AOvVaw1U0AJcW4tTU-L8Fffv7jMe) |
| 239 | [▪ RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22706&sa=D&source=editors&ust=1779048536930578&usg=AOvVaw2OEwvhtxdd73GmO4Cn0tNA) |
| 240 | [▪ FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks (Xu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22485&sa=D&source=editors&ust=1779048536930662&usg=AOvVaw1M9tKkeEI3URhJH4hyW0EQ) |
| 241 | [▪ A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy (Correia et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22240&sa=D&source=editors&ust=1779048536930746&usg=AOvVaw2Jx5P7NL46BFVjJgvj9l6v) |
| 242 | [▪ ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense (Ding, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22182&sa=D&source=editors&ust=1779048536930828&usg=AOvVaw3eEWiHRYLWJdjaFLGXHjNz) |
| 243 | [▪ SafeSearch: Automated Red-Teaming of LLM-Based Search Agents (Dong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23694&sa=D&source=editors&ust=1779048536930906&usg=AOvVaw2oLs8WIlhaGyYoQUezElar) |
| 244 | [▪ FIT: Defying Catastrophic Forgetting in Continual LLM Unlearning (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21682&sa=D&source=editors&ust=1779048536930987&usg=AOvVaw2RLrFrgLNuU-PDi8oFelpO) |
| 245 | [▪ RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21380&sa=D&source=editors&ust=1779048536931068&usg=AOvVaw3r8mW-E1mUAYosDieqsXHI) |
| 246 | [▪ Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses (Hatami et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20184&sa=D&source=editors&ust=1779048536931150&usg=AOvVaw3VO_IuBIN7FxLCXK9G5QE4) |
| 247 | [▪ Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications (Shahin and Alsmadi, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19970&sa=D&source=editors&ust=1779048536931234&usg=AOvVaw2wd6_CZOOaPvST-DTvu_0f) |
| 248 | [▪ Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges (Patel et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.02032&sa=D&source=editors&ust=1779048536931323&usg=AOvVaw38K1IsWQE6f5uU13qgrAul) |
| 249 | [▪ GAVEL: Towards rule-based safety through activation monitoring (Rozenfeld et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19768&sa=D&source=editors&ust=1779048536931404&usg=AOvVaw2NJzpSVUpTayeLICIm31y7) |
| 250 | [▪ RvB: Automating AI System Hardening via Iterative Red-Blue Games (Huang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19726&sa=D&source=editors&ust=1779048536931483&usg=AOvVaw22Y1Os5ywKDm63xViz1Pdc) |
| 251 | [▪ LLM-Assisted Authentication and Fraud Detection (Chan and Chan, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19684&sa=D&source=editors&ust=1779048536931561&usg=AOvVaw3RYUqamdrJg5q5L4qHDR4W) |
| 252 | [▪ SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks (Sivaroopan et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19174&sa=D&source=editors&ust=1779048536931643&usg=AOvVaw2nlS159OCmfCTdmssWQDSp) |
| 253 | [▪ Proactive Hardening of LLM Defenses with HASTE (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19051&sa=D&source=editors&ust=1779048536931721&usg=AOvVaw2erYucr7Ey6dRGinR7vTx2) |
| 254 | [▪ Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models (Chu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15801&sa=D&source=editors&ust=1779048536931803&usg=AOvVaw0uGUo5XPRzd0E2rWVR3K7p) |
| 255 | [▪ $\\alpha^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks (Ferrag, Lakas, and Debbah, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.18754&sa=D&source=editors&ust=1779048536931889&usg=AOvVaw1yNPRSugbBisVnd4r-OB8n) |
| 256 | [▪ Mitigating the OWASP Top 10 For Large Language Models Applications using Intelligent Agents (Fasha et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.18105&sa=D&source=editors&ust=1779048536931993&usg=AOvVaw2IL7tsZWSZwZpnNQFBfXai) |
| 257 | [▪ PatchIsland: Orchestration of LLM Agents for Continuous Vulnerability Repair (Kim et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17471&sa=D&source=editors&ust=1779048536932074&usg=AOvVaw1Qdygv9bB2GHuoZwsVG57J) |
| 258 | [▪ ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMs (Fayyazi et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.17406&sa=D&source=editors&ust=1779048536932155&usg=AOvVaw3ndxJM9I2NjjiIgsv5qNbJ) |
| 259 | [▪ Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks (Tsiolakis et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.16150&sa=D&source=editors&ust=1779048536932237&usg=AOvVaw1RDK241O-PwPFajXMa4QUl) |
| 260 | [▪ PAL\*M: Property Attestation for Large Generative Models (Chantasantitam et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16199&sa=D&source=editors&ust=1779048536932330&usg=AOvVaw3i9iRSlVfvkkk3FJ50xbQ1) |
| 261 | [▪ Introducing the Generative Application Firewall (GAF) (Farreny et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15824&sa=D&source=editors&ust=1779048536932419&usg=AOvVaw1EhpVpTY0tW1JioCgrwsY_) |
| 262 | [▪ The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses (G{\\l}uch et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.08864&sa=D&source=editors&ust=1779048536932504&usg=AOvVaw1tSndOvO67MM6877EYom-I) |
| 263 | [▪ Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM) (Biswas and Talukdar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14298&sa=D&source=editors&ust=1779048536932587&usg=AOvVaw1xsPgDqa-mlSUKUmxQjGg5) |
| 264 | [▪ Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification (Meng et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21099&sa=D&source=editors&ust=1779048536932672&usg=AOvVaw3QM8Z-INY2x6cmo2TvIABo) |
| 265 | [▪ HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13864&sa=D&source=editors&ust=1779048536932754&usg=AOvVaw3nWmPE5GsylCjKmSbp2bs2) |
| 266 | [▪ Serverless AI Security: Attack Surface Analysis and Runtime Protection Mechanisms for FaaS-Based Machine Learning (Pathade et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11664&sa=D&source=editors&ust=1779048536932840&usg=AOvVaw0kIYdo34Ml4UXaWmZAwUAA) |
| 267 | [▪ SecMLOps: A Comprehensive Framework for Integrating Security Throughout the MLOps Lifecycle (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10848&sa=D&source=editors&ust=1779048536932924&usg=AOvVaw1TkRKKfU7TVqrWp9iRqR-8) |
| 268 | [▪ Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10589&sa=D&source=editors&ust=1779048536933007&usg=AOvVaw3YtRZAAd-_2oSxsoqi19uE) |
| 269 | [▪ Blue Teaming Function-Calling Agents (Dolcetti, Zizzo, and Maffeis, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09292&sa=D&source=editors&ust=1779048536933089&usg=AOvVaw2kN2WqNh6dsmMbOYbgdqKz) |
| 270 | [▪ Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis (Song et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08196&sa=D&source=editors&ust=1779048536933171&usg=AOvVaw1sVRt6uxJe2XZYXUOmBurT) |
| 271 | [▪ AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation (Kumar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06197&sa=D&source=editors&ust=1779048536933254&usg=AOvVaw3CzOPPsVxhLZqah9Jwu-ER) |
| 272 | [▪ How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape (Kelley et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06033&sa=D&source=editors&ust=1779048536933343&usg=AOvVaw3rwrPIsRd-PDA4hq3-kdaW) |
| 273 | [▪ BlindU: Blind Machine Unlearning without Revealing Erasing Data (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07214&sa=D&source=editors&ust=1779048536933426&usg=AOvVaw1z5lCg_0jSxDJp15_9P6Pz) |
| 274 | [▪ Defenses Against Prompt Attacks Learn Surface Heuristics (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07185&sa=D&source=editors&ust=1779048536933506&usg=AOvVaw1VN8crxoPC2rYLCrUWtk0N) |
| 275 | [▪ Safe-FedLLM: Delving into the Safety of Federated Large Language Models (Tao et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07177&sa=D&source=editors&ust=1779048536933588&usg=AOvVaw25F0u8NH-SoGIqh8TO6d_z) |
| 276 | [▪ United We Defend: Collaborative Membership Inference Defenses in Federated Learning (Bai et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06866&sa=D&source=editors&ust=1779048536933669&usg=AOvVaw0k86LRi_DTnp9GrMG6alND) |
| 277 | [▪ Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model (M et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06213&sa=D&source=editors&ust=1779048536933754&usg=AOvVaw0JxqeOgwX1QQNuTiNDs8M_) |
| 278 | [▪ AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06177&sa=D&source=editors&ust=1779048536933839&usg=AOvVaw1Gq7Sr5-ZuJGL0JNYgmuYQ) |
| 279 | [▪ VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit (Lin et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05755&sa=D&source=editors&ust=1779048536933921&usg=AOvVaw2g6_Vuy2LmrsgKG8BAtGph) |
| 280 | [▪ HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04034&sa=D&source=editors&ust=1779048536934003&usg=AOvVaw3z0WRN4FqnEU9secFUzLmC) |
| 281 | [▪ AI-Driven Cybersecurity Threats: A Survey of Emerging Risks and Defensive Strategies (Erukude, Marella, and Veluru, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03304&sa=D&source=editors&ust=1779048536934088&usg=AOvVaw1yTjxdQ_PlFNgkUnAGDut9) |
| 282 | [▪ Autonomous Threat Detection and Response in Cloud Security: A Comprehensive Survey of AI-Driven Strategies (Sarraf and Pal, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03303&sa=D&source=editors&ust=1779048536934171&usg=AOvVaw1Q28tpKWWI-0rCDCeNWQvZ) |
| 283 | [▪ Automated Post-Incident Policy Gap Analysis via Threat-Informed Evidence Mapping using Large Language Models (Oh et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03287&sa=D&source=editors&ust=1779048536934256&usg=AOvVaw1dVGWyvo0zhSMNbbHpD4R-) |
| 284 | [▪ SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones (Wu and Cao, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03242&sa=D&source=editors&ust=1779048536934342&usg=AOvVaw1oOe3cDrwVHcSQ-6pyEJUq) |
| 285 | [▪ How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System (Zuo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.17215&sa=D&source=editors&ust=1779048536934427&usg=AOvVaw3Hv5QyU5GHbAS1Oca6k0D8) |
| 286 | [▪ Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment (Xie et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05755&sa=D&source=editors&ust=1779048536934509&usg=AOvVaw32nGd_dfXicq8fc8ahesJf) |
| 287 | [▪ Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models (Rosario, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00848&sa=D&source=editors&ust=1779048536934591&usg=AOvVaw2lnGmn4BkdC8366do7QObJ) |
| 288 | [▪ Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection (Iyer, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01723&sa=D&source=editors&ust=1779048536934671&usg=AOvVaw35u9hW6-S3ZpSZjN0Ghv7Q) |
| 289 | [▪ One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models (Gu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.07167&sa=D&source=editors&ust=1779048536934754&usg=AOvVaw3nKM1YbnhKGaJqVUMKGpiR) |
| 290 | [▪ Improving LLM-Assisted Secure Code Generation through Retrieval-Augmented-Generation and Multi-Tool Feedback (Sriram et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00509&sa=D&source=editors&ust=1779048536934839&usg=AOvVaw1HakMSg_prbrSwJrvgX6Y_) |
| 291 | [▪ PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices (Chattopadhyay et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00367&sa=D&source=editors&ust=1779048536934922&usg=AOvVaw0sp668g_JB0-n59lTESNSi) |
| 292 | [▪ Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing (Bhatt et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00042&sa=D&source=editors&ust=1779048536935001&usg=AOvVaw2RbB0GhA6H8AJ_dpYwqFws) |
| 293 | [▪ Towards Provably Secure Generative AI: Reliable Consensus Sampling (Cui et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.24925&sa=D&source=editors&ust=1779048536935080&usg=AOvVaw1NEdZ6-BnJHj4YopUOraPB) |
| 294 | [▪ How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench (Sajadi, Damevski, and Chatterjee, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02976&sa=D&source=editors&ust=1779048536935166&usg=AOvVaw0qzjSgA4YZDRcnR-rmcTqh) |
| 295 | [▪ Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses (Wang, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22128&sa=D&source=editors&ust=1779048536935248&usg=AOvVaw20Ru4D1rp5KmYPDBZx45N0) |
| 296 | [▪ Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems (Foundjem et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23132&sa=D&source=editors&ust=1779048536935337&usg=AOvVaw1Ju4ZxRPJX9swxrQNdDK41) |
| 297 | [▪ Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management (Arora and Hastings, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22060&sa=D&source=editors&ust=1779048536935427&usg=AOvVaw2oYw0yUzpgGCNw8jpVaKjS) |
| 298 | [▪ Assessing the Software Security Comprehension of Large Language Models (Siddiq et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21238&sa=D&source=editors&ust=1779048536935514&usg=AOvVaw2679_t4EL9DY0ma2eGjtRL) |
| 299 | [▪ AIAuditTrack: A Framework for AI Security system (Luo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20649&sa=D&source=editors&ust=1779048536935595&usg=AOvVaw0jm1Tu9XS1Lk-XuxiS5f-o) |
| 300 | [▪ AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20986&sa=D&source=editors&ust=1779048536935675&usg=AOvVaw0n3Ru8kZaKR4mwvBAcShgU) |
| 301 | [▪ On the Effectiveness of Instruction-Tuning Local LLMs for Identifying Software Vulnerabilities (Park, Ko, and Cho, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20062&sa=D&source=editors&ust=1779048536935757&usg=AOvVaw37UfswIr4RdCE_fTs5bxgr) |
| 302 | [▪ IoT-based Android Malware Detection Using Graph Neural Network With Adversarial Defense (Yumlembam et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20004&sa=D&source=editors&ust=1779048536935841&usg=AOvVaw3lxFW6WbmKNCao217D4JCR) |
| 303 | [▪ Certified Defense on the Fairness of Graph Neural Networks (Dong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2311.02757&sa=D&source=editors&ust=1779048536935921&usg=AOvVaw34kW-QlTYHdXMmCQe6ZhCv) |
| 304 | [▪ DREAM: Dynamic Red-teaming across Environments for AI Models (Lu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19016&sa=D&source=editors&ust=1779048536936000&usg=AOvVaw24qUB-rUc4EWlsz4iu-SA9) |
| 305 | [▪ Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection (Pan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18733&sa=D&source=editors&ust=1779048536936082&usg=AOvVaw3cwjhGZ0WTLXvw3Sb_WGzx) |
| 306 | [▪ AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs (Gaikwad, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17251&sa=D&source=editors&ust=1779048536936161&usg=AOvVaw1tHB15lgmFfv37h9APnC1K) |
| 307 | [▪ Prefix Probing: Lightweight Harmful Content Detection for Large Language Models (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16650&sa=D&source=editors&ust=1779048536936243&usg=AOvVaw3GicEVQmCo-xMhjBuLtGbM) |
| 308 | [▪ Adversarial Robustness in Financial Machine Learning: Defenses, Economic Impact, and Governance Evidence (Baviskar, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15780&sa=D&source=editors&ust=1779048536936333&usg=AOvVaw0eJNR2Da43dZkFOBy-23wT) |
| 309 | [▪ PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy (Kundu, Ahmed, and Hoque, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16851&sa=D&source=editors&ust=1779048536936417&usg=AOvVaw2vkB6joO4njHcJrq2TQFKq) |
| 310 | [▪ Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks (Shaheer et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16307&sa=D&source=editors&ust=1779048536936497&usg=AOvVaw1RITIIMmj_rH9Z85QBAxfU) |
| 311 | [▪ Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection (Song et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16123&sa=D&source=editors&ust=1779048536936578&usg=AOvVaw2VWUOys1Lqlh9jV1Ol7u4b) |
| 312 | [▪ Auto-Tuning Safety Guardrails for Black-Box Large Language Models (Abdulkadir, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15782&sa=D&source=editors&ust=1779048536936656&usg=AOvVaw0C3FMc0wghlgePNnu4yEX8) |
| 313 | [▪ Quantifying Return on Security Controls in LLM Systems (Moulton, O'Brien, and Hastings, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15081&sa=D&source=editors&ust=1779048536936736&usg=AOvVaw0r0BR-WHWZYFUqd75SMBod) |
| 314 | [▪ GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients (Razmjoo, Sharifian, and Shouraki, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12827&sa=D&source=editors&ust=1779048536936818&usg=AOvVaw2zO7RHsF1mfODfO0K7Fz_W) |
| 315 | [▪ Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS (Ennaji, Benkhelifa, and Mancini, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13501&sa=D&source=editors&ust=1779048536936904&usg=AOvVaw3k3T4EF-nMsWkjeFBIYQEm) |
| 316 | [▪ SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models (Afane et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10998&sa=D&source=editors&ust=1779048536936985&usg=AOvVaw2C97TCFMVFOckJBzvlx3St) |
| 317 | [▪ Toward Intelligent and Secure Cloud: Large Language Model Empowered Proactive Defense (Zhou et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.21051&sa=D&source=editors&ust=1779048536937066&usg=AOvVaw25PvXLZPndDZPDqEBMyoEf) |
| 318 | [▪ CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models (Hall, Ungaro, and Eiers, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09957&sa=D&source=editors&ust=1779048536937149&usg=AOvVaw1On7rB_0mkhMdhmEQ6Q2JA) |
| 319 | [▪ Adaptive Intrusion Detection System Leveraging Dynamic Neural Models with Adversarial Learning for 5G/6G Networks (Neha and Bhatia, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10637&sa=D&source=editors&ust=1779048536937235&usg=AOvVaw1fjONXw43DPmj5a6IlOGex) |
| 320 | [▪ From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection (Lu and Lagaisse, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10485&sa=D&source=editors&ust=1779048536937325&usg=AOvVaw27vkSiEf2Xcq2zolNBAasr) |
| 321 | [▪ Chasing Shadows: Pitfalls in LLM Security Research (Evertz et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09549&sa=D&source=editors&ust=1779048536937405&usg=AOvVaw2n4abnVlHFLojCnW95XyDJ) |
| 322 | [▪ AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.18123&sa=D&source=editors&ust=1779048536937485&usg=AOvVaw0Ekv6-eNbtOTD2Ppz6IGCC) |
| 323 | [▪ Evaluating the robustness of adversarial defenses in malware detection systems (Jafari and Shameli-Sendi, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.09342&sa=D&source=editors&ust=1779048536937566&usg=AOvVaw3OYCY1BrxX2R3qprUzu4Lg) |
| 324 | [▪ Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples (Zhou et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06123&sa=D&source=editors&ust=1779048536937649&usg=AOvVaw0WJsSukZIMHaBK5o2oDo1g) |
| 325 | [▪ VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection (Nie et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07533&sa=D&source=editors&ust=1779048536937730&usg=AOvVaw1frJtHHaTgUi7TszA1GZ12) |
| 326 | [▪ Securing the Model Context Protocol: Defending LLMs Against Tool Poisoning and Adversarial Attacks (Jamshidi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06556&sa=D&source=editors&ust=1779048536937812&usg=AOvVaw3P6SKoiqv87JbL1Oxf5_Rb) |
| 327 | [▪ IF-GUIDE: Influence Function-Guided Detoxification of LLMs (Coalson et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.01790&sa=D&source=editors&ust=1779048536937891&usg=AOvVaw14d4RYls6N1ANh_zsYWsS3) |
| 328 | [▪ Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model (Liu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17162&sa=D&source=editors&ust=1779048536937974&usg=AOvVaw12UnTX0LxygDZUmnKEj3fy) |
| 329 | [▪ ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior (Lu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.05745&sa=D&source=editors&ust=1779048536938056&usg=AOvVaw2HVs4TyfekibGVjya2y9e8) |
| 330 | [▪ Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models (Cretu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.05707&sa=D&source=editors&ust=1779048536938139&usg=AOvVaw1gKBcWMa2XmNU1q2eEyxHc) |
| 331 | [▪ AutoGuard: A Self-Healing Proactive Security Layer for DevSecOps Pipelines Using Reinforcement Learning (Anugula et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04368&sa=D&source=editors&ust=1779048536938222&usg=AOvVaw2ENbZoJ7rsPiOcolz6l4Lc) |
| 332 | [▪ Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs (Ma et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03720&sa=D&source=editors&ust=1779048536938316&usg=AOvVaw15xW9NSVddAm9HmyVYntU-) |
| 333 | [▪ Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks (Fu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03100&sa=D&source=editors&ust=1779048536938406&usg=AOvVaw1eEY6Z0Lb4kZdA07vt9lEx) |
| 334 | [▪ VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs (Liao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21393&sa=D&source=editors&ust=1779048536938489&usg=AOvVaw1qQRtOA7_n9i8gxxoxr3MZ) |
| 335 | [▪ OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning (Zhu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02306&sa=D&source=editors&ust=1779048536938569&usg=AOvVaw3PUT1QYFGv2Gtp3WFl9Q7w) |
| 336 | [▪ COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02318&sa=D&source=editors&ust=1779048536938649&usg=AOvVaw10dR27CdmyUnX87jPg1wHA) |
| 337 | [▪ Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI (Sandoval and Rushing, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02157&sa=D&source=editors&ust=1779048536938730&usg=AOvVaw1pWjkV8IB2HLRfzxOmPeIY) |
| 338 | [▪ Large Language Model based Smart Contract Auditing with LLMBugScanner (Yuan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02069&sa=D&source=editors&ust=1779048536938812&usg=AOvVaw2-BqkU8uOELjLryFYovNK5) |
| 339 | [▪ Teleportation-Based Defenses for Privacy in Approximate Machine Unlearning (Maheri et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00272&sa=D&source=editors&ust=1779048536938892&usg=AOvVaw2VhcGlCveBQflmPG0k3CEh) |
| 340 | [▪ Benchmarking and Understanding Safety Risks in AI Character Platforms (Wei, Zhang, and Tyson, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01247&sa=D&source=editors&ust=1779048536938973&usg=AOvVaw3PgUKImu7E9-1T5ejf68PL) |
| 341 | [▪ Red Teaming Large Reasoning Models (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00412&sa=D&source=editors&ust=1779048536939050&usg=AOvVaw3cuks7pbrlo3veFL_fkaWm) |
| 342 | [▪ An Empirical Study on the Security Vulnerabilities of GPTs (Wu, Wu, and Zheng, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00136&sa=D&source=editors&ust=1779048536939129&usg=AOvVaw2o91xXrVUbm65jSACJvJYd) |
| 343 | [▪ ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning (Chen, Kang, and Li, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.22738&sa=D&source=editors&ust=1779048536939210&usg=AOvVaw2QfCWSO2iR2iULfPakIJzk) |
| 344 | [▪ A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures (Kong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.19676&sa=D&source=editors&ust=1779048536939311&usg=AOvVaw13EjbwaK1s5E3zUM1X0am7) |
| 345 | [▪ Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities (Garg et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.23408&sa=D&source=editors&ust=1779048536939398&usg=AOvVaw2FlT3P-p-MydPhwpBUJQJp) |
| 346 | [▪ Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks (Young, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22047&sa=D&source=editors&ust=1779048536939485&usg=AOvVaw3cjI9VBdVhNr9Hfnh_PlVZ) |
| 347 | [▪ Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance (Huwyler, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21901&sa=D&source=editors&ust=1779048536939567&usg=AOvVaw2tIAWdDmSqWVU9F4JoSeEC) |
| 348 | [▪ GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision (Xiang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20994&sa=D&source=editors&ust=1779048536939651&usg=AOvVaw010IIKBCxAy97gHLKVyktS) |
| 349 | [▪ DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation (Pathak et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20709&sa=D&source=editors&ust=1779048536939735&usg=AOvVaw1rzs67NbkdigjLNhRWPUdy) |
| 350 | [▪ VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00882&sa=D&source=editors&ust=1779048536939814&usg=AOvVaw1wUpgKL5QAenzy4mlfcLq1) |
| 351 | [▪ SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models (Alam et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19558&sa=D&source=editors&ust=1779048536939897&usg=AOvVaw3kfLEaMWbWgUoYb90MJ7zT) |
| 352 | [▪ EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering (Gungor et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19523&sa=D&source=editors&ust=1779048536939980&usg=AOvVaw1nWq9eVejQ1Hn2KL-PdalR) |
| 353 | [▪ Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19218&sa=D&source=editors&ust=1779048536940064&usg=AOvVaw2xKa5HHrspRAX1U8fhyMNi) |
| 354 | [▪ Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18467&sa=D&source=editors&ust=1779048536940149&usg=AOvVaw2-DebN54r_aJIiFC5QvxhH) |
| 355 | [▪ SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense (Krukowski et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.08255&sa=D&source=editors&ust=1779048536940230&usg=AOvVaw1j27pmiXJ_MeGBZ4VNf2e_) |
| 356 | [▪ Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks (Ji et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15203&sa=D&source=editors&ust=1779048536940316&usg=AOvVaw2QD_a7D-anj4bX1oyAnxD2) |
| 357 | [▪ DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00447&sa=D&source=editors&ust=1779048536940400&usg=AOvVaw1yKXbCLRyK6CDSh2OWj-2l) |
| 358 | [▪ N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator (Lin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14195&sa=D&source=editors&ust=1779048536940480&usg=AOvVaw197fBVtlPUTqk6bJkY7Xcq) |
| 359 | [▪ Certified but Fooled! Breaking Certified Defences with Ghost Certificates (Vo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14003&sa=D&source=editors&ust=1779048536940563&usg=AOvVaw3uxa-Ro6IzRrmrhDRQ5Q_I) |
| 360 | [▪ ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning (Guan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13771&sa=D&source=editors&ust=1779048536940647&usg=AOvVaw1kgaHFIDYYWqLneCh7p478) |
| 361 | [▪ AI Kill Switch for malicious web-based LLM agent (Lee and Park, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13725&sa=D&source=editors&ust=1779048536940726&usg=AOvVaw2P35Qfducu10M2V0B7s4eC) |
| 362 | [▪ DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.09353&sa=D&source=editors&ust=1779048536940808&usg=AOvVaw2mmnyPU0Qvm9YZoTpdIpuz) |
| 363 | [▪ Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning (Zimmer and Karame, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13654&sa=D&source=editors&ust=1779048536940893&usg=AOvVaw3hgeGMlLCrZ3e9-eiCAcRI) |
| 364 | [▪ SGuard-v1: Safety Guardrail for Large Language Models (Lee et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12497&sa=D&source=editors&ust=1779048536940972&usg=AOvVaw2tIQxPFOvzmW7NxIu-KyDc) |
| 365 | [▪ Defending Unauthorized Model Merging via Dual-Stage Weight Protection (Chen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11851&sa=D&source=editors&ust=1779048536941057&usg=AOvVaw2ajh3ToIk2JmHFdyzXOF2D) |
| 366 | [▪ On the Trade-Off Between Transparency and Security in Adversarial Machine Learning (Fenaux, Srinivasa, and Kerschbaum, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11842&sa=D&source=editors&ust=1779048536941141&usg=AOvVaw3NDEarHkB8aMuJbR1eal6x) |
| 367 | [▪ TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13717&sa=D&source=editors&ust=1779048536941223&usg=AOvVaw2K4mKnkuZEvDL32Y0Ttmj6) |
| 368 | [▪ InfoDecom: Decomposing Information for Defending against Privacy Leakage in Split Inference (Deng, Lu, and Duan, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13365&sa=D&source=editors&ust=1779048536941311&usg=AOvVaw3VnuJwXqCmwaCsRAwWle7I) |
| 369 | [▪ DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13248&sa=D&source=editors&ust=1779048536941393&usg=AOvVaw3Nk22Mzkr2yovVLhMfopUc) |
| 370 | [▪ SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization (Rong et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12982&sa=D&source=editors&ust=1779048536941473&usg=AOvVaw3EAQZBIp5pePyl77E8OMsG) |
| 371 | [▪ AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework (Nathanson et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12668&sa=D&source=editors&ust=1779048536941556&usg=AOvVaw3ss2nKS12HYnu1OJ9tif5G) |
| 372 | [▪ VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization (Li, Yu, and Wang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11896&sa=D&source=editors&ust=1779048536941637&usg=AOvVaw1OznBQ92fOWhTRoHyA2AKy) |
| 373 | [▪ Securing Generative AI in Healthcare: A Zero-Trust Architecture Powered by Confidential Computing on Google Cloud (Amanna and Shinde, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11836&sa=D&source=editors&ust=1779048536941722&usg=AOvVaw3iOAXVWWT6enraFa_iF586) |
| 374 | [▪ PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities (Wei et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11019&sa=D&source=editors&ust=1779048536941803&usg=AOvVaw2vmPnP35m2M2Z_ztoNHQd0) |
| 375 | [▪ Rethinking the Evaluation of Secure Code Generation (Dai, Xu, and Tao, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.15554&sa=D&source=editors&ust=1779048536941882&usg=AOvVaw1xBscVw1BUIWmPehLKkLHQ) |
| 376 | [▪ AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness (Huang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09316&sa=D&source=editors&ust=1779048536941984&usg=AOvVaw2O5ZfcPFngWpm_RdR1ENRd) |
| 377 | [▪ iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification (Xiong et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08905&sa=D&source=editors&ust=1779048536942066&usg=AOvVaw2nRGq0FbjpJZXc7rhsMlBV) |
| 378 | [▪ Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement (Ma et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07741&sa=D&source=editors&ust=1779048536942150&usg=AOvVaw0gAQH-ywqwnqbCw2IeDF-l) |
| 379 | [▪ Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.08875&sa=D&source=editors&ust=1779048536942234&usg=AOvVaw21Fbywt8TKEM_z8Teel_Y2) |
| 380 | [▪ Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models (Ahn and Jung, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18877&sa=D&source=editors&ust=1779048536942324&usg=AOvVaw0xTkrdhHfflmCyyJVHdGfz) |
| 381 | [▪ Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks (Chen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02735&sa=D&source=editors&ust=1779048536942407&usg=AOvVaw1ScW4ZDkTM28iJaASLcWp7) |
| 382 | [▪ LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors (Lu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07192&sa=D&source=editors&ust=1779048536942488&usg=AOvVaw1RlI1bsBGIrMMeRDdGAa9t) |
| 383 | [▪ Efficient LLM Safety Evaluation through Multi-Agent Debate (Lin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06396&sa=D&source=editors&ust=1779048536942567&usg=AOvVaw0ZF7K6IEzc3jB9FcG-xMKy) |
| 384 | [▪ Preserving security in a world with powerful AI Considerations for the future Defense Architecture (Generous, Cook, and Pruet, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05714&sa=D&source=editors&ust=1779048536942649&usg=AOvVaw1jVEE__c3hkczR78C2t6IW) |
| 385 | [▪ Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks (Waheed et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.12386&sa=D&source=editors&ust=1779048536942730&usg=AOvVaw33-DM44u-RfA2Z4iGh1P88) |
| 386 | [▪ ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (Gomaa, Salem, and Abdelnabi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05359&sa=D&source=editors&ust=1779048536942815&usg=AOvVaw3aQLD_n-I9LJoUzNJ1Y4TQ) |
| 387 | [▪ Specification-Guided Vulnerability Detection with Large Language Models (Zhu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04014&sa=D&source=editors&ust=1779048536942896&usg=AOvVaw0EC7rQzkbXewUOZTO3YgcK) |
| 388 | [▪ SME-TEAM: Leveraging Trust and Ethics for Secure and Responsible Use of AI and LLMs in SMEs (Sarker et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.10594&sa=D&source=editors&ust=1779048536942981&usg=AOvVaw02SmBi9l14VNAUSMMiq99J) |
| 389 | [▪ LLM-Driven SAST-Genius: A Hybrid Static Analysis Framework for Comprehensive and Actionable Security (Agrawal and Ahi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.15433&sa=D&source=editors&ust=1779048536943064&usg=AOvVaw1ygzD2Isgi1l0C6QRXsw_L) |
| 390 | [▪ SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories (Shen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21205&sa=D&source=editors&ust=1779048536943148&usg=AOvVaw1X-nfeT0UT7KFqoD6CQvhu) |
| 391 | [▪ LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models (Tahmasivand et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02866&sa=D&source=editors&ust=1779048536943229&usg=AOvVaw1Qt3dZ_shtO6uFIFqfdT5q) |
| 392 | [▪ RepoMark: A Data-Usage Auditing Framework for Code Large Language Models (Qu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21432&sa=D&source=editors&ust=1779048536943314&usg=AOvVaw3SVVPKaYTtmp_yRnB0Lcxl) |
| 393 | [▪ AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding (Anwar and Liu, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00265&sa=D&source=editors&ust=1779048536943399&usg=AOvVaw0LRUU4nk-Pr1snUJN12qJY) |
| 394 | [▪ Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations (Okatan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00973&sa=D&source=editors&ust=1779048536943481&usg=AOvVaw0U1KEfbz1nv3dUWGF4weal) |
| 395 | [▪ DRIP: Defending Prompt Injection via De-instruction Training and Residual Fusion Model Architecture (Liu, Lin, and Dong, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00447&sa=D&source=editors&ust=1779048536943566&usg=AOvVaw0NGKXfpoF0Z0owhWpkVbBB) |
| 396 | [▪ SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents (Yin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.13178&sa=D&source=editors&ust=1779048536943645&usg=AOvVaw0fCZN3AP5taq3DrNPsunY2) |
| 397 | [▪ On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection (Hannan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27675&sa=D&source=editors&ust=1779048536943725&usg=AOvVaw3U9mdKPPiotYq_CVKgWeQk) |
| 398 | [▪ SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation (Su et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26830&sa=D&source=editors&ust=1779048536943806&usg=AOvVaw1CFd9rGHL4ieRtxcW8OIeZ) |
| 399 | [▪ LLM-based Multi-class Attack Analysis and Mitigation Framework in IoT/IIoT Networks (Ikbarieh, Gupta, and Mahalal, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26941&sa=D&source=editors&ust=1779048536943888&usg=AOvVaw2A0Z9rznrjiSO_6Ps8xjQl) |
| 400 | [▪ The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.11094&sa=D&source=editors&ust=1779048536943976&usg=AOvVaw2lLfC9ARQtL29AtzXSJ_ZR) |
| 401 | [▪ Improving LLM Safety Alignment with Dual-Objective Optimization (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.03710&sa=D&source=editors&ust=1779048536944058&usg=AOvVaw0nIB4rHcxBQaA-vOdS1rcd) |
| 402 | [▪ IRCopilot: Automated Incident Response with Large Language Models (Lin et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.20945&sa=D&source=editors&ust=1779048536944137&usg=AOvVaw1kqjvKwxa-auD_8pVp16ZF) |
| 403 | [▪ ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models (Jin et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26096&sa=D&source=editors&ust=1779048536944220&usg=AOvVaw00JtyDy-NZ5eJUR2iidAZs) |
| 404 | [▪ Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control (Cai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26212&sa=D&source=editors&ust=1779048536944307&usg=AOvVaw1NGUaC0PhX6oKVK2yE4-LQ) |
| 405 | [▪ SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning (Zhou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26037&sa=D&source=editors&ust=1779048536944388&usg=AOvVaw06U5XpVjxnZ0xLrCeOGt1E) |
| 406 | [▪ SoK: Honeypots & LLMs, More Than the Sum of Their Parts? (Bridges et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25939&sa=D&source=editors&ust=1779048536944469&usg=AOvVaw0Yo_riLHTkCuCFr2c_925r) |
| 407 | [▪ OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models (Wang and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19169&sa=D&source=editors&ust=1779048536944554&usg=AOvVaw3s2pqVeN0Tm_O8knJ5Ny19) |
| 408 | [▪ Secure Retrieval-Augmented Generation against Poisoning Attacks (Cheng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25025&sa=D&source=editors&ust=1779048536944633&usg=AOvVaw2ArpKl6NdgP6mRyPN3fKuB) |
| 409 | [▪ FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models (Nazari et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24985&sa=D&source=editors&ust=1779048536944716&usg=AOvVaw1IVAlw93PmnDygt3wa3YTr) |
| 410 | [▪ SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability (Xu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23960&sa=D&source=editors&ust=1779048536944798&usg=AOvVaw0bz2zDFnvjHAxSN59uxja5) |
| 411 | [▪ Adversarially-Aware Architecture Design for Robust Medical AI Systems (Gerhart and Iyangar, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23622&sa=D&source=editors&ust=1779048536944878&usg=AOvVaw0ZRjmD0wuED2V4j-BSrwLJ) |
| 412 | [▪ Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents (Sanz-G\\'omez et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24317&sa=D&source=editors&ust=1779048536944964&usg=AOvVaw2hpphgMSsXAqv8xaIS6BMG) |
| 413 | [▪ SAGE: A Generic Framework for LLM Safety Evaluation (Jindal et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.19674&sa=D&source=editors&ust=1779048536945043&usg=AOvVaw32D6zhuXaFCYZMsaU-X_VU) |
| 414 | [▪ T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22300&sa=D&source=editors&ust=1779048536945125&usg=AOvVaw0qqLH3J9G9IPjdE6vkn3s3) |
| 415 | [▪ SecureLearn - An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks (Paracha et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22274&sa=D&source=editors&ust=1779048536945208&usg=AOvVaw2YS20FNhupclcmP_NJp_U-) |
| 416 | [▪ Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs (Davies et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.14828&sa=D&source=editors&ust=1779048536945296&usg=AOvVaw375waBKo08jlqsqEI2_g9x) |
| 417 | [▪ SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots (Adebimpe, Neukirchen, and Welsh, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21459&sa=D&source=editors&ust=1779048536945380&usg=AOvVaw3GZBcuzW2QlYrDZw_s8bBV) |
| 418 | [▪ FLAMES: Fine-tuning LLMs to Synthesize Invariants for Smart Contract Security (Eshghie et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21401&sa=D&source=editors&ust=1779048536945464&usg=AOvVaw1IugxaxJdNEr1OoRLkXyMQ) |
| 419 | [▪ Soft Instruction De-escalation Defense (Walter et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21057&sa=D&source=editors&ust=1779048536945543&usg=AOvVaw2wp-XdbhX8HrEXeLs6IYgR) |
| 420 | [▪ Towards Strong Certified Defense with Universal Asymmetric Randomization (Hong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19977&sa=D&source=editors&ust=1779048536945623&usg=AOvVaw3bgcAKTGfMysm_wmoanAfQ) |
| 421 | [▪ Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses (Yichao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20314&sa=D&source=editors&ust=1779048536945707&usg=AOvVaw1sqBtKafMNvkUZYNbEp-pi) |
| 422 | [▪ SAID: Empowering Large Language Models with Self-Activating Internal Defense (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20129&sa=D&source=editors&ust=1779048536945788&usg=AOvVaw1nRZJgxyn38UddHm_u3V41) |
| 423 | [▪ SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks (Lee et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.11791&sa=D&source=editors&ust=1779048536945868&usg=AOvVaw2Uq0-eRyme0l7vkt6S2NZd) |
| 424 | [▪ AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices (Zhan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19462&sa=D&source=editors&ust=1779048536945950&usg=AOvVaw0mb93qEGkEMraWFah9DrY3) |
| 425 | [▪ Monitoring LLM-based Multi-Agent Systems Against Corruptions via Node Evaluation (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19420&sa=D&source=editors&ust=1779048536946030&usg=AOvVaw0E-A2HAGjtH-Gk4VmuF-_2) |
| 426 | [▪ Defending Against Prompt Injection with DataFilter (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19207&sa=D&source=editors&ust=1779048536946109&usg=AOvVaw2ieNgKaRpUEk0_YIx8f8na) |
| 427 | [▪ OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform (Wang and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.19169&sa=D&source=editors&ust=1779048536946188&usg=AOvVaw04Sru2fIO2IEi_-wioA018) |
| 428 | [▪ Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization (Haddad et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18508&sa=D&source=editors&ust=1779048536946270&usg=AOvVaw2eOlO-FGVQmvDTVq8cmeWL) |
| 429 | [▪ SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering (Lin et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17633&sa=D&source=editors&ust=1779048536946367&usg=AOvVaw35dNrmFJQdhm9TSSVzwODV) |
| 430 | [▪ Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems (Jha et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17276&sa=D&source=editors&ust=1779048536946451&usg=AOvVaw2E7sRNpA-csEf5N3D_E0A_) |
| 431 | [▪ VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search (Li, Zhao, and Liu, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15948&sa=D&source=editors&ust=1779048536946533&usg=AOvVaw1jZY8EymYtey1DhpQ1WEkt) |
| 432 | [▪ Watermark Robustness and Radioactivity May Be at Odds in Federated Learning (Huang, Shao, and Baluta, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17033&sa=D&source=editors&ust=1779048536946617&usg=AOvVaw07pdkszwOnfcMxFxzQKBSM) |
| 433 | [▪ Verifiable Fine-Tuning for LLMs: Zero-Knowledge Training Proofs Bound to Data Provenance and Policy (Akgul et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16830&sa=D&source=editors&ust=1779048536946699&usg=AOvVaw146sGP7OtMVAlamUoYfwNL) |
| 434 | [▪ SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection (Feng and Pan, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16219&sa=D&source=editors&ust=1779048536946781&usg=AOvVaw0npZFZs5w1ZOOg1tsRfa6w) |
| 435 | [▪ Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers (Bertollo, Bodemir, and Burgess, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16005&sa=D&source=editors&ust=1779048536946865&usg=AOvVaw33KOFY9xmHEUDDNKwfVfze) |
| 436 | [▪ A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks (Swanda et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.20639&sa=D&source=editors&ust=1779048536946947&usg=AOvVaw19UjSi4G_D2MT68JN9p5NS) |
| 437 | [▪ GuardReasoner: Towards Reasoning-based LLM Safeguards (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18492&sa=D&source=editors&ust=1779048536947025&usg=AOvVaw1eHVgbT5FcbXIOdI1ml_nL) |
| 438 | [▪ Towards Proactive Defense Against Cyber Cognitive Attacks (Rushing, Umeokolo, and Xu, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15801&sa=D&source=editors&ust=1779048536947104&usg=AOvVaw1hACPF_P-CuKz8cD_QCOss) |
| 439 | [▪ Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13361&sa=D&source=editors&ust=1779048536947185&usg=AOvVaw21qhov2yCGC-DDZPD-gd0R) |
| 440 | [▪ GRIDAI: Generating and Repairing Intrusion Detection Rules via Collaboration among Multiple LLM-based Agents (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13257&sa=D&source=editors&ust=1779048536947267&usg=AOvVaw0njja3NxrcRF5cQDDA2TFL) |
| 441 | [▪ Countermind: A Multi-Layered Security Architecture for Large Language Models (Schwarz, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11837&sa=D&source=editors&ust=1779048536947354&usg=AOvVaw2Cl_aH_E12RemhIjL0MuCp) |
| 442 | [▪ TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11203&sa=D&source=editors&ust=1779048536947436&usg=AOvVaw2-Q6OUn07HW_SAaAWT9_Pk) |
| 443 | [▪ Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10085&sa=D&source=editors&ust=1779048536947520&usg=AOvVaw2i31sGGV5jSjANozd_3wpN) |
| 444 | [▪ SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents (Ying et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10073&sa=D&source=editors&ust=1779048536947600&usg=AOvVaw0QLTFk5ghOHpPpCKd9YQ7r) |
| 445 | [▪ Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices (Patir et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09682&sa=D&source=editors&ust=1779048536947683&usg=AOvVaw1gfrLrg8kGAXe4H2huhiB2) |
| 446 | [▪ Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability (Mo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09620&sa=D&source=editors&ust=1779048536947764&usg=AOvVaw39q9uPoqA6NrwnOQfVl-ap) |
| 447 | [▪ Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models (Meng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21099&sa=D&source=editors&ust=1779048536947844&usg=AOvVaw28qUxT8Gx8XtfJsNNd8EsH) |
| 448 | [▪ From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses (Meng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07968&sa=D&source=editors&ust=1779048536947925&usg=AOvVaw1v49B15WJPs9V0S1cE3KCw) |
| 449 | [▪ RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning (Horal et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06994&sa=D&source=editors&ust=1779048536948006&usg=AOvVaw2ml-SCSb_6kCPnbygpc5T7) |
| 450 | [▪ A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens (Dobre et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16366&sa=D&source=editors&ust=1779048536948088&usg=AOvVaw3GDzV5N0O--j6Uj1I1oW-p) |
| 451 | [▪ DoomArena: A framework for Testing AI Agents Against Evolving Security Threats (Boisvert et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.14064&sa=D&source=editors&ust=1779048536948168&usg=AOvVaw0YuB_Ilbao0k-uVp-d4EST) |
| 452 | [▪ VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation (Miculicich et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05156&sa=D&source=editors&ust=1779048536948250&usg=AOvVaw3SSiOsy0AJpkBbMQ6xIgkc) |
| 453 | [▪ Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling (Llewellyn et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05709&sa=D&source=editors&ust=1779048536948344&usg=AOvVaw3Lk5AZ5HLRw5vHrGRaqLOb) |
| 454 | [▪ SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models (Qi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05173&sa=D&source=editors&ust=1779048536948427&usg=AOvVaw3CSkgOtKx4ak-a-dF5cnpB) |
| 455 | [▪ AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation (Liu and Anwar, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.13118&sa=D&source=editors&ust=1779048536948508&usg=AOvVaw0Sc-qofkHl8tKsuLF5xC5W) |
| 456 | [▪ Thought Purity: A Defense Framework For Chain-of-Thought Attack (Xue et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12314&sa=D&source=editors&ust=1779048536948587&usg=AOvVaw0mGtheQPfGzsiLBEu_D5-c) |
| 457 | [▪ SteerDiff: Steering towards Safe Text-to-Image Diffusion Models (Zhang, He, and Chen, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02710&sa=D&source=editors&ust=1779048536948668&usg=AOvVaw3dpS_vdBPb5wSmpDkvjTxj) |
| 458 | [▪ Proactive defense against LLM Jailbreak (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05052&sa=D&source=editors&ust=1779048536948745&usg=AOvVaw08PGrzvtX9mcHJ9eJRKld6) |
| 459 | [▪ MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection (Nguyen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04397&sa=D&source=editors&ust=1779048536948830&usg=AOvVaw3ibDMTcZXLzdrjt6OY0Wm0) |
| 460 | [▪ SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition (Luo, Yau, and Song, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03319&sa=D&source=editors&ust=1779048536948912&usg=AOvVaw1YKEI5fUZAWyhWg1mFCUD3) |
| 461 | [▪ Permissioned LLMs: Enforcing Access Control in Large Language Models (Jayaraman et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.22860&sa=D&source=editors&ust=1779048536948993&usg=AOvVaw3UJYt0z2TAouLonLpMrbtW) |
| 462 | [▪ A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory (Wei et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02373&sa=D&source=editors&ust=1779048536949075&usg=AOvVaw1CdwacnnZzngeuvZpCLNhI) |
| 463 | [▪ Defend LLMs Through Self-Consciousness (Huang and Paula, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02961&sa=D&source=editors&ust=1779048536949154&usg=AOvVaw3VIf1m_FRdqznf6HkHrVPB) |
| 464 | [▪ UpSafe$^\\circ$C: Upcycling for Controllable Safety in Large Language Models (Sun et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02194&sa=D&source=editors&ust=1779048536949236&usg=AOvVaw3EliFqqSQboaajT8tB18S6) |
| 465 | [▪ TAIBOM: Bringing Trustworthiness to AI-Enabled Systems (Safronov et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02169&sa=D&source=editors&ust=1779048536949324&usg=AOvVaw3WRaekkVLmNobJWfgHW7gY) |
| 466 | [▪ Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks (Palit, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01261&sa=D&source=editors&ust=1779048536949406&usg=AOvVaw3BSV9UUJiluGDAMjk-3mf1) |
| 467 | [▪ A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks (Hossain et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.14285&sa=D&source=editors&ust=1779048536949486&usg=AOvVaw2wAn5JuwkHrAJFryRAvquk) |
| 468 | [▪ A Call to Action for a Secure-by-Design Generative AI Paradigm (Alharthi and Garcia, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00451&sa=D&source=editors&ust=1779048536949565&usg=AOvVaw2SuMhFXFDjElH9cymaFrDS) |
| 469 | [▪ MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00317&sa=D&source=editors&ust=1779048536949645&usg=AOvVaw1-PmXdbeBmMHgdZ4ZSiu5w) |
| 470 | [▪ Detecting Instruction Fine-tuning Attacks on Language Models using Influence Function (Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.09026&sa=D&source=editors&ust=1779048536949725&usg=AOvVaw1oPmYtj--Q8KVHuQH4gQCQ) |
| 471 | [▪ Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.05021&sa=D&source=editors&ust=1779048536949808&usg=AOvVaw2Bf-TPLygJn9Vo5fqMXXWN) |
| 472 | [▪ Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation (Ndayipfukamiye et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.20411&sa=D&source=editors&ust=1779048536949891&usg=AOvVaw06mglOw37JVWK_79lqlfOU) |
| 473 | [▪ QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety (Lee et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12299&sa=D&source=editors&ust=1779048536949970&usg=AOvVaw1XyPMzuAb1Zvqm5c0dn6oZ) |
| 474 | [▪ Federated Learning Resilient to Byzantine Attacks and Data Heterogeneity (Zuo et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.13374&sa=D&source=editors&ust=1779048536950050&usg=AOvVaw19DZjvDsg8ifXinxb0Xp0c) |
| 475 | [▪ SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents (Dong et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23694&sa=D&source=editors&ust=1779048536950129&usg=AOvVaw3NkOwwLnT583aY8okekFMx) |
| 476 | [▪ GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners (Li et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24418&sa=D&source=editors&ust=1779048536950209&usg=AOvVaw1zHuNNWrq38xW7Z_h97pI2) |
| 477 | [▪ Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting (Meng et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23571&sa=D&source=editors&ust=1779048536950295&usg=AOvVaw3KwChnZlY2XbkKR4DH-VFI) |
| 478 | [▪ ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search (Shen et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23519&sa=D&source=editors&ust=1779048536950382&usg=AOvVaw0UqYg-myUW2Psx1Cp60lgQ) |
| 479 | [▪ Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment (Kim et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22745&sa=D&source=editors&ust=1779048536950467&usg=AOvVaw3CAlQ5fipOSIQHXNNGqfTj) |
| 480 | [▪ Think Broad, Act Narrow: CWE Identification with Multi-Agent Large Language Models (Sayagh and Ghafari, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01451&sa=D&source=editors&ust=1779048536950548&usg=AOvVaw0whafXfHBQjhzRVLjgi-SY) |
| 481 | [▪ ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models (Wang et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01365&sa=D&source=editors&ust=1779048536950630&usg=AOvVaw3cEIzIoGgjyeSgHgrNq5Q7) |
| 482 | [▪ Provably Secure Retrieval-Augmented Generation (Zhou, Feng, and Yang, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01084&sa=D&source=editors&ust=1779048536950708&usg=AOvVaw0sVBByPn2z6kHSb4kuW5wt) |
| 483 | [▪ FedGuard: A Diverse-Byzantine-Robust Mechanism for Federated Learning with Major Malicious Clients (Jiang et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00636&sa=D&source=editors&ust=1779048536950790&usg=AOvVaw294w6j-LrENMZR1spYW8Za) |
| 484 | [▪ CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation (Yang et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.13201&sa=D&source=editors&ust=1779048536950872&usg=AOvVaw0zgxmNm3NoHFWsPVqr0-Xs) |
| 485 | [▪ Bridging Privacy and Robustness for Trustworthy Machine Learning (Zhang and Chen, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.16591&sa=D&source=editors&ust=1779048536950952&usg=AOvVaw1lW5IumSPELBU9sn6t8JCk) |
| 486 | [▪ Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems (Sharshar et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22239&sa=D&source=editors&ust=1779048536951035&usg=AOvVaw1CxB1Nu0dYy3wj_-Fku9Sl) |
| 487 | [▪ Strategic Deflection: Defending LLMs from Logit Manipulation (Rachidy et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22160&sa=D&source=editors&ust=1779048536951113&usg=AOvVaw1K74OKV02rU8ZweBuWPmVh) |
| 488 | [▪ Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security (Dai et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22037&sa=D&source=editors&ust=1779048536951196&usg=AOvVaw2qH6eYbRPvAr5I53-2a1pV) |
| 489 | [▪ GUARD-CAN: Graph-Understanding and Recurrent Architecture for CAN Anomaly Detection (Kim and Kim, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21640&sa=D&source=editors&ust=1779048536951276&usg=AOvVaw3ZkpA4TNN8GQgvTqsy-Gke) |
| 490 | [▪ SDD: Self-Degraded Defense against Malicious Fine-tuning (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21182&sa=D&source=editors&ust=1779048536951361&usg=AOvVaw19W2rvTjCIp5JQJO8sfuUs) |
| 491 | [▪ OneShield -- the Next Generation of LLM Guardrails (DeLuca et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21170&sa=D&source=editors&ust=1779048536951441&usg=AOvVaw1pgDBFvv6dviHYz1wIImUT) |
| 492 | [▪ Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards (Madhavan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08610&sa=D&source=editors&ust=1779048536951524&usg=AOvVaw3rooLtwMd-nLyeHkbMQb1A) |
| 493 | [▪ Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs (Camporese and Massacci, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20977&sa=D&source=editors&ust=1779048536951634&usg=AOvVaw18WnGc1By0daD1EisS2m5q) |
| 494 | [▪ PurpCode: Reasoning for Safer Code Generation (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19060&sa=D&source=editors&ust=1779048536951718&usg=AOvVaw3nKLByxcQAQW8wzPt9_4DU) |
| 495 | [▪ SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysis (Lingxiang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17798&sa=D&source=editors&ust=1779048536951802&usg=AOvVaw0-tLIl1w1lol9-Ad080Clu) |
| 496 | [▪ Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18631&sa=D&source=editors&ust=1779048536951884&usg=AOvVaw0aBB33P9vQW35PRN5H2OFZ) |
| 497 | [▪ CASCADE: LLM-Powered JavaScript Deobfuscator at Google (Jiang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17691&sa=D&source=editors&ust=1779048536951986&usg=AOvVaw2qrXtJCHIdKCMypR4NBK4y) |
| 498 | [▪ LLM Meets the Sky: Heuristic Multi-Agent Reinforcement Learning for Secure Heterogeneous UAV Networks (Zheng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17188&sa=D&source=editors&ust=1779048536952071&usg=AOvVaw10uEQvQuyIEThvGSQYMW4N) |
| 499 | [▪ CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems (Hu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.22890&sa=D&source=editors&ust=1779048536952157&usg=AOvVaw2nI_4114JsGtwtHGHUunck) |
| 500 | [▪ DP-TLDM: Differentially Private Tabular Latent Diffusion Model (Zhu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.07842&sa=D&source=editors&ust=1779048536952239&usg=AOvVaw0ZNtavNxNe5c58qEiCeKVr) |
| 501 | [▪ Recent Advances in Malware Detection: Graph Learning and Explainability (Shokouhinejad et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.10556&sa=D&source=editors&ust=1779048536952326&usg=AOvVaw00evqv4t-i6JUGHDHeeN3-) |
| 502 | [▪ "We Need a Standard": Toward an Expert-Informed Privacy Label for Differential Privacy (Dibia et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15997&sa=D&source=editors&ust=1779048536952409&usg=AOvVaw2HsthofzITiIKND63mlL6M) |
| 503 | [▪ ACFIX: Guiding LLMs with Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart Contracts (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.06838&sa=D&source=editors&ust=1779048536952500&usg=AOvVaw1tSncngjJVVpUcDFQzbSFK) |
| 504 | [▪ OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting (Cheng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.03108&sa=D&source=editors&ust=1779048536952585&usg=AOvVaw0rL5i7B-9jFiwA4WpC9DuP) |
| 505 | [▪ Defending Against Unforeseen Failure Modes with Latent Adversarial Training (Casper et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.05030&sa=D&source=editors&ust=1779048536952665&usg=AOvVaw1s8oYwmY0m2bjC-ATJEum7) |
| 506 | [▪ FedStrategist: A Meta-Learning Framework for Adaptive and Robust Aggregation in Federated Learning (Haque, Kamal, and Hossain, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14322&sa=D&source=editors&ust=1779048536952748&usg=AOvVaw3i2HUtuxvzDrp8BvGrWr7u) |
| 507 | [▪ PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15393&sa=D&source=editors&ust=1779048536952832&usg=AOvVaw3X7_viPkdC22JjLIfMEdqQ) |
| 508 | [▪ PromptArmor: Simple yet Effective Prompt Injection Defenses (Shi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15219&sa=D&source=editors&ust=1779048536952913&usg=AOvVaw326SOn0bB0xjVaZ-3Sycn7) |
| 509 | [▪ A Privacy-Centric Approach: Scalable and Secure Federated Learning Enabled by Hybrid Homomorphic Encryption (Nguyen, Khan, and Michalas, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14853&sa=D&source=editors&ust=1779048536952997&usg=AOvVaw2zhyaBsCrdoO3ts6bXc9UQ) |
| 510 | [▪ PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training (Du, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14202&sa=D&source=editors&ust=1779048536953077&usg=AOvVaw0lKmVmEKYLOpMr8b_YiUO1) |
| 511 | [▪ Defense Against Prompt Injection Attack by Leveraging Attack Techniques (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.00459&sa=D&source=editors&ust=1779048536953156&usg=AOvVaw19MVKZzZRHesp9SdICOclc) |
| 512 | [▪ GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention (Abdalla et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13598&sa=D&source=editors&ust=1779048536953237&usg=AOvVaw0ZaP5w2YIrMQyKO3khqTj-) |
| 513 | [▪ Risks of ignoring uncertainty propagation in AI-augmented security pipelines (Mezzi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.14540&sa=D&source=editors&ust=1779048536953330&usg=AOvVaw0uozLoGCWwtchtPUw2pea7) |
| 514 | [▪ JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model (Nian et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.03770&sa=D&source=editors&ust=1779048536953434&usg=AOvVaw1LRIkhd_cvrg28CtKIVHq_) |
| 515 | [▪ SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks (Uddin et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13170&sa=D&source=editors&ust=1779048536953530&usg=AOvVaw18O59HsdW73qc8rXveW4_q) |
| 516 | [▪ Safeguarding Federated Learning-based Road Condition Classification (Liu and Papadimitratos, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12568&sa=D&source=editors&ust=1779048536953626&usg=AOvVaw1oRa2PLzOg8SzG72DyLT9y) |
| 517 | [▪ Thought Purity: Defense Paradigm For Chain-of-Thought Attack (Xue et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12314&sa=D&source=editors&ust=1779048536953721&usg=AOvVaw1kZPmqDxeq-a2QHqvHfZvQ) |
| 518 | [▪ Expanding ML-Documentation Standards For Better Security (Appel, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12003&sa=D&source=editors&ust=1779048536953812&usg=AOvVaw3YpDEjGZ909cnCU3SWZ2tQ) |
| 519 | [▪ A Generative Approach to LLM Harmfulness Detection with Special Red Flag Tokens (Xhonneux et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16366&sa=D&source=editors&ust=1779048536953902&usg=AOvVaw2JI58265j8hBy94rW7xWIY) |
| 520 | [▪ A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation (Trudslev et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11324&sa=D&source=editors&ust=1779048536954038&usg=AOvVaw2Y5jD6eYkSeXR22JbegBtB) |
| 521 | [▪ From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection (Sun et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10873&sa=D&source=editors&ust=1779048536954121&usg=AOvVaw3_06fL5nOgrElSNkMK0u85) |
| 522 | [▪ Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data (Alikhani and Kazemi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10808&sa=D&source=editors&ust=1779048536954206&usg=AOvVaw1xYQX6P3ZGTCLznYOByXjF) |
| 523 | [▪ Defense-as-a-Service: Black-box Shielding against Backdoored Graph Models (Yang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.04916&sa=D&source=editors&ust=1779048536954324&usg=AOvVaw3fuDgXF_nihQi9SU0e4B_X) |
| 524 | [▪ Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators (Xue et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.03278&sa=D&source=editors&ust=1779048536954411&usg=AOvVaw0kPeG5rFJOL3f95zBJ74of) |
| 525 | [▪ SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework (Brito et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10489&sa=D&source=editors&ust=1779048536954493&usg=AOvVaw0Q5ecTa5gnmWB8QpQVxRmu) |
| 526 | [▪ AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.14697&sa=D&source=editors&ust=1779048536954573&usg=AOvVaw0ysK_Gqa1rqJ6CRznq2tkW) |
| 527 | [▪ Entangled Threats: A Unified Kill Chain Model for Quantum Machine Learning Security (Debus et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08623&sa=D&source=editors&ust=1779048536954656&usg=AOvVaw3HMZnfQS6TrmukkKXxDMWy) |
| 528 | [▪ ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware Detection (Alam, Piplai, and Rastogi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08597&sa=D&source=editors&ust=1779048536954738&usg=AOvVaw2GnYHSxtpfYrm3w_E96-zC) |
| 529 | [▪ Beyond the Worst Case: Extending Differential Privacy Guarantees to Realistic Adversaries (Swanberg et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08158&sa=D&source=editors&ust=1779048536954821&usg=AOvVaw3vU867iVaA9HAWK1-BruS6) |
| 530 | [▪ A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning (Gluch and Goldwasser, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.20310&sa=D&source=editors&ust=1779048536954902&usg=AOvVaw2JRsBW1Z4SSk8SYRPbvXei) |
| 531 | [▪ Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences (Lyu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.10427&sa=D&source=editors&ust=1779048536955004&usg=AOvVaw2K4s_ApTKM_UYHpPwNJaZu) |
| 532 | [▪ Adversarial Defenses via Vector Quantization (Dong and Mao, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2305.13651&sa=D&source=editors&ust=1779048536955085&usg=AOvVaw3jnkW_nDZAWyC9WXJuEylK) |
| 533 | [▪ Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking (Wu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07483&sa=D&source=editors&ust=1779048536955170&usg=AOvVaw2px5SU9diQxExVfNx1-CXC) |
| 534 | [▪ Defending Against Prompt Injection With a Few DefensiveTokens (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07974&sa=D&source=editors&ust=1779048536955251&usg=AOvVaw1YLuah4IaiZl0mZaDzrEiH) |
| 535 | [▪ Can Large Language Models Improve Phishing Defense? A Large-Scale Controlled Experiment on Warning Dialogue Explanations (Cau et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07916&sa=D&source=editors&ust=1779048536955341&usg=AOvVaw0J3zxtjC_cClU7FdbUMwIa) |
| 536 | [▪ Hybrid LLM-Enhanced Intrusion Detection for Zero-Day Threats in IoT Networks (Al-Hammouri et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07413&sa=D&source=editors&ust=1779048536955422&usg=AOvVaw26az_WHwi4Jz011Z6R_KjZ) |
| 537 | [▪ Saffron-1: Safety Inference Scaling (Qiu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06444&sa=D&source=editors&ust=1779048536955499&usg=AOvVaw2jk9F-1tVDKvfRtEYT0AWY) |
| 538 | [▪ LDP$^3$: An Extensible and Multi-Threaded Toolkit for Local Differential Privacy Protocols and Post-Processing Methods (Balioglu, Khodaie, and Gursoy, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05872&sa=D&source=editors&ust=1779048536955584&usg=AOvVaw1H_pvGfUG6iTVAmMWAhcS1) |
| 539 | [▪ TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data (Cheruvu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05660&sa=D&source=editors&ust=1779048536955665&usg=AOvVaw2hZl8tEH-bompIl3OKOXTc) |
| 540 | [▪ PROTEAN: Federated Intrusion Detection in Non-IID Environments through Prototype-Based Knowledge Sharing (Chennoufi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05524&sa=D&source=editors&ust=1779048536955747&usg=AOvVaw03W2NauCrWeuSZ0edflgCT) |
| 541 | [▪ MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamin\<br>\<br> (Guo et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17147&sa=D&source=editors&ust=1779048536955830&usg=AOvVaw3VCT3SriFnwnj1PPkA8FGD) |
| 542 | [▪ Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration\<br>\<br> (Tsmindashvili et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17066&sa=D&source=editors&ust=1779048536955911&usg=AOvVaw2aCqbSKF-eP8-b7pjc2uTs) |
| 543 | [▪ LLM-Based Threat Detection and Prevention Framework for IoT Ecosystems\<br>\<br> (Otoum, Asad, and Nayak, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.00240&sa=D&source=editors&ust=1779048536955993&usg=AOvVaw3O3D3Thq0ji-gWCXedWajM) |
| 544 | [▪ Securing RAG: A Risk Assessment and Mitigation Framework\<br>\<br> (Ammann et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.08728&sa=D&source=editors&ust=1779048536956072&usg=AOvVaw2e4xOs4f9asCz1cLXrQTdT) |
| 545 | [▪ Large Language Model Sentinel: LLM Agent for Adversarial Purification (Lin, Tanaka, and Zhao, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20770&sa=D&source=editors&ust=1779048536956153&usg=AOvVaw17nK2r2hRLSprRE4Qa63XM) |
| 546 | [▪ JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks (Zhang et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.10766&sa=D&source=editors&ust=1779048536956231&usg=AOvVaw3v-csf-W0Ttlc_KYEei0P2) |

|     |
| --- |
| Defenses |

**>**

**<**

‍

### Threat to AI Models

#### General Approaches

We added this subsection to cover research that broadly looks at AI security.

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : General Approaches

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks (Iyer et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.15118&sa=D&source=editors&ust=1779048536035018&usg=AOvVaw2SH4J14XYg0ga_Amkvqnts) |
| 3 | [▪ From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World (Conde et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10834&sa=D&source=editors&ust=1779048536035298&usg=AOvVaw3OpeeQX2lnLE5I5Xl1Dchc) |
| 4 | [▪ Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights (Pourhanifeh, AbdulGhaffar, and Matrawy, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10808&sa=D&source=editors&ust=1779048536035494&usg=AOvVaw1cvkZlNmsTFV6tCA-ErAk0) |
| 5 | [▪ CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios (Lim et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07830&sa=D&source=editors&ust=1779048536035647&usg=AOvVaw0qioGpITVnjj3BaJ_4SzyY) |
| 6 | [▪ Evaluating the Reliability of Multiple Large Language Models in Risk Assessment: A CIS Controls Based Approach (Pinto, Labaki, and Miani, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05424&sa=D&source=editors&ust=1779048536035814&usg=AOvVaw2schCXelPNgkJMx6gGvKnB) |
| 7 | [▪ QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing (Ameen, Alam, and Islam, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01885&sa=D&source=editors&ust=1779048536035974&usg=AOvVaw3erDzyoYk8_BeHOVVBAdC2) |
| 8 | [▪ GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems (Mateo-Torrej\\'on and S\\'anchez-Maci\\'an, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24477&sa=D&source=editors&ust=1779048536036146&usg=AOvVaw1a4dcA0EXuwN_cWfyYswuh) |
| 9 | [▪ Training a General Purpose Automated Red Teaming Model (Padmakumar et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23067&sa=D&source=editors&ust=1779048536036296&usg=AOvVaw0qYS5YWq03armdDyoCl9wv) |
| 10 | [▪ CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents (Fu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21308&sa=D&source=editors&ust=1779048536036440&usg=AOvVaw3girf1ckpLMa4jrEjhDkg_) |
| 11 | [▪ Synthesizing Multi-Agent Harnesses for Vulnerability Discovery (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20801&sa=D&source=editors&ust=1779048536036594&usg=AOvVaw2Y9vy4Tuc2AdhhpCCwWP48) |
| 12 | [▪ WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks (Ramesh et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06367&sa=D&source=editors&ust=1779048536036746&usg=AOvVaw3ky5KV0dE5FsXeqJ3lmur_) |
| 13 | [▪ SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement (Duan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04989&sa=D&source=editors&ust=1779048536036897&usg=AOvVaw0cUCfXOzrYseDmjQJP74_e) |
| 14 | [▪ Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities (Mouzouni, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04561&sa=D&source=editors&ust=1779048536037082&usg=AOvVaw0pv3KaW0Gho_P3Y0k4PZ7T) |
| 15 | [▪ AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective (Wang and Luan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24857&sa=D&source=editors&ust=1779048536037242&usg=AOvVaw0gZk0OqqGVuuFjV0pN2e8p) |
| 16 | [▪ Network- and Device-Level Cyber Deception for Contested Environments Using RL and LLMs (Sahu, Paul, and Macwan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17272&sa=D&source=editors&ust=1779048536037396&usg=AOvVaw2VVvsaayVpCA6-Ch3xTvM-) |
| 17 | [▪ ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models (Owiredu-Ashley, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10068&sa=D&source=editors&ust=1779048536037555&usg=AOvVaw2lkKP-k8QYD8wQUltY03Gr) |
| 18 | [▪ CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research? (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09452&sa=D&source=editors&ust=1779048536037725&usg=AOvVaw3pcHgChqavDE6lMM6q5Twr) |
| 19 | [▪ SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection (Hasan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05689&sa=D&source=editors&ust=1779048536037888&usg=AOvVaw2Y4UDG9-ArqtDBzgIEtvhV) |
| 20 | [▪ From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architectures (Bonfanti et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03911&sa=D&source=editors&ust=1779048536038051&usg=AOvVaw1e5yxgPUQZY2bO3eEpdX-R) |
| 21 | [▪ What Makes a Good LLM Agent for Real-world Penetration Testing? (Deng et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.17622&sa=D&source=editors&ust=1779048536038198&usg=AOvVaw2buBUnhjnWcsRNDh5nIbqm) |
| 22 | [▪ Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP (Anbiaee et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11327&sa=D&source=editors&ust=1779048536038358&usg=AOvVaw3gayCl_k5zgxJlZoBd48L5) |
| 23 | [▪ LLMs + Security = Trouble (Livshits, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08422&sa=D&source=editors&ust=1779048536038499&usg=AOvVaw2W_Y6Bn5Kj3oRveW2KHiZe) |
| 24 | [▪ Efficient LLM Moderation with Multi-Layer Latent Prototypes (Chrab\\k{a}szcz et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16174&sa=D&source=editors&ust=1779048536038653&usg=AOvVaw1VuAO7O4E39pjsVrbV9Qpl) |
| 25 | [▪ TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06443&sa=D&source=editors&ust=1779048536038798&usg=AOvVaw0aSexJxGqqwwFV768B5yfU) |
| 26 | [▪ From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2407.16928&sa=D&source=editors&ust=1779048536038943&usg=AOvVaw20tXVMK_88zrGj5HyV0fOn) |
| 27 | [▪ Risk Assessment and Security Analysis of Large Language Models (Zhang, Lyu, and Li, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.17329&sa=D&source=editors&ust=1779048536039092&usg=AOvVaw35ictIhPMyoOMKifnU451g) |
| 28 | [▪ LogicScan: An LLM-driven Framework for Detecting Business Logic Vulnerabilities in Smart Contracts (Gao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03271&sa=D&source=editors&ust=1779048536039244&usg=AOvVaw3XKL12IimYrY4pQVVNj-TV) |
| 29 | [▪ RedSage: A Cybersecurity Generalist LLM (Suryanto et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22159&sa=D&source=editors&ust=1779048536039411&usg=AOvVaw0pmBGMdB2kZKCKxBYuCL-B) |
| 30 | [▪ Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks (Jackson et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17225&sa=D&source=editors&ust=1779048536039561&usg=AOvVaw0Pn2i0-EulRrY1T16kj8KX) |
| 31 | [▪ LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers (Okada, Oba, and Yanai, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03013&sa=D&source=editors&ust=1779048536039692&usg=AOvVaw2ccUwLhN_OrN38fJFvIuaM) |
| 32 | [▪ Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework (Peng et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07122&sa=D&source=editors&ust=1779048536039825&usg=AOvVaw0cKm1GKOH92PDpWpHGSpgi) |
| 33 | [▪ CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation (Iyer, Bobadilla, and Iyengar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06779&sa=D&source=editors&ust=1779048536039942&usg=AOvVaw0_TwNHnZtdKPrjGa9a29pW) |
| 34 | [▪ Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense (Mayoral-Vilches et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05887&sa=D&source=editors&ust=1779048536040080&usg=AOvVaw3iXIjuL4aCOoZYetqWqSp0) |
| 35 | [▪ Automated Red-Teaming Framework for Large Language Model Security Assessment: A Comprehensive Attack Generation and Detection System (Wei et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.20677&sa=D&source=editors&ust=1779048536040191&usg=AOvVaw0dD0A9DPVS1XDOBqpOAjoD) |
| 36 | [▪ Bounty Hunter: Autonomous, Comprehensive Emulation of Multi-Faceted Adversaries (Hackl\\"ander-Jansen, Uetz, and Henze, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15275&sa=D&source=editors&ust=1779048536040291&usg=AOvVaw1RVzZuy6bVXhi9fiB3lwxS) |
| 37 | [▪ PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14233&sa=D&source=editors&ust=1779048536040392&usg=AOvVaw3SHUc2tYeB7Ga7cKCpkiCe) |
| 38 | [▪ Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing (Lin et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09882&sa=D&source=editors&ust=1779048536040505&usg=AOvVaw0h2mVugk3dBuIXb1kwerZf) |
| 39 | [▪ SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security (Zhao, Li, and Sun, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04841&sa=D&source=editors&ust=1779048536040607&usg=AOvVaw0lIlh1C7weCYgkzAFrI9me) |
| 40 | [▪ Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks (Singer et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16466&sa=D&source=editors&ust=1779048536040698&usg=AOvVaw3ORF1I5GthB2ZELn6zO4_w) |
| 41 | [▪ SoK: The Security-Safety Continuum of Multimodal Foundation Models through Information Flow and Global Game-Theoretic Analysis of Asymmetric Threats (Sun et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.11195&sa=D&source=editors&ust=1779048536040790&usg=AOvVaw3kJDCVmSPBRd33vABmr7Np) |
| 42 | [▪ Large Language Models for Cyber Security (Somani and Cherukuri, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04508&sa=D&source=editors&ust=1779048536040870&usg=AOvVaw3qcdexNlAley2KG8Tg3CS9) |
| 43 | [▪ The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies (He et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.19354&sa=D&source=editors&ust=1779048536040957&usg=AOvVaw0SOt8vdltP8Zo1O7Mla5LQ) |
| 44 | [▪ Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation (Borah, Alam, and Rastogi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27080&sa=D&source=editors&ust=1779048536041050&usg=AOvVaw3TeWzwrUMoTgUZwOW9WDiq) |
| 45 | [▪ Ask What Your Country Can Do For You: Towards a Public Red Teaming Model (Kennedy et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20061&sa=D&source=editors&ust=1779048536041137&usg=AOvVaw2Shegr5IZrBfmhM8qpDC9i) |
| 46 | [▪ SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models (Hong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15476&sa=D&source=editors&ust=1779048536041221&usg=AOvVaw3Pj1Wpb4q9K5jhulwS1GPM) |
| 47 | [▪ BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing (Kaplan, Warnecke, and Archibald, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11823&sa=D&source=editors&ust=1779048536041330&usg=AOvVaw3mEAAN0VV5fwbVTCsn3I3K) |
| 48 | [▪ Frontier AI's Impact on the Cybersecurity Landscape (Potter et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.05408&sa=D&source=editors&ust=1779048536041434&usg=AOvVaw1MYYMLx-NEBmkeCsHX53on) |
| 49 | [▪ PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11688&sa=D&source=editors&ust=1779048536041531&usg=AOvVaw3T-VDs9h5UGDq3nLF-Fflz) |
| 50 | [▪ Living Off the LLM: How LLMs Will Change Adversary Tactics (Oesch et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11398&sa=D&source=editors&ust=1779048536041618&usg=AOvVaw2ALbZasL0YLhTelHGJiaSX) |
| 51 | [▪ CREST-Search: Comprehensive Red-teaming for Evaluating Safety Threats in Large Language Models Powered by Web Search (Ou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09689&sa=D&source=editors&ust=1779048536041704&usg=AOvVaw1s-gK7CYWCrJusbxGEYTBk) |
| 52 | [▪ MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols (Yang, Wu, and Chen, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.13220&sa=D&source=editors&ust=1779048536041795&usg=AOvVaw1ZLnltOkVsN0Mm401fe6Gv) |
| 53 | [▪ Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions (Hou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.23278&sa=D&source=editors&ust=1779048536041877&usg=AOvVaw1YmDXgsuxx59wfCdXHUKtd) |
| 54 | [▪ Quantifying Risks in Multi-turn Conversation with Large Language Models (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03969&sa=D&source=editors&ust=1779048536041965&usg=AOvVaw141yvQTKjOU7al24eTldLJ) |
| 55 | [▪ LLAMAFUZZ: Large Language Model Enhanced Greybox Fuzzing (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.07714&sa=D&source=editors&ust=1779048536042086&usg=AOvVaw1D_PMNGAGR6w6_2V3_7CNC) |
| 56 | [▪ MALF: A Multi-Agent LLM Framework for Intelligent Fuzzing of Industrial Control Protocols (Ning, Zong, and He, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02694&sa=D&source=editors&ust=1779048536042193&usg=AOvVaw0f-PHwoB9JS-0-6N1unN11) |
| 57 | [▪ Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B (Lin et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23882&sa=D&source=editors&ust=1779048536042294&usg=AOvVaw2SHXITDBU8EYV9Bb8A3I7Y) |
| 58 | [▪ Generative AI-Empowered Secure Communications in Space-Air-Ground Integrated Networks: A Survey and Tutorial (Hu et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01983&sa=D&source=editors&ust=1779048536042378&usg=AOvVaw0qPVQMXQcAoDqt86MoiqLI) |
| 59 | [▪ AI-Driven Cybersecurity Threat Detection: Building Resilient Defense Systems Using Predictive Analytics (Das et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01422&sa=D&source=editors&ust=1779048536042462&usg=AOvVaw1bck4N7aGWXsGxoeqXkHG9) |
| 60 | [▪ BlockA2A: Towards Secure and Verifiable Agent-to-Agent Interoperability (Zou et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01332&sa=D&source=editors&ust=1779048536042562&usg=AOvVaw1zQYL0GEB3h6JeAznHIwR7) |
| 61 | [▪ From Cloud-Native to Trust-Native: A Protocol for Verifiable Multi-Agent Systems (Li, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22077&sa=D&source=editors&ust=1779048536042679&usg=AOvVaw3rKdfjp-WEfnulFBZeXcg1) |
| 62 | [▪ Prompt Optimization and Evaluation for LLM Automated Red Teaming (Freenor et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22133&sa=D&source=editors&ust=1779048536042832&usg=AOvVaw3N_m4NwnuTDtwcgB0ushsM) |
| 63 | [▪ Interpretable Anomaly-Based DDoS Detection in AI-RAN with XAI and LLMs (Chatzimiltis et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21193&sa=D&source=editors&ust=1779048536042948&usg=AOvVaw19m4H3dOpS_RcQgiSvpcV0) |
| 64 | [▪ Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems (Sharma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21146&sa=D&source=editors&ust=1779048536043055&usg=AOvVaw2ys7tBvaxMIacwfTaDo8Ip) |
| 65 | [▪ Leveraging Trustworthy AI for Automotive Security in Multi-Domain Operations: Towards a Responsive Human-AI Multi-Domain Task Force for Cyber Social Security (Barletta et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21145&sa=D&source=editors&ust=1779048536043156&usg=AOvVaw2OVrOwQyMhs04cYFv1dwod) |
| 66 | [▪ Security practices in AI development (Spelda and Stritecky, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21061&sa=D&source=editors&ust=1779048536043237&usg=AOvVaw2z_RQYd4B2Lvk5q8ubYdDP) |
| 67 | [▪ PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI (Lin et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14433&sa=D&source=editors&ust=1779048536043331&usg=AOvVaw3-BfHSDJ4sh2vKcKHBmpMZ) |
| 68 | [▪ Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint (Ahi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12088&sa=D&source=editors&ust=1779048536043462&usg=AOvVaw3l5odqn71_MOOxI8_RrSjC) |
| 69 | [▪ A Large Language Model-Supported Threat Modeling Framework for Transportation Cyber-Physical Systems (Salek et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00831&sa=D&source=editors&ust=1779048536043560&usg=AOvVaw2vhVHuekjUPEPVfEspcGNn) |
| 70 | [▪ Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation (Chandra and Navneet, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20014&sa=D&source=editors&ust=1779048536043650&usg=AOvVaw1i1mO3rm0q3r3dRTwCQcaZ) |
| 71 | [▪ Regression-aware Continual Learning for Android Malware Detection (Ghiani et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18313&sa=D&source=editors&ust=1779048536043737&usg=AOvVaw14shUO2-gYLpeTrPppPxHC) |
| 72 | [▪ Your ATs to Ts: MITRE ATT&CK Attack Technique to P-SSCRM Task Mapping (Hamer et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18037&sa=D&source=editors&ust=1779048536043819&usg=AOvVaw2HsRd-TfRs-nfACCWlTSEV) |
| 73 | [▪ Information Security Based on LLM Approaches: A Review (Gong, Li, and Li, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18215&sa=D&source=editors&ust=1779048536043915&usg=AOvVaw2BoJCBn_UB04tJxWKlKMbB) |
| 74 | [▪ Optimizing Privacy-Utility Trade-off in Decentralized Learning with Generalized Correlated Noise (Rodio, Chen, and Larsson, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.14644&sa=D&source=editors&ust=1779048536044012&usg=AOvVaw3Ay7aGP8aWVj60FyWc60RY) |
| 75 | [▪ Enabling Cyber Security Education through Digital Twins and Generative AI (Barletta et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17518&sa=D&source=editors&ust=1779048536044115&usg=AOvVaw2w9DGWR8CbuwWhE0tyyiKe) |
| 76 | [▪ Threshold-Protected Searchable Sharing: Privacy Preserving Aggregated-ANN Search for Collaborative RAG (Guo, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17199&sa=D&source=editors&ust=1779048536044224&usg=AOvVaw1_vsuacQ-Lj-qNDOFBOPkw) |
| 77 | [▪ Large Language Models are Autonomous Cyber Defenders (Castro et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.04843&sa=D&source=editors&ust=1779048536044323&usg=AOvVaw1S8fTZgM8g0jce8OZTCu0e) |
| 78 | [▪ Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis (Tsegai et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2211.05403&sa=D&source=editors&ust=1779048536044421&usg=AOvVaw0qk197lcA5QizHjY0D48fw) |
| 79 | [▪ Adaptive Network Security Policies via Belief Aggregation and Rollout (Hammar et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15163&sa=D&source=editors&ust=1779048536044513&usg=AOvVaw08EgQB-Skk3ghECUrTE_o1) |
| 80 | [▪ Learning to Communicate in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence (Contractor, Li, and Mallah, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14658&sa=D&source=editors&ust=1779048536044615&usg=AOvVaw2nE6izOuGIFO_2ga9ufG1b) |
| 81 | [▪ ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation (Wu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14201&sa=D&source=editors&ust=1779048536044719&usg=AOvVaw2O_HtGt0V4djAhywAcPamM) |
| 82 | [▪ Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques (Jaffal, Alkhanafseh, and Mohaisen, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13629&sa=D&source=editors&ust=1779048536044804&usg=AOvVaw0MIloIJAqjyYBY8iwlKvPW) |
| 83 | [▪ Challenges in GenAI and Authentication: a scoping review (Bezerra, Bezerra, and Westphall, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11775&sa=D&source=editors&ust=1779048536044891&usg=AOvVaw3rmBEZfqz_EnZHdmtQdJLC) |
| 84 | [▪ Towards Effective Complementary Security Analysis using Large Language Models (Wagner et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.16899&sa=D&source=editors&ust=1779048536044970&usg=AOvVaw1e3WrwRkrqwB6Ceq07H9le) |
| 85 | [▪ PREAMBLE: Private and Efficient Aggregation via Block Sparse Vectors (Asi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.11897&sa=D&source=editors&ust=1779048536045068&usg=AOvVaw0leVsQmvruh5OBYJkcnLGv) |
| 86 | [▪ Domain Borders Are There to Be Crossed With Federated Few-Shot Adaptation (R\\"oder, Raab, and Schleif, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10160&sa=D&source=editors&ust=1779048536045161&usg=AOvVaw03_oY7JPDazrAZsVYpHuZj) |
| 87 | [▪ ARBoids: Adaptive Residual Reinforcement Learning With Boids Model for Cooperative Multi-USV Target Defense (Tao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.18549&sa=D&source=editors&ust=1779048536045269&usg=AOvVaw1rF7JMcS26Nw36i8Yjckat) |
| 88 | [▪ Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs) (Verma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.14937&sa=D&source=editors&ust=1779048536045360&usg=AOvVaw2NTxc19xMTrm2LxiadC_Oa) |
| 89 | [▪ BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15216&sa=D&source=editors&ust=1779048536045449&usg=AOvVaw0rVcCrg4clGSGsLHbjLayp) |
| 90 | [▪ Autonomous AI-based Cybersecurity Framework for Critical Infrastructure: Real-Time Threat Mitigation (Paulraj et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07416&sa=D&source=editors&ust=1779048536045551&usg=AOvVaw10vOu6E1cdAA6XjTQh6ftZ) |
| 91 | [▪ Automated Attack Testflow Extraction from Cyber Threat Report using BERT for Contextual Analysis (Ahmadou et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07244&sa=D&source=editors&ust=1779048536045652&usg=AOvVaw2iLaGdXyCcAHzKTna3dMB-) |
| 92 | [▪ Automated Reasoning for Vulnerability Management by Design (Shaked and Messe, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05794&sa=D&source=editors&ust=1779048536045742&usg=AOvVaw1lSKxFi4hv3Tn-lac2Q0fr) |
| 93 | [▪ LLMs on support of privacy and security of mobile apps: state of the art and research directions (Nguyen, Carminati, and Ferrari, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.11679&sa=D&source=editors&ust=1779048536045828&usg=AOvVaw3BJl13ZELxKHRqXNGzI-xw) |
| 94 | [▪ Red Teaming AI Red Teaming (Majumdar, Pendleton, and Gupta, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05538&sa=D&source=editors&ust=1779048536045905&usg=AOvVaw19uQHRxT5_2sf62SBOxLyW) |
| 95 | [▪ Taming Data Challenges in ML-based Security Tasks: Lessons from Integrating Generative AI (Kanchi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06092&sa=D&source=editors&ust=1779048536045985&usg=AOvVaw0FmIoOoN9XmBaFPv3FfvPg) |
| 96 | [▪ Phare: A Safety Probe for Large Language Models (Le Jeune et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11365&sa=D&source=editors&ust=1779048536046097&usg=AOvVaw3tnDkwLcl5klBzZmlaboBz) |
| 97 | [▪ aiXamine: Simplified LLM Safety and Security (Deniz et al, Jun 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.14985&sa=D&source=editors&ust=1779048536046192&usg=AOvVaw0u57LrspniFMRxbTpv1HzL) |
| 98 | [▪ Safety at Scale: A Comprehensive Survey of Large Model Safety (Ma et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05206&sa=D&source=editors&ust=1779048536046274&usg=AOvVaw2z3e3vMAT7C0L0zxb-o_Jl) |
| 99 | [▪ Emerging Security Challenges of Large Language Models (Debar et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.17614&sa=D&source=editors&ust=1779048536046350&usg=AOvVaw0jM1WJmKdQh3Y0gTFMaOQc) |
| 100 | [▪ Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs (Pasini et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.18216&sa=D&source=editors&ust=1779048536046430&usg=AOvVaw2YRGd94w_fxaTvd3jVMsTp) |
| 101 | [▪ Blockchain for Large Language Model Security and Safety: A Holistic Survey (Geren et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20181&sa=D&source=editors&ust=1779048536046512&usg=AOvVaw0guvzV5f78Tl17K3tz8clh) |
| 102 | [▪ One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks (Guo et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22725&sa=D&source=editors&ust=1779048536046594&usg=AOvVaw0pBB2ZH_Zt7Xsz2AVYgRWX) |
| 103 | [▪ How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries (Banerjee et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.15302&sa=D&source=editors&ust=1779048536046677&usg=AOvVaw1lWRtKRE6S2cTs6vEM4I-o) |
| 104 | [▪ Defending Large Language Models Against Attacks With Residual Stream Activation Analysis (Kawasaki, Davis, and Abbas, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.03230&sa=D&source=editors&ust=1779048536046762&usg=AOvVaw3iVEOwtLcYb0YUR6feoYAG) |
| 105 | [▪ LProtector: An LLM-driven Vulnerability Detection System (Sheng et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.06493&sa=D&source=editors&ust=1779048536046844&usg=AOvVaw0TX2N7u0EPoXsCXmTLtMO6) |
| 106 | [▪ SECURE: Benchmarking Large Language Models for Cybersecurity (Bhusal et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20441&sa=D&source=editors&ust=1779048536046938&usg=AOvVaw3IAEf3ykXlI6sd7xhOz2gs) |
| 107 | [▪ Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis (Brokman et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.16527&sa=D&source=editors&ust=1779048536047033&usg=AOvVaw3u5GGOiaXmjLTA5UPjN5m6) |
| 108 | [▪ Safety Layers in Aligned Large Language Models: The Key to LLM Security (Li et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.17003&sa=D&source=editors&ust=1779048536047113&usg=AOvVaw3Dg-Dvurp5yJ8OSSRGIZ2t) |
| 109 | [▪ AttackBench: Evaluating Gradient-based Attacks for Adversarial Examples (Cinà et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.19460&sa=D&source=editors&ust=1779048536047200&usg=AOvVaw1Er_j_U8eA9_s3iDlq-tci) |
| 110 | [▪ Advancing Cyber Incident Timeline Analysis Through Rule Based AI and Large Language Models (Loumachi and Ghanem, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.02572&sa=D&source=editors&ust=1779048536047287&usg=AOvVaw3yU8cpqhHQeGr7dKviYavv) |
| 111 | [▪ Real-world Adversarial Defense against Patch Attacks based on Diffusion Model (Wei et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.09406&sa=D&source=editors&ust=1779048536047367&usg=AOvVaw3XuhXrvxXgpiHjWX9Qc2p-) |
| 112 | [▪ LLM Honeypot: Leveraging Large Language Models as Advanced Interactive Honeypot Systems (Otal and Canbaz, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.08234&sa=D&source=editors&ust=1779048536047447&usg=AOvVaw1oblOxL4smuL45XGnEIZL9) |
| 113 | [▪ SECURE: Benchmarking Large Language Models for Cybersecurity Advisory (Bhusal et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.20441&sa=D&source=editors&ust=1779048536047532&usg=AOvVaw3cf0dm8uoDICKbauF3QOLO) |
| 114 | [▪ Continual Adversarial Defense (Wang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.09481&sa=D&source=editors&ust=1779048536047623&usg=AOvVaw0XZK7ibYFoD3Isg4JqjqLE) |
| 115 | [▪ Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection (Song et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08902%23&sa=D&source=editors&ust=1779048536047789&usg=AOvVaw3ZMTvCTEEZy2b3j45dYFLI) |
| 116 | [▪ Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey (Truong, Dan, and Le, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.03400%23&sa=D&source=editors&ust=1779048536047882&usg=AOvVaw0e_KLoFehPSzAda2QjSvYO) |
| 117 | [▪ Blockchain for Large Language Model Security and Safety: A Holistic Survey (Geren et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20181&sa=D&source=editors&ust=1779048536047963&usg=AOvVaw01WpjRh7-eCX-dn4fsEfB8) |

|     |
| --- |
| General Approaches |

**>**

**<**

‍

#### Data Poisoning and Simulated Publication of Poisoned Public Datasets

Covers:

- OWASP LLM 03: Training Data Poisoning
- OWASP ML 02: Data Poisoning Attack
- MITRE ATLAS Resource Development

We moved this subsection from 'Threats Using AI Models' to this section as poisoned data is a threat _to_ AI.

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Data Poisoning and Simulated Publication of Poisoned Public Datasets

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ DSTAN-Med: Dual-Channel Spatiotemporal Attention with Physiological Plausibility Filtering for False Data Injection Attack Detection in IoT-Based Medical Devices (Hasan, Islam, and Hossain, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14165&sa=D&source=editors&ust=1779048536031332&usg=AOvVaw2mAoge78BUERerCZxT4pNi) |
| 3 | [▪ SoK: Unlearnability and Unlearning for Model Dememorization (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11592&sa=D&source=editors&ust=1779048536031549&usg=AOvVaw3fYuO0FObPfYdsCh3auQtg) |
| 4 | [▪ Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing (Mao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10146&sa=D&source=editors&ust=1779048536031821&usg=AOvVaw1bQJ7zB4Cd5vlKW4sgFtVD) |
| 5 | [▪ Generate "Normal", Edit Poisoned: Branding Injection via Hint Embedding in Image Editing (Sun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10600&sa=D&source=editors&ust=1779048536031991&usg=AOvVaw2d6u19fYskJHu0s_lewqA3) |
| 6 | [▪ Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation (Yang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10253&sa=D&source=editors&ust=1779048536032122&usg=AOvVaw3QbRW40g4jRCBR-XQq4f9t) |
| 7 | [▪ Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning (Kereopa-Yorke et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09822&sa=D&source=editors&ust=1779048536032233&usg=AOvVaw0B4OxunMXa0STzEVV4UQgF) |
| 8 | [▪ ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts (Luo et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09033&sa=D&source=editors&ust=1779048536032348&usg=AOvVaw2eMzN5gpkHuLlJwZy4zJe8) |
| 9 | [▪ When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06731&sa=D&source=editors&ust=1779048536032469&usg=AOvVaw3i4e3VMRBIm1oGBopQ7rOK) |
| 10 | [▪ Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection (Zeng et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.21938&sa=D&source=editors&ust=1779048536032615&usg=AOvVaw3QZh40k04APgsNptHJGOmT) |
| 11 | [▪ Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05224&sa=D&source=editors&ust=1779048536032740&usg=AOvVaw3RDPZ_HkBeXZWdwNWgqTu_) |
| 12 | [▪ LoopTrap: Termination Poisoning Attacks on LLM Agents (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05846&sa=D&source=editors&ust=1779048536032845&usg=AOvVaw2jAikjoAUfrq5McIernfh6) |
| 13 | [▪ Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning (Korn, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05632&sa=D&source=editors&ust=1779048536032974&usg=AOvVaw3Nww87RMiWDtTG2H5X1nCM) |
| 14 | [▪ Gray-Box Poisoning of Continuous Malware Ingestion Pipelines (Dolej\\v{s}, Jure\\v{c}ek, and L\\'orencz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04698&sa=D&source=editors&ust=1779048536033113&usg=AOvVaw1I1Xn9FZHPON1VW8SJgNZE) |
| 15 | [▪ MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents (California and Berkeley), May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03482&sa=D&source=editors&ust=1779048536033231&usg=AOvVaw22fXA3ilw9HTs_1vIlGcPx) |
| 16 | [▪ Adversarial Update-Based Federated Unlearning for Poisoned Model Recovery (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02110&sa=D&source=editors&ust=1779048536033333&usg=AOvVaw3D2w9u280xx0LUW6kPVVEU) |
| 17 | [▪ Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training (Duan, Zhou, and Li, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02374&sa=D&source=editors&ust=1779048536033455&usg=AOvVaw2PNZ0EY25lpj80x9tZ-4n0) |
| 18 | [▪ Repurposing and Evaluating the (In)Feasibility of Dataset Poisoning enabled Watermarking for Contrastive Learning (Dai et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01834&sa=D&source=editors&ust=1779048536033561&usg=AOvVaw1jrjy29pPoaVMigJRpT1Sb) |
| 19 | [▪ Needle-in-RAG: Prompt-Conditioned Character-Level Traceback of Poisoned Spans in Retrieved Evidence (Cui and Liu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01782&sa=D&source=editors&ust=1779048536033670&usg=AOvVaw1iGp0giHEfHWWDbKiUvRgO) |
| 20 | [▪ From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception (Zhang, Zhang, and Mao, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01301&sa=D&source=editors&ust=1779048536033774&usg=AOvVaw2mtYdAZYWYL-jyvrzeENSs) |
| 21 | [▪ Repurposing Image Diffusion Models for Adversarial Synthetic Structured Data: A Case Study of Ground Truth Drift (Arthur and Schwartz, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00788&sa=D&source=editors&ust=1779048536033887&usg=AOvVaw0r8iVJuhxymSEjDt6b7Qjv) |
| 22 | [▪ Defense against Poisoning Attacks under Shuffle-DP (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00625&sa=D&source=editors&ust=1779048536034027&usg=AOvVaw0DmaOmTbfcOQbU1SOfWH1D) |
| 23 | [▪ CleanBase: Detecting Malicious Documents in RAG Knowledge Databases (Jin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00460&sa=D&source=editors&ust=1779048536034159&usg=AOvVaw3tE7f9-LeuCBfsXbSwgXYg) |
| 24 | [▪ SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation (Rezakhani et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27238&sa=D&source=editors&ust=1779048536034285&usg=AOvVaw3D_mdtlFnVU3LD_-b3Wrmp) |
| 25 | [▪ Poisoning Learned Index Structures: Static and Dynamic Adversarial Attacks on ALEX (Jue, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24975&sa=D&source=editors&ust=1779048536034396&usg=AOvVaw24sXlvXGkPmiWs5_nkpHTE) |
| 26 | [▪ RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents (Xiao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22888&sa=D&source=editors&ust=1779048536034562&usg=AOvVaw3WG5BK2UKfwih2PakDmd2c) |
| 27 | [▪ Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets (Xiao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22291&sa=D&source=editors&ust=1779048536034700&usg=AOvVaw3poxAIiOpzUf1BafRmiXqf) |
| 28 | [▪ CSC: Turning the Adversary's Poison against Itself (Shi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21416&sa=D&source=editors&ust=1779048536034817&usg=AOvVaw1aNX9abSVOBSqcch6cIuvg) |
| 29 | [▪ From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04278&sa=D&source=editors&ust=1779048536034936&usg=AOvVaw2-CPz7lR8etdfwpIy0U3Ir) |
| 30 | [▪ Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories (Bercovich et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17596&sa=D&source=editors&ust=1779048536035099&usg=AOvVaw3zDwdTFuFXJGUdVe1Zu6uU) |
| 31 | [▪ Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning (Qian, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16966&sa=D&source=editors&ust=1779048536035246&usg=AOvVaw2HvB6y6B643M-TnFr6sitO) |
| 32 | [▪ Privacy-Aware Machine Unlearning with SISA for Reinforcement Learning-Based Ransomware Detection (Ferdous, Islam, and Islam, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16760&sa=D&source=editors&ust=1779048536035400&usg=AOvVaw3ZNJswEYOkPoHSgR0ZfJSi) |
| 33 | [▪ FedIDM: Achieving Fast and Stable Convergence in Byzantine Federated Learning through Iterative Distribution Matching (Yang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15115&sa=D&source=editors&ust=1779048536035571&usg=AOvVaw3uk9YO0CkYuDQc0ksZttI8) |
| 34 | [▪ Robustness Analysis of Machine Learning Models for IoT Intrusion Detection Under Data Poisoning Attacks (Wulnye et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.14444&sa=D&source=editors&ust=1779048536035725&usg=AOvVaw15p2wI-J-IGKpGw1o-8EeT) |
| 35 | [▪ PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction (Wadekar et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13153&sa=D&source=editors&ust=1779048536035880&usg=AOvVaw3psnvaIcoEqkTFt7pr5pPc) |
| 36 | [▪ Poisoning with A Pill: Circumventing Detection in Federated Learning (Guo et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2407.15389&sa=D&source=editors&ust=1779048536036018&usg=AOvVaw2ihN-KciRuUllUIWP50HSv) |
| 37 | [▪ DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10611&sa=D&source=editors&ust=1779048536036143&usg=AOvVaw0oUC1Ah2p0zqK6ZUjB8uov) |
| 38 | [▪ XFED: Non-Collusive Model Poisoning Attack Against Byzantine-Robust Federated Classifiers (Mouri, Ridowan, and Adnan, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09489&sa=D&source=editors&ust=1779048536036259&usg=AOvVaw1_XEXU2nC_3RypQBpSBK_g) |
| 39 | [▪ One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems (Chang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11548&sa=D&source=editors&ust=1779048536036367&usg=AOvVaw2YwN7n0GW-U6yzVL_YGyJ7) |
| 40 | [▪ TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation (Ye et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07536&sa=D&source=editors&ust=1779048536036480&usg=AOvVaw2Vv-mZbWQHfAy1h_YDQpC9) |
| 41 | [▪ RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement (Wang, Wang, and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07403&sa=D&source=editors&ust=1779048536036594&usg=AOvVaw2qbbKrXqHC80lbtkFMqXPY) |
| 42 | [▪ FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization (Zhu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06833&sa=D&source=editors&ust=1779048536036775&usg=AOvVaw2izuwtOMxtARthtCyj-tKP) |
| 43 | [▪ Can You Trust the Vectors in Your Vector Database? Black-Hole Attack from Embedding Space Defects (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05480&sa=D&source=editors&ust=1779048536036887&usg=AOvVaw0k5C54DPiBI-wC-S7-o4_Y) |
| 44 | [▪ Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6 (Lorenzo, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04289&sa=D&source=editors&ust=1779048536036988&usg=AOvVaw2SkBet8nGU9JH9nkRmqXc5) |
| 45 | [▪ Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents (Zou et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02623&sa=D&source=editors&ust=1779048536037083&usg=AOvVaw1wKv8T3PY1kCtmrQRWPxCt) |
| 46 | [▪ Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation (Zarkadis and Douligeris, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17717&sa=D&source=editors&ust=1779048536037185&usg=AOvVaw1nWjIP7S0BdT9wKJwxJN4v) |
| 47 | [▪ S-DAPT-2026: A Stage-Aware Synthetic Dataset for Advanced Persistent Threat Detection (Tijjani et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06690&sa=D&source=editors&ust=1779048536037283&usg=AOvVaw3IyuhtD3v20Mc-NAkpaIoa) |
| 48 | [▪ Towards Explainable Privacy Preservation in Federated Learning via Shapley Value-Guided Noise Injection (Li, Gui, and Wu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.12958&sa=D&source=editors&ust=1779048536037377&usg=AOvVaw1u-R3vlVWHDsH_2c2X3nTc) |
| 49 | [▪ RAGShield: Provenance-Verified Defense-in-Depth Against Knowledge Base Poisoning in Government Retrieval-Augmented Generation Systems (Patil, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00387&sa=D&source=editors&ust=1779048536037482&usg=AOvVaw1245qETbx3XtMlpX6vpoGX) |
| 50 | [▪ Unveiling the Security Risks of Federated Learning in the Wild: From Research to Practice (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20615&sa=D&source=editors&ust=1779048536037575&usg=AOvVaw1U18iE7saZ7d4X6kwuJ0rR) |
| 51 | [▪ Memory poisoning and secure multi-agent systems (Torra and Bras-Amor\\'os, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20357&sa=D&source=editors&ust=1779048536037667&usg=AOvVaw34gk75WmXAvv-0ayp4_ddz) |
| 52 | [▪ A Model Consistency-Based Countermeasure to GAN-Based Data Poisoning Attack in Federated Learning (Sun et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2405.11440&sa=D&source=editors&ust=1779048536037762&usg=AOvVaw3cbmmWu4GiNk2oCCn44O0Q) |
| 53 | [▪ FedTrident: Resilient Road Condition Classification Against Poisoning Attacks in Federated Learning (Liu and Papadimitratos, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19101&sa=D&source=editors&ust=1779048536037872&usg=AOvVaw2RWN1ynif_KwnpoYSoD8YA) |
| 54 | [▪ Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems (Thornton, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18034&sa=D&source=editors&ust=1779048536037969&usg=AOvVaw0reec6Mu6scZ4ELHSWwwXE) |
| 55 | [▪ Towards Unsupervised Adversarial Document Detection in Retrieval Augmented Generation Systems (Levi, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17176&sa=D&source=editors&ust=1779048536038071&usg=AOvVaw2SM7HefsLIsl1TMqxTsUXa) |
| 56 | [▪ Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning (Yan et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17174&sa=D&source=editors&ust=1779048536038170&usg=AOvVaw0utX-57GXjlBaLMPd0bJzY) |
| 57 | [▪ KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11501&sa=D&source=editors&ust=1779048536038263&usg=AOvVaw31Y7At7nCl7VQUcWoonW4M) |
| 58 | [▪ Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems (Habler et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22427&sa=D&source=editors&ust=1779048536038355&usg=AOvVaw1zG41wFq3Gjvt3tO15KWJe) |
| 59 | [▪ SPOILER: TEE-Shielded DNN Partitioning of On-Device Secure Inference with Poison Learning (Kang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.06263&sa=D&source=editors&ust=1779048536038452&usg=AOvVaw3HzqRSmqcdms--LTm64yEt) |
| 60 | [▪ SuperLocalMemory: Privacy-Preserving Multi-Agent Memory with Bayesian Trust Defense Against Memory Poisoning (Bhardwaj, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02240&sa=D&source=editors&ust=1779048536038544&usg=AOvVaw1_5P2GWgbR3nvuTa9X4o0e) |
| 61 | [▪ Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs (Xie et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02262&sa=D&source=editors&ust=1779048536038634&usg=AOvVaw08Uwr9PoB4ZqkR5BjjARNl) |
| 62 | [▪ Aggressive or Imperceptible, or Both: Network Pruning Assisted Hybrid Byzantines in Federated Learning (Ozfatura et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2404.06230&sa=D&source=editors&ust=1779048536038724&usg=AOvVaw1cEQ9N7OrZCcd-r4_PH7Cs) |
| 63 | [▪ Sparsification Under Siege: Dual-Level Defense Against Poisoning in Communication-Efficient Federated Learning (Jin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.01454&sa=D&source=editors&ust=1779048536038813&usg=AOvVaw1pke-rwr2qfHpJNWIAft81) |
| 64 | [▪ Turning Black Box into White Box: Dataset Distillation Leaks (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01053&sa=D&source=editors&ust=1779048536038898&usg=AOvVaw39Ar_UxbfaeR9EWytcuVIr) |
| 65 | [▪ Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation (Edemacu and Shokri, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00172&sa=D&source=editors&ust=1779048536038994&usg=AOvVaw3WRGdcLcgJ4IclNjDFsiIo) |
| 66 | [▪ Enhancing Continual Learning for Software Vulnerability Prediction: Addressing Catastrophic Forgetting via Hybrid-Confidence-Aware Selective Replay for Temporal LLM Fine-Tuning (Dou, Bahsi, and Guerra-Manzanares, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.23834&sa=D&source=editors&ust=1779048536039102&usg=AOvVaw3PNa9k5Q-ECz0T5HCI--0U) |
| 67 | [▪ HubScan: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems (Habler et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22427&sa=D&source=editors&ust=1779048536039192&usg=AOvVaw1i2XKCo3iz333lnUpFvtqY) |
| 68 | [▪ Poisoned Acoustics (Dahme, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22258&sa=D&source=editors&ust=1779048536039274&usg=AOvVaw2qt5zUXk3BOuMBiE2xXm9T) |
| 69 | [▪ PenTiDef: Enhancing Privacy and Robustness in Decentralized Federated Intrusion Detection Systems against Poisoning Attacks (Duy et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.17973&sa=D&source=editors&ust=1779048536039371&usg=AOvVaw1P1EPELMD6Zffjpvyzb-tx) |
| 70 | [▪ Intent Laundering: AI Safety Datasets Are Not What They Seem (Golchin and Wetter, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16729&sa=D&source=editors&ust=1779048536039491&usg=AOvVaw10UwBxJF-DKxqxRBlKQPJc) |
| 71 | [▪ SRFed: Mitigating Poisoning Attacks in Privacy-Preserving Federated Learning with Heterogeneous Data (Lu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16480&sa=D&source=editors&ust=1779048536039591&usg=AOvVaw0AzZYvtK_SitCwUEgJwpr1) |
| 72 | [▪ Efficient Semi-Supervised Adversarial Training via Latent Clustering-Based Data Reduction (Ghosh, Xu, and Zhang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.10466&sa=D&source=editors&ust=1779048536039689&usg=AOvVaw3XdeowU6adIypjrE4rc1Ns) |
| 73 | [▪ Closing the Distribution Gap in Adversarial Training for LLMs (Hu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15238&sa=D&source=editors&ust=1779048536039783&usg=AOvVaw2r8G__kpjfHmfiPjNM_1Rd) |
| 74 | [▪ Towards Privacy-Guaranteed Label Unlearning in Vertical Federated Learning: Few-Shot Forgetting without Disclosure (Gu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2410.10922&sa=D&source=editors&ust=1779048536039881&usg=AOvVaw39ZaIwU9lAIbebx9tb6ih9) |
| 75 | [▪ One RNG to Rule Them All: How Randomness Becomes an Attack Vector in Machine Learning (Prabhu, Gan, and Ghodsi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09182&sa=D&source=editors&ust=1779048536039982&usg=AOvVaw3gxah5oBjTt59dUN5XG2GS) |
| 76 | [▪ Confundo: Learning to Generate Robust Poison for Practical RAG Systems (Hu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06616&sa=D&source=editors&ust=1779048536040078&usg=AOvVaw3k0n_Qk5xrlguhybl0ik82) |
| 77 | [▪ VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems (Guan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06409&sa=D&source=editors&ust=1779048536040184&usg=AOvVaw1HCZR3_lUsAiP6CDoFxRCo) |
| 78 | [▪ ADCA: Attention-Driven Multi-Party Collusion Attack in Federated Self-Supervised Learning (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.05612&sa=D&source=editors&ust=1779048536040282&usg=AOvVaw02XDLZyp1R2DEdhseY0yt2) |
| 79 | [▪ Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning (Draganov et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04899&sa=D&source=editors&ust=1779048536040385&usg=AOvVaw3w5cpXCbLMlCpTn-ItTu8D) |
| 80 | [▪ Statistical MIA: Rethinking Membership Inference Attack for Reliable Unlearning Auditing (Sun et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01150&sa=D&source=editors&ust=1779048536040484&usg=AOvVaw3DSUNpwDIvOA5CDtHN1F0m) |
| 81 | [▪ Spattack: Subgroup Poisoning Attacks on Federated Recommender Systems (Yan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06258&sa=D&source=editors&ust=1779048536040571&usg=AOvVaw1UDbVhkZ0_nhTLcLjuZT45) |
| 82 | [▪ Stealthy Poisoning Attacks Bypass Defenses in Regression Settings (Carnerero-Cano et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22308&sa=D&source=editors&ust=1779048536040656&usg=AOvVaw2oRr-xo1QrQ_Roa7z6QBfM) |
| 83 | [▪ On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21531&sa=D&source=editors&ust=1779048536040739&usg=AOvVaw20MbIsuj6jbDU2043_fAxZ) |
| 84 | [▪ Robust Federated Learning for Malicious Clients using Loss Trend Deviation Detection (Bhaskar, B, and P, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20915&sa=D&source=editors&ust=1779048536040825&usg=AOvVaw0wpmQd6waZl0QFbs_GUVt4) |
| 85 | [▪ Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models (Chaudhari et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.19061&sa=D&source=editors&ust=1779048536040911&usg=AOvVaw3QK88Owkj1HvO-ld4_X4-h) |
| 86 | [▪ Attacks on Approximate Caches in Text-to-Image Diffusion Models (Sun, Jie, and Liu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20424&sa=D&source=editors&ust=1779048536041012&usg=AOvVaw0J4m4QXrG1roW2KOA_MuFC) |
| 87 | [▪ Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2406.18944&sa=D&source=editors&ust=1779048536041095&usg=AOvVaw1RMGkGjB_8mX6GVEm9Jjez) |
| 88 | [▪ Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12042&sa=D&source=editors&ust=1779048536041183&usg=AOvVaw09toawfs19vCXjXhXzn5xV) |
| 89 | [▪ Distributional Machine Unlearning via Selective Data Removal (Allouah, Guerraoui, and Koyejo, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15112&sa=D&source=editors&ust=1779048536041265&usg=AOvVaw1rHP6sE9bVu2LSK62zHl6f) |
| 90 | [▪ MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20412&sa=D&source=editors&ust=1779048536041348&usg=AOvVaw1cEG0psOBdUoqb0e15b20a) |
| 91 | [▪ MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07395&sa=D&source=editors&ust=1779048536041445&usg=AOvVaw3JiLzqfFFGAb6-9oKJSE9e) |
| 92 | [▪ Memory Poisoning Attack and Defense on Memory Based LLM-Agents (Sunil et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05504&sa=D&source=editors&ust=1779048536041533&usg=AOvVaw2QWyAN9QAL8aBQrNZnbAcX) |
| 93 | [▪ Practical Poisoning Attacks against Retrieval-Augmented Generation (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.03957&sa=D&source=editors&ust=1779048536041617&usg=AOvVaw0zVrhMyAokN8Fa8xGP8giC) |
| 94 | [▪ Knowledge-to-Data: LLM-Driven Synthesis of Structured Network Traffic for Testbed-Free IDS Evaluation (Kampourakis et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05022&sa=D&source=editors&ust=1779048536041704&usg=AOvVaw1W3PLHqorcGJN4hw-dJU8i) |
| 95 | [▪ Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs (P et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04275&sa=D&source=editors&ust=1779048536041789&usg=AOvVaw2TzLm7lFPte7czLV_TxvwL) |
| 96 | [▪ VFEFL: Privacy-Preserving Federated Learning against Malicious Clients via Verifiable Functional Encryption (Cai, Han, and Meng, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12846&sa=D&source=editors&ust=1779048536041876&usg=AOvVaw1ouuT9YWQQAp0348ugLT2k) |
| 97 | [▪ Quality Degradation Attack in Synthetic Data (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02947&sa=D&source=editors&ust=1779048536041959&usg=AOvVaw2-wxohobWRfa5zkBKXAI3H) |
| 98 | [▪ Low Rank Comes with Low Security: Gradient Assembly Poisoning Attacks against Distributed LoRA-based LLM Systems (Dong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00566&sa=D&source=editors&ust=1779048536042047&usg=AOvVaw0OKngHIRDPVdmWquISyKX0) |
| 99 | [▪ Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against (Enkhbayar, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22293&sa=D&source=editors&ust=1779048536042133&usg=AOvVaw0sIAA9NijTebE8FlobsMHP) |
| 100 | [▪ Certifying the Right to Be Forgotten: Primal-Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning (Jiang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23171&sa=D&source=editors&ust=1779048536042221&usg=AOvVaw33JEH67Jee0oItq41hMo3-) |
| 101 | [▪ Look Twice before You Leap: A Rational Agent Framework for Localized Adversarial Anonymization (Duan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06713&sa=D&source=editors&ust=1779048536042306&usg=AOvVaw1JY_AVkCM-GMLCTaTC5om3) |
| 102 | [▪ GShield: Mitigating Poisoning Attacks in Federated Learning (M. et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19286&sa=D&source=editors&ust=1779048536042387&usg=AOvVaw31YbqEmbcqMYk-B2t-ikgg) |
| 103 | [▪ The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation (Jia et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.19025&sa=D&source=editors&ust=1779048536042479&usg=AOvVaw3ACHu3R3FJk8p56mMH-43a) |
| 104 | [▪ SecureCode v2.0: A Production-Grade Dataset for Training Security-Aware Code Generation Models (Thornton, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18542&sa=D&source=editors&ust=1779048536042569&usg=AOvVaw2j7dbENPTiWYQVlU_1cVUz) |
| 105 | [▪ A Certified Unlearning Approach without Access to Source Data (Basaran et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06486&sa=D&source=editors&ust=1779048536042653&usg=AOvVaw1KeQZuUL-m-mONR4-mqnb_) |
| 106 | [▪ Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias (Bernstein et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.17361&sa=D&source=editors&ust=1779048536042750&usg=AOvVaw0OuYxkQSKTgtLPJ2WSC-gk) |
| 107 | [▪ Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15335&sa=D&source=editors&ust=1779048536042838&usg=AOvVaw33qt6fGmDNc1pxHa7JkDUd) |
| 108 | [▪ Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space (Zhou and Wang, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14448&sa=D&source=editors&ust=1779048536042926&usg=AOvVaw25Z3WGyvH0ADuRt7OiBRxg) |
| 109 | [▪ Evaluating Adversarial Attacks on Federated Learning for Temperature Forecasting (Chichifoi, Merizzi, and Colajanni, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13207&sa=D&source=editors&ust=1779048536043012&usg=AOvVaw3d7Oi3p7JDVNdkrMRtyOYm) |
| 110 | [▪ CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation (Yang and Ardakanian, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12086&sa=D&source=editors&ust=1779048536043096&usg=AOvVaw28nVQovwGiYZCjzW6bTNTF) |
| 111 | [▪ Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization (He et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.03752&sa=D&source=editors&ust=1779048536043180&usg=AOvVaw1fjW-eZWmqxQ5MFk-CMsCq) |
| 112 | [▪ Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search (Coalson et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.06073&sa=D&source=editors&ust=1779048536043263&usg=AOvVaw3P1cueKreFCppAJW5W7OJc) |
| 113 | [▪ MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks (Chen et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08289&sa=D&source=editors&ust=1779048536043349&usg=AOvVaw3TtnzMY1u46sDOSJs38XK7) |
| 114 | [▪ Data Taggants: Dataset Ownership Verification via Harmless Targeted Data Poisoning (Bouaziz, Usunier, and El-Mhamdi, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.09101&sa=D&source=editors&ust=1779048536043501&usg=AOvVaw3sT4nWem8qKPWM60NgiWcF) |
| 115 | [▪ DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification (Liu and Papadimitratos, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06172&sa=D&source=editors&ust=1779048536043615&usg=AOvVaw3rShbhlCcVkbXI_an6m3yN) |
| 116 | [▪ Adversarial Robustness of Traffic Classification under Resource Constraints: Input Structure Matters (Chehade et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02276&sa=D&source=editors&ust=1779048536043713&usg=AOvVaw3sjsOJ_QSpu6oBhEY_eWza) |
| 117 | [▪ Bias Injection Attacks on RAG Databases and Sanitization Defenses (Wu and Saxena, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00804&sa=D&source=editors&ust=1779048536043807&usg=AOvVaw2ojbbpKTTP1-9be3IjfMam) |
| 118 | [▪ Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models (Tice et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.01784&sa=D&source=editors&ust=1779048536043899&usg=AOvVaw0mDiTPg3lsJTD9nJwBYz0K) |
| 119 | [▪ Dataset Poisoning Attacks on Behavioral Cloning Policies (Kalra et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20992&sa=D&source=editors&ust=1779048536043991&usg=AOvVaw384WKiIgidrVEF9IybeZPt) |
| 120 | [▪ Synthetic Data: AI's New Weapon Against Android Malware (Nogueira et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19649&sa=D&source=editors&ust=1779048536044087&usg=AOvVaw1zg5BYTgfqUzOz8y0S5VVT) |
| 121 | [▪ FedPoisonTTP: A Threat Model and Poisoning Attack for Federated Test-Time Personalization (Iftee et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19248&sa=D&source=editors&ust=1779048536044188&usg=AOvVaw03OvnasOq1x8LZfBECEJsk) |
| 122 | [▪ Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks (G, Govindarajulu, and Shah, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22060&sa=D&source=editors&ust=1779048536044274&usg=AOvVaw3u5eV9qg5C2QbIY_nGfGWH) |
| 123 | [▪ One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image (Shereen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.02132&sa=D&source=editors&ust=1779048536044358&usg=AOvVaw3wi9kVJvsAT4KOmnF4NdUv) |
| 124 | [▪ Robustness of LLM-enabled vehicle trajectory prediction under data security threats (Wang and Liu, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13753&sa=D&source=editors&ust=1779048536044446&usg=AOvVaw38AgiD_qE3A2BEBqd2223h) |
| 125 | [▪ Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System (He et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06059&sa=D&source=editors&ust=1779048536044533&usg=AOvVaw32ZxTH7ZD7toX8eDHClIlk) |
| 126 | [▪ CITADEL: A Semi-Supervised Active Learning Framework for Malware Detection Under Continuous Distribution Drift (Haque et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11979&sa=D&source=editors&ust=1779048536044622&usg=AOvVaw2n3A4FRO8gNVTsQKdQAiL1) |
| 127 | [▪ AMUN: Adversarial Machine UNlearning (Ebrahimpour-Boroojeny, Sundaram, and Chandrasekaran, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.00917&sa=D&source=editors&ust=1779048536044709&usg=AOvVaw0pYwZMOd7WvsEkdUEmtazq) |
| 128 | [▪ Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis (Abtahi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11020&sa=D&source=editors&ust=1779048536044794&usg=AOvVaw0QyUa7tcl_445luWWK2N60) |
| 129 | [▪ Taught by the Flawed: How Dataset Insecurity Breeds Vulnerable AI Code (Xia and Alalfi, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09879&sa=D&source=editors&ust=1779048536044877&usg=AOvVaw0YMl6Ac1AMPja-akfIA6HJ) |
| 130 | [▪ DP-GENG : Differentially Private Dataset Distillation Guided by DP-Generated Data (Shi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09876&sa=D&source=editors&ust=1779048536044964&usg=AOvVaw3j8HFqjLTAj4HjacUX4rHz) |
| 131 | [▪ Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06151&sa=D&source=editors&ust=1779048536045049&usg=AOvVaw0PZ2K8n3iYZJ60IIiBfoRe) |
| 132 | [▪ PrometheusFree: Concurrent Detection of Laser Fault Injection Attacks in Optical Neural Networks (Nishida et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14741&sa=D&source=editors&ust=1779048536045133&usg=AOvVaw0-TGT06vfaT8DaDRKe8oc1) |
| 133 | [▪ Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data (Song et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07242&sa=D&source=editors&ust=1779048536045216&usg=AOvVaw1HPT6wqeMChxfu_3obTOB9) |
| 134 | [▪ Adversarial Node Placement in Decentralized Federated Learning: Maximum Spanning-Centrality Strategy and Performance Analysis (Piaseczny et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06742&sa=D&source=editors&ust=1779048536045303&usg=AOvVaw1QnzTitVr8MZbT-eKV8Asf) |
| 135 | [▪ IndirectAD: Practical Data Poisoning Attacks against Recommender Systems for Item Promotion (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05845&sa=D&source=editors&ust=1779048536045388&usg=AOvVaw3jw8_QuiL0MbzLHkL7pZje) |
| 136 | [▪ Retrieval-Augmented Review Generation for Poisoning Recommender Systems (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.15252&sa=D&source=editors&ust=1779048536045488&usg=AOvVaw3J5eI1lnVJHrxOk-HZlfPp) |
| 137 | [▪ Adaptive and Robust Data Poisoning Detection and Sanitization in Wearable IoT Systems using Large Language Models (Mithsara et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02894&sa=D&source=editors&ust=1779048536045575&usg=AOvVaw3Rsbq8P_OOfXdS6QexXfbX) |
| 138 | [▪ On The Dangers of Poisoned LLMs In Security Automation (Karlsen and Eilertsen, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02600&sa=D&source=editors&ust=1779048536045672&usg=AOvVaw2srQ2LneizwxiYg1SnY9U5) |
| 139 | [▪ Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems (Kim, Lee, and Koo, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01268&sa=D&source=editors&ust=1779048536045761&usg=AOvVaw3G7d30BfJVGgEFpnbC52Ow) |
| 140 | [▪ Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning (Churina, Chebrolu, and Jaidka, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26829&sa=D&source=editors&ust=1779048536045846&usg=AOvVaw1FKnlaamizeHpsdGmGz3Dk) |
| 141 | [▪ VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes (Yu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26833&sa=D&source=editors&ust=1779048536045935&usg=AOvVaw2OPaB8uhQiQ0Pf_qY93D4O) |
| 142 | [▪ PEEL: A Poisoning-Exposing Encoding Theoretical Framework for Local Differential Privacy (Shuai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26102&sa=D&source=editors&ust=1779048536046021&usg=AOvVaw10g99aXZH6TRWUFoo4lgf2) |
| 143 | [▪ On the Fragility of Contribution Score Computation in Federated Learning (Pejo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.19921&sa=D&source=editors&ust=1779048536046110&usg=AOvVaw0bsRJEhpqe4S0mV7cei5Us) |
| 144 | [▪ RAGRank: Using PageRank to Counter Poisoning in CTI LLM Pipelines (Jia et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20768&sa=D&source=editors&ust=1779048536046200&usg=AOvVaw1ap6h48b_cI67XNr5jIW9T) |
| 145 | [▪ The Black Tuesday Attack: how to crash the stock market with adversarial examples to financial forecasting models (Hofweber et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18990&sa=D&source=editors&ust=1779048536046285&usg=AOvVaw3XL3-8bSnGPCMgh_g6bCUB) |
| 146 | [▪ Delta-Influence: Unlearning Poisons via Influence Functions (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.13731&sa=D&source=editors&ust=1779048536046367&usg=AOvVaw0nJOgJRXgNRe-QqKprk5v4) |
| 147 | [▪ Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.13772&sa=D&source=editors&ust=1779048536046457&usg=AOvVaw3bl_SftAi7sX0pHo4YSaKE) |
| 148 | [▪ Byzantine Failures Harm the Generalization of Robust Distributed Learning Algorithms More Than Data Poisoning (Boudou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.18020&sa=D&source=editors&ust=1779048536046543&usg=AOvVaw00Q1nEpWxVTIYQ-W2BvOYK) |
| 149 | [▪ ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13842&sa=D&source=editors&ust=1779048536046624&usg=AOvVaw0IPtItAf9dhmzyw8H1Osfb) |
| 150 | [▪ Cascading Adversarial Bias from Injection to Distillation in Language Models (Chaudhari et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.24842&sa=D&source=editors&ust=1779048536046708&usg=AOvVaw0SY25jOt5IXo51POKeOhcr) |
| 151 | [▪ Tracing Back the Malicious Clients in Poisoning Attacks to Federated Learning (Jia et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.07221&sa=D&source=editors&ust=1779048536046790&usg=AOvVaw2c_UbBx3krI12DhuVMLNrE) |
| 152 | [▪ Fairness-Constrained Optimization Attack in Federated Learning (Kasyap et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12143&sa=D&source=editors&ust=1779048536046871&usg=AOvVaw2CAdYMrJoOrmn_P26nKQEC) |
| 153 | [▪ GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models (Luo et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.10983&sa=D&source=editors&ust=1779048536046958&usg=AOvVaw3bGREoLSoAQZIB_quA_AW5) |
| 154 | [▪ When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge? (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.15267&sa=D&source=editors&ust=1779048536047044&usg=AOvVaw2VtoTmcp56Gzyp8owij8Ax) |
| 155 | [▪ When Vision Fails: Text Attacks Against ViT and OCR (Boucher et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2306.07033&sa=D&source=editors&ust=1779048536047124&usg=AOvVaw3HCaMdvglAIRfbQtNSNFwg) |
| 156 | [▪ RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation (Stambolic, Dhar, and Cavigelli, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11195&sa=D&source=editors&ust=1779048536047213&usg=AOvVaw3oFeX-r3HHE55q7_Lp7mXW) |
| 157 | [▪ How Secure is Forgetting? Linking Machine Unlearning to Machine Learning Attacks (P. et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.20257&sa=D&source=editors&ust=1779048536047301&usg=AOvVaw1tvCIma1eWuKaeS8QoRavL) |
| 158 | [▪ Provable Watermarking for Data Poisoning Attacks (Zhu, Yu, and Gao, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09210&sa=D&source=editors&ust=1779048536047384&usg=AOvVaw3ovVPMeb_wDuiX7ikaBucS) |
| 159 | [▪ MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks (Ha et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.17832&sa=D&source=editors&ust=1779048536047473&usg=AOvVaw2e6H6Oq1gA0juYpFSPByZP) |
| 160 | [▪ A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models (Cai, Li, and Zhang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.02441&sa=D&source=editors&ust=1779048536047560&usg=AOvVaw1-QU8NXgWi8MVBWcHKSfY3) |
| 161 | [▪ Impact of Dataset Properties on Membership Inference Vulnerability of Deep Transfer Learning (Tobaben et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.06674&sa=D&source=editors&ust=1779048536047645&usg=AOvVaw35QvDQiz9TsTVs3nTd6PL3) |
| 162 | [▪ From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse (Jhuma and Faisal, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03636&sa=D&source=editors&ust=1779048536047732&usg=AOvVaw1ySR5e4V1eNQ_FcyAJhJ7R) |
| 163 | [▪ Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training (Yu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.11191&sa=D&source=editors&ust=1779048536047816&usg=AOvVaw0VcYOCATv-mK1ko4v2qRtN) |
| 164 | [▪ Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors (Chen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00586&sa=D&source=editors&ust=1779048536047900&usg=AOvVaw224Hd515zYjrLbitccWDBm) |
| 165 | [▪ From Mean to Extreme: Formal Differential Privacy Bounds on the Success of Real-World Data Reconstruction Attacks (Riess et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.12861&sa=D&source=editors&ust=1779048536047984&usg=AOvVaw06lBOdzUW6KXJpGoY8aN8v) |
| 166 | [▪ Fast Exact Unlearning for In-Context Learning Data for LLMs (Muresanu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.00751&sa=D&source=editors&ust=1779048536048069&usg=AOvVaw2nny3dm-vD5tTqckWPRZSM) |
| 167 | [▪ The Impact of Scaling Training Data on Adversarial Robustness (Zimmerli et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.25927&sa=D&source=editors&ust=1779048536048172&usg=AOvVaw3jHKvh5CAOSd4ErFM-z9Vb) |
| 168 | [▪ Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks? (Sasnauskas, Yalın, and Radanović, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06891&sa=D&source=editors&ust=1779048536048442&usg=AOvVaw0-J_rqtEy891ZKNXuxxj2x) |
| 169 | [▪ FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems (Long and Li, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24408&sa=D&source=editors&ust=1779048536048560&usg=AOvVaw1MdJbJeWRrdWj3JoRYcbxT) |
| 170 | [▪ Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data (Liang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23041&sa=D&source=editors&ust=1779048536048663&usg=AOvVaw033Qrj9iY_Vf8OrrYuFQK1) |
| 171 | [▪ AntiFLipper: A Secure and Efficient Defense Against Label-Flipping Attacks in Federated Learning (Rahman et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22873&sa=D&source=editors&ust=1779048536048759&usg=AOvVaw1Yfh1qdN9siHs2Pbr7_a7Y) |
| 172 | [▪ Defending Against Beta Poisoning Attacks in Machine Learning Models (Gulciftci and Gursoy, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01276&sa=D&source=editors&ust=1779048536048853&usg=AOvVaw0_O86qGT8rb92sXeu2Bi-I) |
| 173 | [▪ UTrace: Poisoning Forensics for Private Collaborative Learning (Rose et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.15126&sa=D&source=editors&ust=1779048536048947&usg=AOvVaw3o2PliT3BXHi0DKtJ3gmwL) |
| 174 | [▪ Graph Representation-based Model Poisoning on Federated Large Language Models (Cai et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.01694&sa=D&source=editors&ust=1779048536049043&usg=AOvVaw3JZpaf3CiSTUwpyvCZuCiS) |
| 175 | [▪ Scalable contribution bounding to achieve privacy (Cohen-Addad et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.23432&sa=D&source=editors&ust=1779048536049144&usg=AOvVaw01H0yYsTV1b0xpTrzBRCiL) |
| 176 | [▪ Privacy-Preserving Federated Learning Scheme with Mitigating Model Poisoning Attacks: Vulnerabilities and Countermeasures (Wu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.23622&sa=D&source=editors&ust=1779048536049249&usg=AOvVaw0DABhg2bz11Fs85b_wiukB) |
| 177 | [▪ Generalizable Targeted Data Poisoning against Varying Physical Objects (Chen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.03908&sa=D&source=editors&ust=1779048536049334&usg=AOvVaw1xgnJO8Q8fVJwF4hDHwA6z) |
| 178 | [▪ CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset Separation (Xu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05113&sa=D&source=editors&ust=1779048536049424&usg=AOvVaw0Ow60x-OsV4p03tMYsZiKU) |
| 179 | [▪ PPFPL: Cross-silo Privacy-preserving Federated Prototype Learning Against Data Poisoning Attacks on Non-IID Data (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.03173&sa=D&source=editors&ust=1779048536049513&usg=AOvVaw2zF1EZHFxYjTT_v6PEtXXl) |
| 180 | [▪ OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud (Yuan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.19334&sa=D&source=editors&ust=1779048536049598&usg=AOvVaw0tL4RP0BFLX-mAFUF0WN7V) |
| 181 | [▪ Sparsification Under Siege: Defending Against Poisoning Attacks in Communication-Efficient Federated Learning (Jin et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.01454&sa=D&source=editors&ust=1779048536049683&usg=AOvVaw3IR3VvP-78ypcq0RIZzirR) |
| 182 | [▪ Rethinking Data Protection in the (Generative) Artificial Intelligence Era (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.03034&sa=D&source=editors&ust=1779048536049766&usg=AOvVaw33vQxPA-u0Ze6hUI8Jf-rJ) |
| 183 | [▪ A Bayesian Incentive Mechanism for Poison-Resilient Federated Learning (Commey et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12439&sa=D&source=editors&ust=1779048536049848&usg=AOvVaw1goCI2Ea7xXcOfaec1odjN) |
| 184 | [▪ A Privacy-Preserving Framework for Advertising Personalization Incorporating Federated Learning and Differential Privacy (Li, Lin, and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12098&sa=D&source=editors&ust=1779048536049936&usg=AOvVaw26hQ6mmNmHipoyQOQQIOu-) |
| 185 | [▪ Effective Fine-Tuning of Vision Transformers with Low-Rank Adaptation for Privacy-Preserving Image Classification (Lin, Imaizumi, and Kiya, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11943&sa=D&source=editors&ust=1779048536050023&usg=AOvVaw3aGYXerkl2QOEdYu0T21AN) |
| 186 | [▪ Adaptive Federated Learning with Functional Encryption: A Comparison of Classical and Quantum-safe Options (Sorbera et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.00563&sa=D&source=editors&ust=1779048536050131&usg=AOvVaw2AqDoQh62LqRmC5RFl0cVV) |
| 187 | [▪ When and Where do Data Poisons Attack Textual Inversion? (Styborski et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10578&sa=D&source=editors&ust=1779048536050223&usg=AOvVaw2K_ZjxA89FxGZvNftCni6d) |
| 188 | [▪ Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts (Zarlenga et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.17921&sa=D&source=editors&ust=1779048536050315&usg=AOvVaw0PJCfYbI8z_FYEQvJ0_lkx) |
| 189 | [▪ RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08862&sa=D&source=editors&ust=1779048536050414&usg=AOvVaw1u6mOI-_bSq30Y_EMr9-M0) |
| 190 | [▪ Q-Detection: A Quantum-Classical Hybrid Poisoning Attack Detection Method (He et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06262&sa=D&source=editors&ust=1779048536050509&usg=AOvVaw0vvF4P-dA-mcKOkqFls7T0) |
| 191 | [▪ Phantom Subgroup Poisoning: Stealth Attacks on Federated Recommender Systems (Yan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06258&sa=D&source=editors&ust=1779048536050601&usg=AOvVaw0hELJUD8tSsivf_96TkTdl) |
| 192 | [▪ Privacy-preserving Machine Learning in Internet of Vehicle Applications: Fundamentals, Recent Advances, and Future Direction (Islam and Zulkernine, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.01089&sa=D&source=editors&ust=1779048536050717&usg=AOvVaw2tlCGEhEbyUpzENMbKaYBP) |
| 193 | [▪ The Impact of Event Data Partitioning on Privacy-aware Process Discovery (Lim et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06008&sa=D&source=editors&ust=1779048536050807&usg=AOvVaw1vsqslfj-Dzx9vYIzyaFmd) |
| 194 | [▪ Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05728&sa=D&source=editors&ust=1779048536050891&usg=AOvVaw0ByRASgVSWVzqZu5XdyI7a) |
| 195 | [▪ DATABench: Evaluating Dataset Auditing in Deep Learning from an Adversarial Perspective (Shao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05622&sa=D&source=editors&ust=1779048536050980&usg=AOvVaw2GE2gdm6Dx_S8dBIKTRdve) |
| 196 | [▪ A Linear Approach to Data Poisoning (Granziol and Flynn, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15175&sa=D&source=editors&ust=1779048536051065&usg=AOvVaw336fq72fgqI8M518kXap2D) |
| 197 | [▪ Traceback of Poisoning Attacks to Retrieval-Augmented Generation (Zhang et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21668&sa=D&source=editors&ust=1779048536051151&usg=AOvVaw2USYpyDRFQc9NDYwYdPUgE) |
| 198 | [▪ Machine Unlearning Fails to Remove Data Poisoning Attacks (Pawelczyk , Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.17216&sa=D&source=editors&ust=1779048536051237&usg=AOvVaw05ulQKTYCxXCKyAPQEqiRE) |
| 199 | [▪ Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning (Wang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.14217&sa=D&source=editors&ust=1779048536051324&usg=AOvVaw1yEKhHSCXh6jdgghzPJz0W) |
| 200 | [▪ Adversarial Poisoning Attack on Quantum Machine Learning Models (Kundu and Ghosh, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14412&sa=D&source=editors&ust=1779048536051419&usg=AOvVaw1RD4Llj52gqNB4izmEvc6L) |
| 201 | [▪ Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws (Bowen et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02946&sa=D&source=editors&ust=1779048536051508&usg=AOvVaw3rcnCWYmJ9dLml67yoSP3y) |
| 202 | [▪ Inverting Gradient Attacks Naturally Makes Data Poisons: An Availability Attack on Neural Networks (Bouaziz, Mhamdi, and Usunier, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.21453&sa=D&source=editors&ust=1779048536051596&usg=AOvVaw1t2GHezPolIe1oOLXZf_9h) |
| 203 | [▪ Certified Robustness to Data Poisoning in Gradient-Based Training (Sosnin et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.05670&sa=D&source=editors&ust=1779048536051680&usg=AOvVaw0pTt1fRZITFfyr9oThps16) |
| 204 | [▪ Controlled Generation of Natural Adversarial Documents for Stealthy Retrieval Poisoning (Zhang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02163&sa=D&source=editors&ust=1779048536051763&usg=AOvVaw1M-S8YVyOnyDj_yYszdQAj) |
| 205 | [▪ Securing Voice Authentication Applications Against Targeted Data Poisoning (Mohammadi et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.17277&sa=D&source=editors&ust=1779048536051844&usg=AOvVaw13RMKfLIuNwqc5bJ0XBcll) |
| 206 | [▪ Data Poisoning-based Backdoor Attack Framework against Supervised Learning Rules of Spiking Neural Networks (Jin et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.15670&sa=D&source=editors&ust=1779048536051928&usg=AOvVaw31oY1fgWAvMHAsm1Kzg8qE) |
| 207 | [▪ Hiding Backdoors within Event Sequence Data via Poisoning Attacks (Ermilova et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2308.102&sa=D&source=editors&ust=1779048536052011&usg=AOvVaw2bqxXqLyL_HOdJpPCpztL9) |
| 208 | [▪ ConfusedPilot: Confused Deputy Risks in RAG-based LLMs (RoyChowdhury et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.04870&sa=D&source=editors&ust=1779048536052096&usg=AOvVaw1kg5Q3OWjJijedz_LMWSlk) |
| 209 | [▪ PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models (Zou et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.07867%23&sa=D&source=editors&ust=1779048536052183&usg=AOvVaw1wL4CXIvOBec5RFZjotp2C) |
| 210 | [▪ Scaling Laws for Data Poisoning in LLMs (Bowen et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.02946&sa=D&source=editors&ust=1779048536052268&usg=AOvVaw0TshdQ8PAOucheTvFhCdpn) |
| 211 | [▪ Threats, Attacks, and Defenses in Machine Unlearning: A Survey (Liu et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.13682&sa=D&source=editors&ust=1779048536052373&usg=AOvVaw2avNeoZ43kxjfiNs8BmwDa) |
| 212 | [▪ Debiased Graph Poisoning Attack via Contrastive Surrogate Objective (Yoon et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.19155&sa=D&source=editors&ust=1779048536052477&usg=AOvVaw3CTkyleb2YVbPaXCiPHR8A) |

|     |
| --- |
| Data Poisoning and Simulated Publication of Poisoned Public Datasets |

**>**

**<**

‍

#### Model (Mis)Interpretability

Added this subsection to cover cybersecurity issues that arise from interpretability issues.

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model (Mis)Interpretability

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ When the Ruler is Broken: Parsing-Induced Suppression in LLM-Based Security Log Evaluation (Garware and Zisad, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07293&sa=D&source=editors&ust=1779048536729172&usg=AOvVaw066t0yNrYMqjquV7G5HzUK) |
| 3 | [▪ How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27714&sa=D&source=editors&ust=1779048536729446&usg=AOvVaw1sfDDvKathH2zB8biNZN8K) |
| 4 | [▪ Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models (Min, Pham, and Sun, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24542&sa=D&source=editors&ust=1779048536729683&usg=AOvVaw3ZoSG9jQ6IExYoLfUKge-e) |
| 5 | [▪ Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways (Lin et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21083&sa=D&source=editors&ust=1779048536729886&usg=AOvVaw0QUdVj_Esl0afVvNEKtTDO) |
| 6 | [▪ Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs (Agarwal et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20945&sa=D&source=editors&ust=1779048536730056&usg=AOvVaw0PIgfte_oFbcUJiSYfRJyI) |
| 7 | [▪ Attacks Meet Interpretability (AmI) Evaluation and Findings (Ma, Ye, and Mehnaz, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2310.08808&sa=D&source=editors&ust=1779048536730184&usg=AOvVaw1Zp7i5WofxC4TBmTJALltt) |
| 8 | [▪ Attribution-Driven Explainable Intrusion Detection with Encoder-Based Large Language Models (Biswas et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06266&sa=D&source=editors&ust=1779048536730289&usg=AOvVaw3m0dBMeGffYgVXTSlWAH3y) |
| 9 | [▪ Systematic Integration of Digital Twins and Constrained LLMs for Interpretable Cyber-Physical Anomaly Detection (Kampourakis, Gkioulos, and Katsikas, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03790&sa=D&source=editors&ust=1779048536730391&usg=AOvVaw3TIQYndPjzCpqeFVz7lkHb) |
| 10 | [▪ Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models (Biskupski and Kleber, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.22214&sa=D&source=editors&ust=1779048536730484&usg=AOvVaw1r2piU9NsKUtaRG_UDLOvw) |
| 11 | [▪ Manifold of Failure: Behavioral Attraction Basins in Language Models (Munshi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22291&sa=D&source=editors&ust=1779048536730574&usg=AOvVaw2J9znxi1LRDitmauvSVLs5) |
| 12 | [▪ Layer-Targeted Multilingual Knowledge Erasure in Large Language Models (Li, Chandrasekaran, and Yu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22562&sa=D&source=editors&ust=1779048536730694&usg=AOvVaw1dvxbRyx2eLbkiCdscl9J-) |
| 13 | [▪ MARVEL: Multi-Agent RTL Vulnerability Extraction using Large Language Models (Collini et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11963&sa=D&source=editors&ust=1779048536730782&usg=AOvVaw2IyLekYXzZiFlBz_TqYHtQ) |
| 14 | [▪ Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling (Srisumrith and Sodsee, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19087&sa=D&source=editors&ust=1779048536730934&usg=AOvVaw0qtWUcO7uDbthHeKYw7ID0) |
| 15 | [▪ LLM-FS: Zero-Shot Feature Selection for Effective and Interpretable Malware Detection (Gill, K, and D, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09634&sa=D&source=editors&ust=1779048536731063&usg=AOvVaw15m5wqJS3YXPY7OXNc9diE) |
| 16 | [▪ Empirical Analysis of Adversarial Robustness and Explainability Drift in Cybersecurity Classifiers (Rajhans and Khawarey, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06395&sa=D&source=editors&ust=1779048536731195&usg=AOvVaw167eWbX5aaAYM39FIZzKXT) |
| 17 | [▪ Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts (Rajhans, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22653&sa=D&source=editors&ust=1779048536731337&usg=AOvVaw2GejPVjSp-Xn-x2l1Yn-hl) |
| 18 | [▪ The Semantic Trap: Do Fine-tuned LLMs Learn Vulnerability Root Cause or Just Functional Pattern? (Huang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22655&sa=D&source=editors&ust=1779048536731491&usg=AOvVaw2ZQN_dbl6GQJCjrsA4LMdI) |
| 19 | [▪ Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents (Goel et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11708&sa=D&source=editors&ust=1779048536731633&usg=AOvVaw3n7CVrt5fCxX1ZlaZMBKzO) |
| 20 | [▪ Attesting Model Lineage by Consisted Knowledge Evolution with Fine-Tuning Trajectory (Shang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11683&sa=D&source=editors&ust=1779048536731786&usg=AOvVaw1z82QOfNYFOGL5xB-XfFk3) |
| 21 | [▪ Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test (Haque et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03640&sa=D&source=editors&ust=1779048536731991&usg=AOvVaw1Qx5tvfSDDq632UAZ0sKl5) |
| 22 | [▪ Explain First, Trust Later: LLM-Augmented Explanations for Graph-Based Crypto Anomaly Detection (Watson, Richards, and Schiff, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.14933&sa=D&source=editors&ust=1779048536732149&usg=AOvVaw1bNdTVtn4kVilvRBkGJEj_) |
| 23 | [▪ BEACON: A Unified Behavioral-Tactical Framework for Explainable Cybercrime Analysis with Large Language Models (Sachdeva et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06555&sa=D&source=editors&ust=1779048536732320&usg=AOvVaw2CYJ6IaTyMxsR6DrendFVR) |
| 24 | [▪ Can VLMs Detect and Localize Fine-Grained AI-Edited Images? (Sun et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15644&sa=D&source=editors&ust=1779048536732480&usg=AOvVaw3ReJ-mtltft_u755b-vhY0) |
| 25 | [▪ The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution (Tihanyi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10493&sa=D&source=editors&ust=1779048536732650&usg=AOvVaw0LaVPdju_1DfJdocciF1jK) |
| 26 | [▪ LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.09443&sa=D&source=editors&ust=1779048536732789&usg=AOvVaw0g4ZvgAsPjOfI-XHgnO6tm) |
| 27 | [▪ Interpretable Ransomware Detection Using Hybrid Large Language Models: A Comparative Analysis of BERT, RoBERTa, and DeBERTa Through LIME and SHAP (Ngoie et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13517&sa=D&source=editors&ust=1779048536732952&usg=AOvVaw2UgqSA5KP0AYQxE1PqL62c) |
| 28 | [▪ Interpretable LLM Guardrails via Sparse Representation Steering (He et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.16851&sa=D&source=editors&ust=1779048536733106&usg=AOvVaw3nhscHC_Ho3gD-vVP4dHNk) |
| 29 | [▪ RepoMark: A Code Usage Auditing Framework for Code Large Language Models (Qu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21432&sa=D&source=editors&ust=1779048536733244&usg=AOvVaw1gJw4vpandR5KcEvIAQovU) |
| 30 | [▪ Model Provenance Testing for Large Language Models (Nikolic, Baluta, and Saxena, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.00706&sa=D&source=editors&ust=1779048536733407&usg=AOvVaw23Q3lKRhzC4ugW4zB0GSmn) |
| 31 | [▪ Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications (Mia and Pritom, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03623&sa=D&source=editors&ust=1779048536733552&usg=AOvVaw21bCZx0UwFal8iFMI0NEMq) |
| 32 | [▪ Reconstructing Trust Embeddings from Siamese Trust Scores: A Direct-Sum Approach with Fixed-Point Semantics (Alpay, Alpay, and Kilictas, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01479&sa=D&source=editors&ust=1779048536733699&usg=AOvVaw0hfWyGZfGhZw4FmxH9ku3H) |
| 33 | [▪ Preliminary Investigation into Uncertainty-Aware Attack Stage Classification (Gaudenzi et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00368&sa=D&source=editors&ust=1779048536733828&usg=AOvVaw11CvB1DlTixuNC0U_KbjYk) |
| 34 | [▪ Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems (Liu et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.23453&sa=D&source=editors&ust=1779048536733934&usg=AOvVaw2d_kiWVt_NqQ2Tw-8z8Zfz) |
| 35 | [▪ CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation (Manuel and Rad, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22066&sa=D&source=editors&ust=1779048536734061&usg=AOvVaw0cdaju_Q40pR9YDAn2zsYs) |
| 36 | [▪ Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection (Liang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22447&sa=D&source=editors&ust=1779048536734295&usg=AOvVaw1l759DsXRAEuKhreOeRZms) |
| 37 | [▪ POLARIS: Explainable Artificial Intelligence for Mitigating Power Side-Channel Leakage (Mahfuz et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22177&sa=D&source=editors&ust=1779048536734527&usg=AOvVaw097o4mu97d4k0Zlm2mf0px) |
| 38 | [▪ Adversarial attacks and defenses in explainable artificial intelligence: A survey (Baniecki and Biecek, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2306.06123&sa=D&source=editors&ust=1779048536734709&usg=AOvVaw3VEbAmryZrPF1YJlPxw152) |
| 39 | [▪ Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses? (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21817&sa=D&source=editors&ust=1779048536734923&usg=AOvVaw3s9mCYDQOm5gdCdXBbdD0M) |
| 40 | [▪ Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs (Pan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15265&sa=D&source=editors&ust=1779048536735110&usg=AOvVaw3ym6iNKo12lmxs9SM81Yqp) |
| 41 | [▪ KGV: Integrating Large Language Models with Knowledge Graphs for Cyber Threat Intelligence Credibility Assessment (Wu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2408.08088&sa=D&source=editors&ust=1779048536735291&usg=AOvVaw0ZPsyAiwQBklhgrzwYLx2Q) |
| 42 | [▪ Scheduzz: Constraint-based Fuzz Driver Generation with Dual Scheduling (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18289&sa=D&source=editors&ust=1779048536735463&usg=AOvVaw3z7O6czxVQ4eHw7zQiKbZP) |
| 43 | [▪ Lower Bounds for Public-Private Learning under Distribution Shift (Setlur, Thaker, and Ullman, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17895&sa=D&source=editors&ust=1779048536735606&usg=AOvVaw3hOXYo_XwW4l9oo8s6mGMN) |
| 44 | [▪ LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models (Lekssays et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16585&sa=D&source=editors&ust=1779048536735753&usg=AOvVaw3c9wT-CtMuRyxw0TLpLH9C) |
| 45 | [▪ Explainable Vulnerability Detection in C/C++ Using Edge-Aware Graph Attention Networks (Haque et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16540&sa=D&source=editors&ust=1779048536735884&usg=AOvVaw1E8dcRn_xcFXZjAPW8ahIx) |
| 46 | [▪ Attacking interpretable NLP systems (Abdukhamidov et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16164&sa=D&source=editors&ust=1779048536736041&usg=AOvVaw3q9jU5veoHAOEkaLPNkG_n) |
| 47 | [▪ Too Much to Trust? Measuring the Security and Cognitive Impacts of Explainability in AI-Driven SOCs (Rastogi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.02065&sa=D&source=editors&ust=1779048536736201&usg=AOvVaw3ifCgteMduq0bkLbJbVym7) |
| 48 | [▪ Distributional Unlearning: Forgetting Distributions, Not Just Samples (Allouah, Guerraoui, and Koyejo, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15112&sa=D&source=editors&ust=1779048536736358&usg=AOvVaw2A3H9zBUf_KdYQqzLhe_PA) |
| 49 | [▪ Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack (Abdukhamidov et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14248&sa=D&source=editors&ust=1779048536736522&usg=AOvVaw2-Qr82PIadsuFV_2fRnUjF) |
| 50 | [▪ Multi-Granular Discretization for Interpretable Generalization in Precise Cyberattack Identification (Chung, Huang, and Pai, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14223&sa=D&source=editors&ust=1779048536736688&usg=AOvVaw2_riISPBlUs_98y8mOZvCQ) |
| 51 | [▪ GPU-Accelerated Interpretable Generalization for Rapid Cyberattack Detection and Forensics (Huang, Chung, and Pai, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14222&sa=D&source=editors&ust=1779048536736818&usg=AOvVaw1D077h9YiHWTfaJb0Vm0pO) |
| 52 | [▪ SoK: Semantic Privacy in Large Language Models (Ma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.23603&sa=D&source=editors&ust=1779048536736918&usg=AOvVaw340Vs-SkHpq_2a4A9ZIrx1) |
| 53 | [▪ What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift (Chang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21042&sa=D&source=editors&ust=1779048536737012&usg=AOvVaw3H5Pp6oh2xhPRsVartI-Av) |
| 54 | [▪ Interpreting Differential Privacy in Terms of Disclosure Risk (Kazan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09699&sa=D&source=editors&ust=1779048536737094&usg=AOvVaw0Y2TnQ6WP2k6YOqLhZdHvC) |
| 55 | [▪ White-Basilisk: A Hybrid Model for Code Vulnerability Detection (Lamprou et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08540&sa=D&source=editors&ust=1779048536737179&usg=AOvVaw02_enyZ18Es9GsbIS8Zo0h) |
| 56 | [▪ Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models (Tian et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.09446&sa=D&source=editors&ust=1779048536737262&usg=AOvVaw01rVUay31k_btx43Jrkdx1) |
| 57 | [▪ Protecting Classifiers From Attacks (Gallego et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2004.08705&sa=D&source=editors&ust=1779048536737347&usg=AOvVaw2yIXCQjM9-jA0GO2FkI9g1) |
| 58 | [▪ Open Problems in Mechanistic Interpretability (Sharkley et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16496&sa=D&source=editors&ust=1779048536737433&usg=AOvVaw1q85oNDkaNwuluX2hIxBma) |
| 59 | [▪ Fooling SHAP with Output Shuffling Attacks (Yuan and Dasgupta, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.06509&sa=D&source=editors&ust=1779048536737522&usg=AOvVaw06BSVUmYCVxLrvcfs-Qs4F) |
| 60 | [▪ Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability (García-Carrasco, Maté, and Trujillo, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.19842&sa=D&source=editors&ust=1779048536737604&usg=AOvVaw1ob1JR_ulGwgCrXJdYKPx7) |

|     |
| --- |
| Model (Mis)Interpretability |

**>**

**<**

‍

#### Model Collapse

Covers:

- OWASP LLM 03: Training Data Poisoning
- OWASP ML 02: Data Poisoning Attack

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Collapse

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Persona-Model Collapse in Emergent Misalignment (Costa and Vicente, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12850&sa=D&source=editors&ust=1779048536033097&usg=AOvVaw13-xwBEqYLrGwl-2d0U01z) |
| 3 | [▪ Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents (Gamage, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20911&sa=D&source=editors&ust=1779048536033295&usg=AOvVaw2gknKWWIGIUmzwwgvg3qS7) |
| 4 | [▪ SpectralGuard: Detecting Memory Collapse Attacks in State Space Models (Bonetto, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12414&sa=D&source=editors&ust=1779048536033391&usg=AOvVaw2xDwUAutvZYhbJgRlO9A-9) |
| 5 | [▪ Self-Destructive Language Model (Wang, Zhu, and Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.12186&sa=D&source=editors&ust=1779048536033479&usg=AOvVaw0F_eUcTV1EIoIA9Ho-ouF2) |
| 6 | [▪ Forgetting-MarI: LLM Unlearning via Marginal Information Regularization (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11914&sa=D&source=editors&ust=1779048536033570&usg=AOvVaw0_j_jFVIulE_QSf94gNCrw) |
| 7 | [▪ LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring (Li, Phuong, and Siegel, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00943&sa=D&source=editors&ust=1779048536033658&usg=AOvVaw3QuIqn9g9ZhuIOH4P_Q5Yi) |
| 8 | [▪ Threats, Attacks, and Defenses in Machine Unlearning: A Survey (Liu et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.13682&sa=D&source=editors&ust=1779048536033758&usg=AOvVaw0jmfZiF1XWaMoHVV_jK4nb) |
| 9 | [▪ Data Duplication: A Novel Multi-Purpose Attack Paradigm in Machine Unlearning (Ye et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16663&sa=D&source=editors&ust=1779048536033845&usg=AOvVaw1-Esldx6uLcENFebu-jsSz) |
| 10 | [▪ Understanding Implosion in Text-to-Image Generative Models (Ding et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.12314&sa=D&source=editors&ust=1779048536033926&usg=AOvVaw1uFYjVrAfXbkwwC2xRe1hD) |
| 11 | [▪ AI models collapse when trained on recursively generated data (Shumailov et al, Jul 2024)](https://www.google.com/url?q=https://www.nature.com/articles/s41586-024-07566-y&sa=D&source=editors&ust=1779048536034021&usg=AOvVaw2lyMDGT9ofLnWEP9nEWN4a) |

|     |
| --- |
| Model Collapse |

**>**

**<**

‍

#### Model Denial of Service and Chaff Data Spamming

Covers:

- OWASP LLM 04: Model Denial of Service
- MITRE ATLAS Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Denial of Service and Chaff Data Spamming

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13338&sa=D&source=editors&ust=1779048535896241&usg=AOvVaw36WHc0LdDfQEROET_-rs5P) |
| 3 | [▪ AESOP: Adversarial Execution-path Selection to Overload Deep Learning Pipelines (Li et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10987&sa=D&source=editors&ust=1779048535896397&usg=AOvVaw36dwLPt1e8_CM7oxJd-BVT) |
| 4 | [▪ Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection (Liang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11442&sa=D&source=editors&ust=1779048535896510&usg=AOvVaw29vf9lafQL1iHp30McZtWc) |
| 5 | [▪ FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework (B\\'ejar, Romera-Paredes, and Hern\\'andez-Ramos, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02789&sa=D&source=editors&ust=1779048535896683&usg=AOvVaw0MT66jJFl4JFlXvtZSczaI) |
| 6 | [▪ Position Paper: Denial-of-Service against Multi-Round Transaction Simulation (Tang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21169&sa=D&source=editors&ust=1779048535896789&usg=AOvVaw1kgYp1xCsyw_7b3yMAgazy) |
| 7 | [▪ Semantic Denial of Service in LLM-controlled robots (Steinberg and Gal, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.24790&sa=D&source=editors&ust=1779048535896888&usg=AOvVaw0OnxYpnULtkz9OkuS7fOrr) |
| 8 | [▪ Position Paper: Denial-of-Service Against Multi-Round Transaction Simulation (Tang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21169&sa=D&source=editors&ust=1779048535896988&usg=AOvVaw0Xf6_C55NYxHtVPlnhPidy) |
| 9 | [▪ Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems (Yamamoto and Matsuura, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17249&sa=D&source=editors&ust=1779048535897086&usg=AOvVaw2fGPVnma_ViXYShZ3oGcW_) |
| 10 | [▪ Detecting and Mitigating DDoS Attacks with AI: A Survey (Apostu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.17867&sa=D&source=editors&ust=1779048535897182&usg=AOvVaw3wukE7zIv73_ivx6ww9TaC) |
| 11 | [▪ RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08214&sa=D&source=editors&ust=1779048535897281&usg=AOvVaw0ktNBi3keTKLj3OZCPI2qF) |
| 12 | [▪ Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07878&sa=D&source=editors&ust=1779048535897374&usg=AOvVaw1ig6waxHkw-_LAh-YYabEv) |
| 13 | [▪ OverThink: Slowdown Attacks on Reasoning LLMs (Kumar et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.02542&sa=D&source=editors&ust=1779048535897466&usg=AOvVaw3H0g19MaqqA7TU00yF9gOk) |
| 14 | [▪ ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00154&sa=D&source=editors&ust=1779048535897573&usg=AOvVaw15bzkq23xtfngQBroOYIEL) |
| 15 | [▪ Rethinking On-Device LLM Reasoning: Why Analogical Mapping Outperforms Abstract Thinking for IoT DDoS Detection (Pan et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14343&sa=D&source=editors&ust=1779048535897677&usg=AOvVaw0igEA78ikuPd0K7jJDABJI) |
| 16 | [▪ RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress (Huang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23995&sa=D&source=editors&ust=1779048535897776&usg=AOvVaw2fDRO1_UxfW5GXjCYsCQFp) |
| 17 | [▪ Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark (Manu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23779&sa=D&source=editors&ust=1779048535897871&usg=AOvVaw3ERakPr7FTrbPLOvd5Pudx) |
| 18 | [▪ ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07086&sa=D&source=editors&ust=1779048535897973&usg=AOvVaw0cfupv7eE2ERgzcRTAuHre) |
| 19 | [▪ RemedyGS: Defend 3D Gaussian Splatting against Computation Cost Attacks (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22147&sa=D&source=editors&ust=1779048535898083&usg=AOvVaw3Xb4b15uxsF1LiT4gEvZ-j) |
| 20 | [▪ LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07876&sa=D&source=editors&ust=1779048535898257&usg=AOvVaw2DHn5kYuWjdPMGJawxjEuT) |
| 21 | [▪ Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems (Yakubu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04114&sa=D&source=editors&ust=1779048535898365&usg=AOvVaw106TLOGXW9YDk4AXUVY8bF) |
| 22 | [▪ Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models (Swileh and Zhang, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00460&sa=D&source=editors&ust=1779048535898466&usg=AOvVaw3vxAqv7HS8YdAioas-W527) |
| 23 | [▪ AdaDoS: Adaptive DoS Attack via Deep Adversarial Reinforcement Learning in SDN (Shao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20566&sa=D&source=editors&ust=1779048535898564&usg=AOvVaw1_Qe5tSnbxwFYClDEqm2cJ) |
| 24 | [▪ One Token Embedding Is Enough to Deadlock Your Large Reasoning Model (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15965&sa=D&source=editors&ust=1779048535898656&usg=AOvVaw045IFnM4p_EN5SEyJzG-f2) |
| 25 | [▪ From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond (Dayaratne et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06530&sa=D&source=editors&ust=1779048535898750&usg=AOvVaw3QwMWzcMErTWfDUVuZSlEn) |
| 26 | [▪ BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models (Yan et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16670&sa=D&source=editors&ust=1779048535898838&usg=AOvVaw2xB4fT-qxsS-gpiXusM7Qw) |
| 27 | [▪ RECALLED: An Unbounded Resource Consumption Attack on Large Vision-Language Models (Gao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18053&sa=D&source=editors&ust=1779048535898935&usg=AOvVaw1xBQOZ5Ag6OHrDGM_Ge-RY) |
| 28 | [▪ Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World (Sadasivan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06256&sa=D&source=editors&ust=1779048535899031&usg=AOvVaw1s7BiJjLyGTrZwcxhfstZE) |
| 29 | [▪ Per-Row Activation Counting on Real Hardware: Demystifying Performance Overheads (Kim et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05556&sa=D&source=editors&ust=1779048535899124&usg=AOvVaw0g3XxWvcJm8Wya31t0hRcS) |
| 30 | [▪ Impact of White-Box Adversarial Attacks on Convolutional Neural Networks (Podder and Ghosh, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02043&sa=D&source=editors&ust=1779048535899245&usg=AOvVaw0qHmu3zit40w4zlGovc6wm) |
| 31 | [▪ DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models (Yin, Liu and Xu, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10561&sa=D&source=editors&ust=1779048535899334&usg=AOvVaw1VSJx7Kfo6GKbiPymbjiLO) |
| 32 | [▪ Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework (Pisano et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.00029%23&sa=D&source=editors&ust=1779048535899445&usg=AOvVaw1AeW_Ai0INUTmgAIoVf2JS) |
| 33 | [▪ Self-Evaluation as a Defense Against Adversarial Attacks on LLMs (Brown et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.03234&sa=D&source=editors&ust=1779048535899542&usg=AOvVaw3DBDYaUBsMiRQN570ckusH) |
| 34 | [▪ Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacks (Diao et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20836&sa=D&source=editors&ust=1779048535899626&usg=AOvVaw3L24CUQ1eJ3QxlD7ejze-j) |
| 35 | [▪ Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent (Waghela, Sen, and Rakshit, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.21073&sa=D&source=editors&ust=1779048535899711&usg=AOvVaw2OZh51YhqOVen1jtQPqsB3) |

|     |
| --- |
| Model Denial of Service and Chaff Data Spamming |

**>**

**<**

‍

#### Model Modifications

We added this subsection to include security issues that arise from post-hoc model modifications such as fine-tuning, quantization.

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Modifications

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing (Wang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21228&sa=D&source=editors&ust=1779048536049824&usg=AOvVaw1nkV_ftFlXmGtKSQZKxLLQ) |
| 3 | [▪ Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs (Cai et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07961&sa=D&source=editors&ust=1779048536049972&usg=AOvVaw0wXjhdG9G6-CnouICe8Vf9) |
| 4 | [▪ BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs (Karamat, Saif, and Garcia, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22174&sa=D&source=editors&ust=1779048536050054&usg=AOvVaw2FTBqQ39Mf-JW--RwWl8mH) |
| 5 | [▪ Fine-Tuning Integrity for Modern Neural Networks: Structured Drift Proofs via Norm, Rank, and Sparsity Certificates (Shang and Chen, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04738&sa=D&source=editors&ust=1779048536050129&usg=AOvVaw1Y5wrTPQKWurZst6YalfVR) |
| 6 | [▪ Locket: Robust Feature-Locking Technique for Language Models (He, Duddu, and Asokan, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12117&sa=D&source=editors&ust=1779048536050215&usg=AOvVaw3wir3FZxFq9l-CYgiRruYv) |
| 7 | [▪ CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18449&sa=D&source=editors&ust=1779048536050290&usg=AOvVaw3UyhqjmcP2mYNEdw1pD8g2) |
| 8 | [▪ CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models (Xu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13435&sa=D&source=editors&ust=1779048536050363&usg=AOvVaw19yG_Xb1kulqOBwYn9VJMG) |
| 9 | [▪ Reverse-Engineering Model Editing on Language Models (Sun et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10134&sa=D&source=editors&ust=1779048536050434&usg=AOvVaw3h0PFrqCdh_XLuFZic-B7m) |
| 10 | [▪ Making Models Unmergeable via Scaling-Sensitive Loss Landscape (Jang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21898&sa=D&source=editors&ust=1779048536050517&usg=AOvVaw3gtd0lw01cxspgFMx2BF-4) |
| 11 | [▪ FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02092&sa=D&source=editors&ust=1779048536050594&usg=AOvVaw1tVw5lWkcEYWz4DFsW0Jas) |
| 12 | [▪ Are Robust LLM Fingerprints Adversarially Robust? (Nasery et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.26598&sa=D&source=editors&ust=1779048536050686&usg=AOvVaw08f7cbaSqhjY2uBunkXc_D) |
| 13 | [▪ Empirical Evaluation of Concept Drift in ML-Based Android Malware Detection (Sabbah et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22772&sa=D&source=editors&ust=1779048536050761&usg=AOvVaw0Zx1sWWbu5Ie4WiDJW1KcG) |
| 14 | [▪ Generating Adversarial Point Clouds Using Diffusion Model (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21163&sa=D&source=editors&ust=1779048536050841&usg=AOvVaw0sprR8fpZdPrT7GaF84wb0) |
| 15 | [▪ Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16302&sa=D&source=editors&ust=1779048536050907&usg=AOvVaw3SqLP7GNxU2rf1gRaJzNnH) |
| 16 | [▪ FORTA: Byzantine-Resilient FL Aggregation via DFT-Guided Krum (Shahul and Harshan, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14588&sa=D&source=editors&ust=1779048536050971&usg=AOvVaw1hxKIdGsGbfkGIv5aIa0L-) |
| 17 | [▪ Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities (Che et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05209&sa=D&source=editors&ust=1779048536051038&usg=AOvVaw2Yz3vQX27nNS5hkGkfThGj) |
| 18 | [▪ LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents (Yan and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10610&sa=D&source=editors&ust=1779048536051103&usg=AOvVaw3WWPUE-7qb2WZxHpg9Rkpi) |
| 19 | [▪ Score Attack: A Lower Bound Technique for Optimal Differentially Private Learning (Cai, Wang, and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2303.07152&sa=D&source=editors&ust=1779048536051168&usg=AOvVaw33Lfm2pqpfmZwZiOGaYM3O) |
| 20 | [▪ Differentially Private Federated Low Rank Adaptation Beyond Fixed-Matrix (Wen et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09990&sa=D&source=editors&ust=1779048536051233&usg=AOvVaw2hOpeByYsJ5pHnIyVZNcdr) |
| 21 | [▪ Quantum Properties Trojans (QuPTs) for Attacking Quantum Neural Networks (Bhowmik, Humble, and Thapliyal, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08202&sa=D&source=editors&ust=1779048536051297&usg=AOvVaw0-vuGwNgWzUc-PWm0vj_Km) |
| 22 | [▪ Adaptive Diffusion Denoised Smoothing : Certified Robustness via Randomized Smoothing with Differentially Private Guided Denoising Diffusion (Shpilevskiy et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08163&sa=D&source=editors&ust=1779048536051365&usg=AOvVaw3OpqUtW8JpaLq3MIIOsR9B) |
| 23 | [▪ Decomposition-Based Optimal Bounds for Privacy Amplification via Shuffling (Su, Cheng, and Wang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.07414&sa=D&source=editors&ust=1779048536051430&usg=AOvVaw273ypla4eP71vDYxCY_JkT) |
| 24 | [▪ Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05512&sa=D&source=editors&ust=1779048536051501&usg=AOvVaw0oF3t8YDqlC0HhEv866jQ2) |
| 25 | [▪ Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting\<br>\<br> (To and Le, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17160&sa=D&source=editors&ust=1779048536051574&usg=AOvVaw33FGj3K-pGCrjAvpWeLQ-y) |
| 26 | [▪ Finetuning-Activated Backdoors in LLMs (Gloaguen, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16567&sa=D&source=editors&ust=1779048536051653&usg=AOvVaw0aUE9EoVnR33LV6aROBWbe) |
| 27 | [▪ The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning (Xiang et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.04106&sa=D&source=editors&ust=1779048536051727&usg=AOvVaw2mknDmPVHHfLjylHO_Sn3B) |
| 28 | [▪ Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey (Huang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.18169&sa=D&source=editors&ust=1779048536051802&usg=AOvVaw2OqeJfQQtZaQz1cHQL-bz0) |
| 29 | [▪ Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks (Poppi et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.18210%23&sa=D&source=editors&ust=1779048536051876&usg=AOvVaw356IsKTg_z2WCpSOtTPGdS) |
| 30 | [▪ Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack (Chen et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00960&sa=D&source=editors&ust=1779048536051961&usg=AOvVaw3LcZ6jadLsEr4U4ygLiCoo) |
| 31 | [▪ The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs (Chen et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.00787&sa=D&source=editors&ust=1779048536052043&usg=AOvVaw11hpXuELpqhbB0RbBAJj3E) |
| 32 | [▪ BadMerging: Backdoor Attacks Against Model Merging (Zhang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.07362&sa=D&source=editors&ust=1779048536052121&usg=AOvVaw2kS58EdB3z27jff6CDZtfi) |
| 33 | [▪ Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning (Huang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09600&sa=D&source=editors&ust=1779048536052199&usg=AOvVaw16G5ZGsmje31FyZe-3-n_6) |
| 34 | [▪ RLCP: A Reinforcement Learning-based Copyright Protection Method for Text-to-Image Diffusion Model (Shi et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.16634&sa=D&source=editors&ust=1779048536052282&usg=AOvVaw12-wZMtihq_3FcHUMTwbpf) |
| 35 | [▪ Large Language Models as Carriers of Hidden Messages (Hoscilowicz et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.02481&sa=D&source=editors&ust=1779048536052347&usg=AOvVaw3dPLW7ITp3mvsllh-1juFH) |
| 36 | [▪ Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence (Chen et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2112.13060&sa=D&source=editors&ust=1779048536052437&usg=AOvVaw0-ArSPgJGIk9W-Xvi28Ro4) |
| 37 | [▪ Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning (Huang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.09600%23&sa=D&source=editors&ust=1779048536052515&usg=AOvVaw1cpIhU7sLzW1PAFHgcVNjf) |
| 38 | [▪ Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data (Baumgärtner et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.05530&sa=D&source=editors&ust=1779048536052586&usg=AOvVaw1Ir0G_8LqXB5ep-fbbE52w) |
| 39 | [▪ Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes (Kumar et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.04392&sa=D&source=editors&ust=1779048536052660&usg=AOvVaw30A0ybKpzJUi1aCEYJuxcG) |
| 40 | [▪ DeepBaR: Fault Backdoor Attack on Deep Neural Network Layers (Martínez-Mejía et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.21220&sa=D&source=editors&ust=1779048536052725&usg=AOvVaw0iHngxzIbDM_5wEBAFu0YD) |
| 41 | [▪ Resilience and Security of Deep Neural Networks Against Intentional and Unintentional Perturbations: Survey and Research Challenges (Sayyed et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.00193&sa=D&source=editors&ust=1779048536052793&usg=AOvVaw1W6frWRqShtHF5sXlocS-T) |

|     |
| --- |
| Model Modifications |

**>**

**<**

‍

#### Inadequate AI Alignment

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Inadequate AI Alignment

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Ghost in the Context: Measuring Policy-Carriage Failures in Decision-Time Assembly (Santos-Grueiro, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12535&sa=D&source=editors&ust=1779048535924127&usg=AOvVaw2UVXdMxf0WNcji910lEOcC) |
| 3 | [▪ No More, No Less: Task Alignment in Terminal Agents (Mavali et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12233&sa=D&source=editors&ust=1779048535924426&usg=AOvVaw2_Q9tay33_O-gLn9a_U5Q5) |
| 4 | [▪ Leveraging RAG for Training-Free Alignment of LLMs (Halloran, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11217&sa=D&source=editors&ust=1779048535924581&usg=AOvVaw1333ypEyKb0w7y9zgyvLuu) |
| 5 | [▪ You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation (Arazzi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04992&sa=D&source=editors&ust=1779048535924703&usg=AOvVaw0XIM3MyVj3-NO6KF_b7459) |
| 6 | [▪ When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models (Hossain et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02914&sa=D&source=editors&ust=1779048535924816&usg=AOvVaw3lLrjqQgZ2li0UeATEMPCM) |
| 7 | [▪ Dependency-Aware Privacy for Multi-turn Agents (Anshumaan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03188&sa=D&source=editors&ust=1779048535924904&usg=AOvVaw3qsb3giuwwfbeqaEiTzhjX) |
| 8 | [▪ Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness (Li and Liu, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.24056&sa=D&source=editors&ust=1779048535924991&usg=AOvVaw1_QgD6u6mTPzD0vtIOoN5W) |
| 9 | [▪ Tatemae: Detecting Alignment Faking via Tool Selection in LLMs (Leonesi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26511&sa=D&source=editors&ust=1779048535925076&usg=AOvVaw36gTVodzctrxCT5UZdsxE8) |
| 10 | [▪ Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers (Dubi\\'nski et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25891&sa=D&source=editors&ust=1779048535925164&usg=AOvVaw2mXJSGy-81F30xz44Hwa90) |
| 11 | [▪ Beyond Context: Large Language Models' Failure to Grasp Users' Intent (Hussain and Salahuddin, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21110&sa=D&source=editors&ust=1779048535925282&usg=AOvVaw0DHpDG8At86-xMIn_zeFjp) |
| 12 | [▪ Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems (Wu and Gong, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22879&sa=D&source=editors&ust=1779048535925379&usg=AOvVaw1MwmxGi-obriuKGJS8DWhF) |
| 13 | [▪ SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs (Pan, Wu, and Yao, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20930&sa=D&source=editors&ust=1779048535925465&usg=AOvVaw3XyMX4t2ZXDylD-vmmX2uw) |
| 14 | [▪ Sensitivity Uncertainty Alignment in Large Language Models (Hiremath and Hiremath, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20903&sa=D&source=editors&ust=1779048535925542&usg=AOvVaw32xRJMp4znMjbYZfXiTW5P) |
| 15 | [▪ Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models (Liu, Tang, and Tun, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.10294&sa=D&source=editors&ust=1779048535925630&usg=AOvVaw1QWEXnKI2gxs9KxvG_xtyF) |
| 16 | [▪ Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning (Hui et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16054&sa=D&source=editors&ust=1779048535925724&usg=AOvVaw330-S6eAuE6ZyX3KmNkRvR) |
| 17 | [▪ Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2405.13068&sa=D&source=editors&ust=1779048535925813&usg=AOvVaw2UZrHqMBJttH8X_XFkseiy) |
| 18 | [▪ Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs (Roh and Houmansadr, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16659&sa=D&source=editors&ust=1779048535925894&usg=AOvVaw0xSaRfb9iSoFx13AnNFIJH) |
| 19 | [▪ Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries (Hung et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15717&sa=D&source=editors&ust=1779048535925999&usg=AOvVaw2yYoi7527OAPM13DjDNm9b) |
| 20 | [▪ Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design (Eshuijs, Wang, and Fokkens, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12500&sa=D&source=editors&ust=1779048535926103&usg=AOvVaw2HNICXBmyKxfML8JEmZwHo) |
| 21 | [▪ The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07754&sa=D&source=editors&ust=1779048535926199&usg=AOvVaw1BF7aqN_1daVRTiATNY1nn) |
| 22 | [▪ FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment (Kuznetsov et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04992&sa=D&source=editors&ust=1779048535926312&usg=AOvVaw2oDhXWWnKDujz0I4hVHlxK) |
| 23 | [▪ Understanding the Effects of Safety Unalignment on Large Language Models (Halloran, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02574&sa=D&source=editors&ust=1779048535926399&usg=AOvVaw2rVrubO7iQQzSO9Ckw-QH2) |
| 24 | [▪ PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.18649&sa=D&source=editors&ust=1779048535926478&usg=AOvVaw1prhOD8ZHuHFz3YGT2yKPt) |
| 25 | [▪ Safety, Security, and Cognitive Risks in World Models (Parmar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01346&sa=D&source=editors&ust=1779048535926561&usg=AOvVaw0bDBCHxHjOM5wKaVFeybOo) |
| 26 | [▪ UK AISI Alignment Evaluation Case-Study (Souly et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00788&sa=D&source=editors&ust=1779048535926652&usg=AOvVaw2mDK4SCn5PCJXGno3lr6gJ) |
| 27 | [▪ Security in LLM-as-a-Judge: A Comprehensive SoK (Almasoud et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29403&sa=D&source=editors&ust=1779048535926743&usg=AOvVaw3WTv7HixfD9ohEcVV6RJIH) |
| 28 | [▪ Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment (Ding, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12681&sa=D&source=editors&ust=1779048535926847&usg=AOvVaw3Ld_djgsHyvZAFZ62LXA00) |
| 29 | [▪ SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do (Dhodapkar and Pishori, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27148&sa=D&source=editors&ust=1779048535926965&usg=AOvVaw0CjsQ15QDCYYl73jcdVE8f) |
| 30 | [▪ Internal Safety Collapse in Frontier Large Language Models (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23509&sa=D&source=editors&ust=1779048535927055&usg=AOvVaw1c8NLjwxIewCgtFpSFn5_g) |
| 31 | [▪ Silent Commitment Failure in Instruction-Tuned Language Models: Evidence of Governability Divergence Across Architectures (Ruddell, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21415&sa=D&source=editors&ust=1779048535927158&usg=AOvVaw1vWr_nkU__B9Y51mVG04_0) |
| 32 | [▪ Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models (Tan, Hu, and Lee, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21697&sa=D&source=editors&ust=1779048535927276&usg=AOvVaw353cPoUoz04PM2hhW3GMET) |
| 33 | [▪ MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment (Dou and Yang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18637&sa=D&source=editors&ust=1779048535927361&usg=AOvVaw11ik2xqe3YZ1X3viiQVIwy) |
| 34 | [▪ Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMs (Xia et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2407.16576&sa=D&source=editors&ust=1779048535927453&usg=AOvVaw1HJ_Wv2WjTnZxHTESynNls) |
| 35 | [▪ State-Dependent Safety Failures in Multi-Turn Language Model Interaction (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15684&sa=D&source=editors&ust=1779048535927543&usg=AOvVaw3xWyAGc4KnuRWA15kGWLwR) |
| 36 | [▪ Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks (Chu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11914&sa=D&source=editors&ust=1779048535927639&usg=AOvVaw1qI7H-V-ZbEV1hzOdHMX_J) |
| 37 | [▪ Proof-of-Guardrail in AI Agents and What (Not) to Trust from It (Jin et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05786&sa=D&source=editors&ust=1779048535927725&usg=AOvVaw0SeoV1dlzoEcO4fvqRUnOT) |
| 38 | [▪ When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG (Li et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03919&sa=D&source=editors&ust=1779048535927827&usg=AOvVaw0uxY506ROZlTL-FO0WpU17) |
| 39 | [▪ Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders (Campbell et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01246&sa=D&source=editors&ust=1779048535927931&usg=AOvVaw2xqnJyZMtPLhruIKK0oDZG) |
| 40 | [▪ Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems (Gao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12462&sa=D&source=editors&ust=1779048535928049&usg=AOvVaw3cgBhTQh6sx-hv82XqnCE4) |
| 41 | [▪ Fail-Closed Alignment for Large Language Models (Coalson et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16977&sa=D&source=editors&ust=1779048535928146&usg=AOvVaw3xPMKwYBF2_jeJFWN5Tzuu) |
| 42 | [▪ A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models (Segal et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15689&sa=D&source=editors&ust=1779048535928258&usg=AOvVaw0pdbEhyy0CKWo3RgKNnxv2) |
| 43 | [▪ Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13562&sa=D&source=editors&ust=1779048535928351&usg=AOvVaw1VGHU4370sBE59aH9t9Xgm) |
| 44 | [▪ CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis (Eswaran et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11304&sa=D&source=editors&ust=1779048535928455&usg=AOvVaw02F_LXjEp1VC-vMF1sbLOl) |
| 45 | [▪ Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10161&sa=D&source=editors&ust=1779048535928548&usg=AOvVaw2Ednf_eouoseDh0_cYv-6C) |
| 46 | [▪ Is Reasoning Capability Enough for Safety in Long-Context Language Models? (Fu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08874&sa=D&source=editors&ust=1779048535928636&usg=AOvVaw2eVefpZvsouJpwVtS5bJqC) |
| 47 | [▪ When Evaluation Becomes a Side Channel: Regime Leakage and Structural Mitigations for Alignment Assessment (Santos-Grueiro, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08449&sa=D&source=editors&ust=1779048535928721&usg=AOvVaw31l-2Pwc2DTJub46PnPTvL) |
| 48 | [▪ When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents (Jones et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08235&sa=D&source=editors&ust=1779048535928808&usg=AOvVaw3c6NjIXf-ow_FST47qNk0l) |
| 49 | [▪ RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models (Liang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04448&sa=D&source=editors&ust=1779048535928887&usg=AOvVaw27OxFWudydi07tUqtzq9Uf) |
| 50 | [▪ Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01600&sa=D&source=editors&ust=1779048535928962&usg=AOvVaw1QVB7ZvGYJY52wTpzvqVMy) |
| 51 | [▪ Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures (Su et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.23081&sa=D&source=editors&ust=1779048535929047&usg=AOvVaw0538WRuHoViGf5utIhUjKu) |
| 52 | [▪ SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment (Fang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16506&sa=D&source=editors&ust=1779048535929126&usg=AOvVaw09orjYx4v5Sl-JTEbuPpYv) |
| 53 | [▪ LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions (Hu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08211&sa=D&source=editors&ust=1779048535929207&usg=AOvVaw2l-WVWaUpYwUdXKkoKi4L4) |
| 54 | [▪ Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs (Chua et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05367&sa=D&source=editors&ust=1779048535929292&usg=AOvVaw3kQLsTnVfQyfVyoij2friV) |
| 55 | [▪ Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity (An et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06596&sa=D&source=editors&ust=1779048535929374&usg=AOvVaw1PWlwUoJIZR_s97HR08aAP) |
| 56 | [▪ Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06419&sa=D&source=editors&ust=1779048535929452&usg=AOvVaw0jQ7uKo7XGEIggoSGLeirR) |
| 57 | [▪ What Matters For Safety Alignment? (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03868&sa=D&source=editors&ust=1779048535929544&usg=AOvVaw0nVIT-fKOhpXfEoQFGhga4) |
| 58 | [▪ Beyond Context: Large Language Models Failure to Grasp Users Intent (Hussain, Salahuddin, and Papadimitratos, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21110&sa=D&source=editors&ust=1779048535929628&usg=AOvVaw2cxnoV2Yc3ZZ7EHBiD_bXm) |
| 59 | [▪ Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance (Ludvigsen, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16419&sa=D&source=editors&ust=1779048535929718&usg=AOvVaw2Z1LoilGvCR01WE8mU22Xx) |
| 60 | [▪ PROPS: Progressively Private Self-alignment of Large Language Models (Teku et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06783&sa=D&source=editors&ust=1779048535929809&usg=AOvVaw28NKP_DzqxiOe7HaRzr_Oc) |
| 61 | [▪ Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents (Liang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06716&sa=D&source=editors&ust=1779048535929902&usg=AOvVaw2Nu1dy4gckySFnfLnD-8wq) |
| 62 | [▪ Matching Ranks Over Probability Yields Truly Deep Safety Alignment (Vega and Singh, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.05518&sa=D&source=editors&ust=1779048535929995&usg=AOvVaw1okiZE596iGMh14uobc0T9) |
| 63 | [▪ Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position (Xie, Song, and Luo, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.12398&sa=D&source=editors&ust=1779048535930083&usg=AOvVaw1iTJFsOdx1zlRvs8oXtjSj) |
| 64 | [▪ Can LLMs Make (Personalized) Access Control Decisions? (Groschupp et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20284&sa=D&source=editors&ust=1779048535930172&usg=AOvVaw31QrfixBk7Z7lTG6Cc_gzN) |
| 65 | [▪ Can LLMs Threaten Human Survival? Benchmarking Potential Existential Threats from LLMs via Prefix Completion (Cui et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19171&sa=D&source=editors&ust=1779048535930272&usg=AOvVaw2NU7QtwrQ0TwMJY2s1W446) |
| 66 | [▪ Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19009&sa=D&source=editors&ust=1779048535930354&usg=AOvVaw1EUJAdw0Rmp_Nps2V1DsUg) |
| 67 | [▪ EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models (Wu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09880&sa=D&source=editors&ust=1779048535930434&usg=AOvVaw3cPCqSsYVXCi5g6ZtLwkiC) |
| 68 | [▪ A Self-Improving Architecture for Dynamic Safety in Large Language Models (Slater, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07645&sa=D&source=editors&ust=1779048535930511&usg=AOvVaw0crANDZBoQVQblwXTU_Ffe) |
| 69 | [▪ HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection (Dai, Jiang, and Deng, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06942&sa=D&source=editors&ust=1779048535930591&usg=AOvVaw0Izy7P4blbJzT8adHCeoxk) |
| 70 | [▪ EASE: Practical and Efficient Safety Alignment for Small Language Models (Shi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06512&sa=D&source=editors&ust=1779048535930681&usg=AOvVaw1xtMsN8myWFhSd4dhj-_d2) |
| 71 | [▪ XBreaking: Understanding how LLMs security alignment can be broken (Arazzi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21700&sa=D&source=editors&ust=1779048535930778&usg=AOvVaw3fghVSNA5sZRJV3xW3NzbB) |
| 72 | [▪ Reimagining Safety Alignment with An Image (Xia et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00509&sa=D&source=editors&ust=1779048535930857&usg=AOvVaw2KVCC7eIkWj0r_IdemUwE2) |
| 73 | [▪ Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures (Schwarz, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.27190&sa=D&source=editors&ust=1779048535930947&usg=AOvVaw2V9maUNQRL_mmw9H8hDyTP) |
| 74 | [▪ When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? (Peng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17862&sa=D&source=editors&ust=1779048535931033&usg=AOvVaw2226h57ojbd3bxrwZoSVE6) |
| 75 | [▪ HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15499&sa=D&source=editors&ust=1779048535931126&usg=AOvVaw3jkNOyOKDw-lwZYjBq8Cwj) |
| 76 | [▪ Cross-Modal Safety Alignment: Is textual unlearning all you need? (Chakraborty et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.02575&sa=D&source=editors&ust=1779048535931207&usg=AOvVaw3g9RTXJX3I7qs5DnjVzSzC) |
| 77 | [▪ A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment (Meng and Zhang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09615&sa=D&source=editors&ust=1779048535931317&usg=AOvVaw0v5zb5pVc1JGKTSLWSB1Av) |
| 78 | [▪ LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions (Hu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08211&sa=D&source=editors&ust=1779048535931427&usg=AOvVaw37Dq42iOQROEKqbc1s4sOv) |
| 79 | [▪ Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning? (Yin et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06036&sa=D&source=editors&ust=1779048535931525&usg=AOvVaw2t5_l3eEO3Zx_ec3jvrVwL) |
| 80 | [▪ Superficial Safety Alignment Hypothesis (Li and Kim, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.10862&sa=D&source=editors&ust=1779048535931607&usg=AOvVaw0Mpt6j3_G4pATmBY6YwksD) |
| 81 | [▪ AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14987&sa=D&source=editors&ust=1779048535931692&usg=AOvVaw3jiy5HAtE8xjg1kQ8XSZgj) |
| 82 | [▪ Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework (Dassanayake et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12872&sa=D&source=editors&ust=1779048535931773&usg=AOvVaw15ucc2G7qgH8AzAnN7EViT) |
| 83 | [▪ ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning (Zhao et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11500&sa=D&source=editors&ust=1779048535931851&usg=AOvVaw2VWQ42fyy9JBPohfygtp1h) |
| 84 | [▪ Agent Safety Alignment via Reinforcement Learning (Sha et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08270&sa=D&source=editors&ust=1779048535931938&usg=AOvVaw2qt1BzkQQek_0gxryo8B3K) |
| 85 | [▪ On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment (Ball et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07341&sa=D&source=editors&ust=1779048535932025&usg=AOvVaw2RfPCBzImN5vigHIslV--5) |
| 86 | [▪ Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework (Krishna et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06260&sa=D&source=editors&ust=1779048535932118&usg=AOvVaw2_5v9oFuiQD80ZBfcFSwaV) |
| 87 | [▪ Emergent misalignment as prompt sensitivity: A research note (Wyse et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06253&sa=D&source=editors&ust=1779048535932208&usg=AOvVaw1pki1jxJ6h3DZwnZFCX0hl) |

|     |
| --- |
| Inadequate AI Alignment |

**>**

**<**

‍

‍

#### Discover ML Model Family and Ontology/Model Extraction

Added model extraction to this as it did not have its own category.

Covers:

- MITRE ATLAS Discovery

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Discover ML Model Family and Ontology/Model Extraction

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies (Cotti et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00081&sa=D&source=editors&ust=1779048535945145&usg=AOvVaw3oAH4E9wNn3iyy5VfNCF5j) |
| 3 | [▪ Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs (Liu, Evans, and Xiong, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18697&sa=D&source=editors&ust=1779048535945366&usg=AOvVaw2L2332NuEtWIO62doCdrgb) |
| 4 | [▪ CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models (Lee, Koo, and Kwak, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16363&sa=D&source=editors&ust=1779048535945494&usg=AOvVaw27Ks9mO5zcQ_9BR_C-dEIE) |
| 5 | [▪ AttnDiff: Attention-based Differential Fingerprinting for Large Language Models (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05502&sa=D&source=editors&ust=1779048535945592&usg=AOvVaw2mQropBSGFoT0oozkPx0Rm) |
| 6 | [▪ Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test (Zhu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06975&sa=D&source=editors&ust=1779048535945676&usg=AOvVaw0x-93Sv0AEugUbPeaiFTTD) |
| 7 | [▪ Navigating the Deep: End-to-End Extraction on Deep Neural Networks (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17047&sa=D&source=editors&ust=1779048535945757&usg=AOvVaw0nY6R3q1qilGIGfnSs3zOR) |
| 8 | [▪ A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors (Xu and Sheng, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09434&sa=D&source=editors&ust=1779048535945862&usg=AOvVaw2krwe33mSoOFCP_hj5lJMA) |
| 9 | [▪ FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting (Fu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16029&sa=D&source=editors&ust=1779048535945952&usg=AOvVaw12mx2twyjEnepJatY8WWm0) |
| 10 | [▪ Identifying Models Behind Text-to-Image Leaderboards (Naseh et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09647&sa=D&source=editors&ust=1779048535946049&usg=AOvVaw0x63r-IF90pvyqOolkJzKG) |
| 11 | [▪ Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06390&sa=D&source=editors&ust=1779048535946158&usg=AOvVaw1LPeEgtQ2Ji1M_ZRrHJRJq) |
| 12 | [▪ A Fingerprint for Large Language Models (Yang and Wu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.01235&sa=D&source=editors&ust=1779048535946249&usg=AOvVaw3PPaCeYcS1UBo0ZZT_BKxy) |
| 13 | [▪ SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting (Zhang and Zheng, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03620&sa=D&source=editors&ust=1779048535946343&usg=AOvVaw0CYgF8l2t2ROsOcAWLvrwo) |
| 14 | [▪ A Systematic Study of Model Extraction Attacks on Graph Foundation Models (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11912&sa=D&source=editors&ust=1779048535946434&usg=AOvVaw251N0o41ISF9C0mnSUp73k) |
| 15 | [▪ Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution Approach (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07033&sa=D&source=editors&ust=1779048535946527&usg=AOvVaw1tL-pBOjOChiMCQRX6HN4j) |
| 16 | [▪ Ghost in the Transformer: Tracing LLM Lineage with SVD-Fingerprint (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06390&sa=D&source=editors&ust=1779048535946609&usg=AOvVaw38NfMavpCWghOtf879EFzZ) |
| 17 | [▪ Structuring Security: A Survey of Cybersecurity Ontologies, Semantic Log Processing, and LLMs Application (Louren\\c{c}o et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16610&sa=D&source=editors&ust=1779048535946694&usg=AOvVaw2S-zEeVa3WBELATTqegzZJ) |
| 18 | [▪ MalCVE: Malware Detection and CVE Association Using Large Language Models (Cristea, Molnes, and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15567&sa=D&source=editors&ust=1779048535946778&usg=AOvVaw1uulffoaRMmpleEx0-9ifZ) |
| 19 | [▪ Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation (Shao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06605&sa=D&source=editors&ust=1779048535946863&usg=AOvVaw0sHr4_OFpv6-46aOsErwbH) |
| 20 | [▪ SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From (Tong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.26404&sa=D&source=editors&ust=1779048535946945&usg=AOvVaw2CRIEvVFRBM4JI6mjuIGvJ) |
| 21 | [▪ LLM-Assisted Model-Based Fuzzing of Protocol Implementations (Huang, Wang, and Zhou, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01750&sa=D&source=editors&ust=1779048535947048&usg=AOvVaw2dESkDveUty9WvDFiZAyKM) |
| 22 | [▪ PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models (Gasmi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19185&sa=D&source=editors&ust=1779048535947137&usg=AOvVaw1N-K_4rQof-oRzNqcgSW0p) |
| 23 | [▪ Evaluating Ensemble and Deep Learning Models for Static Malware Detection with Dimensionality Reduction Using the EMBER Dataset (Abedin and Mehrub, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16952&sa=D&source=editors&ust=1779048535947224&usg=AOvVaw3syw-MBVZ9gZoHqbv_fI-P) |
| 24 | [▪ Revisiting Pre-trained Language Models for Vulnerability Detection (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16887&sa=D&source=editors&ust=1779048535947304&usg=AOvVaw193KcbZ-LPvU7QJHiYJJPO) |
| 25 | [▪ Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications (H\\"arer, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10467&sa=D&source=editors&ust=1779048535947388&usg=AOvVaw1uZ5rG13y_iGcVxKnVhUqi) |
| 26 | [▪ TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2304.02838&sa=D&source=editors&ust=1779048535947488&usg=AOvVaw3tDsN6coPzQIMOFO33EvVQ) |
| 27 | [▪ Toward an Intent-Based and Ontology-Driven Autonomic Security Response in Security Orchestration Automation and Response (Huang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12061&sa=D&source=editors&ust=1779048535947607&usg=AOvVaw3v85DF-IajI2oamaM8z43c) |
| 28 | [▪ SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents (Masoor, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10562&sa=D&source=editors&ust=1779048535947701&usg=AOvVaw3lSDLxJoXqLtBsuzlXeau0) |
| 29 | [▪ UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification (Cai et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.12318&sa=D&source=editors&ust=1779048535947792&usg=AOvVaw2I1PV43hiOW0vGrpIs14j-) |
| 30 | [▪ AICrypto: A Comprehensive Benchmark For Evaluating Cryptography Capabilities of Large Language Models (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09580&sa=D&source=editors&ust=1779048535947886&usg=AOvVaw0i2TMkaDoXGsaboY3ttZcV) |
| 31 | [▪ BISON: Blind Identification with Stateless scOped pseudoNyms (Heher, More, and Heimberger, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.01518&sa=D&source=editors&ust=1779048535947978&usg=AOvVaw37R6e-NLHb_HWLfiI6hxrg) |
| 32 | [▪ Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms (Gasmi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06323&sa=D&source=editors&ust=1779048535948089&usg=AOvVaw0fv4GOBx9McCF13BV8IoJ-) |
| 33 | [▪ From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks (Khouna, Ferry, and Vidal, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05325&sa=D&source=editors&ust=1779048535948173&usg=AOvVaw1ZGlctSHGemVV_CaoWwOln) |
| 34 | [▪ One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs\<br>\<br> (Li et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17598&sa=D&source=editors&ust=1779048535948260&usg=AOvVaw0zm8pnbwc3tIF-O6JYIhwJ) |
| 35 | [▪ Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries (Wang et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15420&sa=D&source=editors&ust=1779048535948342&usg=AOvVaw2fOyhqjRyHx-O8cfQyBxnI) |
| 36 | [▪ How to Backdoor the Knowledge Distillation (Wu et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21323&sa=D&source=editors&ust=1779048535948420&usg=AOvVaw2BukCEShlHTsnzFVmpgpkV) |
| 37 | [▪ Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations (Ezzeddine, Ayoub, and Giordano, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21323&sa=D&source=editors&ust=1779048535948505&usg=AOvVaw0HsJvbVld1_WY8mXvaB46n) |
| 38 | [▪ Efficient and Effective Model Extraction (Zhu et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14122&sa=D&source=editors&ust=1779048535948583&usg=AOvVaw3XPtWnOCLB5nCKF7fLQELU) |
| 39 | [▪ CaBaGe: Data-Free Model Extraction using ClAss BAlanced Generator Ensemble (Rosenthal et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10643&sa=D&source=editors&ust=1779048535948666&usg=AOvVaw32QBbXsoqe-nzsCq6gIN0D) |
| 40 | [▪ Alignment-Aware Model Extraction Attacks on Large Language Models (Liang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.02718%23&sa=D&source=editors&ust=1779048535948745&usg=AOvVaw3smbLVyq8kmhZHx5navjR3) |

|     |
| --- |
| Discover ML Model Family and Ontology/Model Extraction |

**>**

**<**

‍

#### Improper Error Handling

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Improper Error Handling

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |

|     |
| --- |
| Improper Error Handling |

**>**

**<**

‍

#### Robust Multi-Prompt and Multi-Model Attacks

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Robust Multi-Prompt and Multi-Model Attacks

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting (Chin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.16769&sa=D&source=editors&ust=1779048536614169&usg=AOvVaw2t8Mbvead8Zi4605B3D9ds) |
| 3 | [▪ When LLMs Team Up: A Coordinated Attack Framework for Automated Cyber Intrusions (Qi et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08763&sa=D&source=editors&ust=1779048536614273&usg=AOvVaw13js4wr74dNFdDhsrkICiI) |
| 4 | [▪ Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation (Liang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04129&sa=D&source=editors&ust=1779048536614357&usg=AOvVaw1nyxf-20359Lw9ytC4Jibl) |
| 5 | [▪ Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines (Ahad et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08608&sa=D&source=editors&ust=1779048536614450&usg=AOvVaw1-NRABWAHX-pxS3BclqQ4r) |
| 6 | [▪ TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07223&sa=D&source=editors&ust=1779048536614533&usg=AOvVaw3N_jNuCmH6HVH---uMpVzj) |
| 7 | [▪ When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.00627&sa=D&source=editors&ust=1779048536614611&usg=AOvVaw1D2thzmIqXbSVUqx2SpTcC) |
| 8 | [▪ ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents (Wu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12614&sa=D&source=editors&ust=1779048536614716&usg=AOvVaw2vMO5AM130xGQo0lFVKsnE) |
| 9 | [▪ Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference (Yang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10091&sa=D&source=editors&ust=1779048536614809&usg=AOvVaw143sO-Dh0MhaeA3e5CpLCA) |
| 10 | [▪ OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01592&sa=D&source=editors&ust=1779048536614888&usg=AOvVaw2oAZ8K1V_TlB0hNXpiojfr) |
| 11 | [▪ Tipping the Dominos: Topology-Aware Multi-Hop Attacks on LLM-Based Multi-Agent Systems (Liang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04129&sa=D&source=editors&ust=1779048536614988&usg=AOvVaw3tU3ta2aUFlkI9BToyC08I) |
| 12 | [▪ Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16110&sa=D&source=editors&ust=1779048536615064&usg=AOvVaw1nKjh9YWNsU1UkfQwMbFSg) |
| 13 | [▪ RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework (Ikbarieh, Aryal, and Gupta, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06212&sa=D&source=editors&ust=1779048536615136&usg=AOvVaw3TyWJzoYiDDMa8woAlWQad) |
| 14 | [▪ Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks (Chang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01746&sa=D&source=editors&ust=1779048536615258&usg=AOvVaw2LMZRcXB_IC3xh-oGKyNFH) |
| 15 | [▪ MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15994&sa=D&source=editors&ust=1779048536615387&usg=AOvVaw2LfklJjALkfdeRXZoeReE0) |
| 16 | [▪ Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment (Shahid et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.05157&sa=D&source=editors&ust=1779048536615528&usg=AOvVaw0bUAWoRht4mu7mvUbSC5VR) |

|     |
| --- |
| Robust Multi-Prompt and Multi-Model Attacks |

**>**

**<**

‍

#### Multi-Modal Attacks

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Multi-Modal Attacks

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Adversarial Hubness in Multi-Modal Retrieval (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2412.14113&sa=D&source=editors&ust=1779048536018412&usg=AOvVaw3t1KvU3Gh3OuuYQuJJZQRY) |
| 3 | [▪ Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving (Ju et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12743&sa=D&source=editors&ust=1779048536018678&usg=AOvVaw0i3YjQHhgbpevJWk30O_ev) |
| 4 | [▪ FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence (Yan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08820&sa=D&source=editors&ust=1779048536018807&usg=AOvVaw0d6GpsyllTr4O7cHwWtXvj) |
| 5 | [▪ STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack (Mao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00699&sa=D&source=editors&ust=1779048536018912&usg=AOvVaw2ATzY66qjAz36yXkFelG-0) |
| 6 | [▪ Cross-Modal Phantom: Coordinated Camera-LiDAR Spoofing Against Multi-Sensor Fusion in Autonomous Vehicles (Khan and Hasan, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21841&sa=D&source=editors&ust=1779048536019010&usg=AOvVaw0IfznEUJIWcAYUQAJ8trb8) |
| 7 | [▪ SoK: The Next Frontier in AV Security: Systematizing Perception Attacks and the Emerging Threat of Multi-Sensor Fusion (Khan, Islam, and Hasan, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20621&sa=D&source=editors&ust=1779048536019118&usg=AOvVaw1Ms7p2hvB3ps7g6-KH3VRM) |
| 8 | [▪ Text Steganography with Dynamic Codebook and Multimodal Large Language Model (Gao, Lei, and Peng, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20269&sa=D&source=editors&ust=1779048536019213&usg=AOvVaw2tvonFNEmWCaEeYHdZWZnu) |
| 9 | [▪ ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19083&sa=D&source=editors&ust=1779048536019300&usg=AOvVaw2PA7TWNBo8OJ_o06112Nfh) |
| 10 | [▪ Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.08140&sa=D&source=editors&ust=1779048536019388&usg=AOvVaw3H7Ya5g68EVziETRMWdTvp) |
| 11 | [▪ From Incomplete Architecture to Quantified Risk: Multimodal LLM-Driven Security Assessment for Cyber-Physical Systems (Huang, Poskitt, and Shar, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05674&sa=D&source=editors&ust=1779048536019483&usg=AOvVaw0ZSNLrtRy1Micxo-c6mCg3) |
| 12 | [▪ See No Evil: Adversarial Attacks Against Linguistic-Visual Association in Referring Multi-Object Tracking Systems (Bouzidi, Liu, and Faruque, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.02028&sa=D&source=editors&ust=1779048536019585&usg=AOvVaw2J723ct9DwNpmZv3ZsDVMp) |
| 13 | [▪ Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey (Jain, Ar{\\i}k, and Thakur, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.27918&sa=D&source=editors&ust=1779048536019681&usg=AOvVaw3b4utXoaBodYGlDNFkk9RD) |
| 14 | [▪ Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.20181&sa=D&source=editors&ust=1779048536019769&usg=AOvVaw3Ev_zgoRvQmjFRBzVZga1i) |
| 15 | [▪ REFORGE: Multi-modal Attacks Reveal Vulnerable Concept Unlearning in Image Generation Models (Zou et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.16576&sa=D&source=editors&ust=1779048536019860&usg=AOvVaw3imkads3FwJHaGdZ5PmyOu) |
| 16 | [▪ Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution (Shi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01784&sa=D&source=editors&ust=1779048536019947&usg=AOvVaw1Uuc056bwG05iZqGL2hNAd) |
| 17 | [▪ MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21394&sa=D&source=editors&ust=1779048536020035&usg=AOvVaw0u3EkMNisd4eyiS66W8-bn) |
| 18 | [▪ Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars (Tehrani et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18079&sa=D&source=editors&ust=1779048536020122&usg=AOvVaw3K5B9indYIZZ3mYH6gi7mG) |
| 19 | [▪ SecureScan: An AI-Driven Multi-Layer Framework for Malware and Phishing Detection Using Logistic Regression and Threat Intelligence Integration (Firdos and Dangi, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10750&sa=D&source=editors&ust=1779048536020220&usg=AOvVaw2SyWrC60hjKZSWBLBmchPj) |
| 20 | [▪ Universal Anti-forensics Attack against Image Forgery Detection via Multi-modal Guidance (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06530&sa=D&source=editors&ust=1779048536020311&usg=AOvVaw2eapOAOcqZctiLll66ipqJ) |
| 21 | [▪ Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models (Yu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.23255&sa=D&source=editors&ust=1779048536020399&usg=AOvVaw0xnBRVUM3tiU706vMfX102) |
| 22 | [▪ FOCA: Multimodal Malware Classification via Hyperbolic Cross-Attention (Choudhury et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17638&sa=D&source=editors&ust=1779048536020494&usg=AOvVaw1YSwFTpRB8_hl5F7RWPkhU) |
| 23 | [▪ Integrating APK Image and Text Data for Enhanced Threat Detection: A Multimodal Deep Learning Approach to Android Malware (Arifin, Rahman, and Eisty, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08959&sa=D&source=editors&ust=1779048536020621&usg=AOvVaw1meIaD0JxL0sU_rKuWX5Ny) |
| 24 | [▪ Failure Analysis of Safety Controllers in Autonomous Vehicles Under Object-Based LiDAR Attacks (Ganiuly, Bolatbek, and Smaiyl, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22244&sa=D&source=editors&ust=1779048536020729&usg=AOvVaw3mgnxyrf8kDIvLpZvXVsUb) |
| 25 | [▪ FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning (Khalil and Hoque, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09872&sa=D&source=editors&ust=1779048536020831&usg=AOvVaw0bYfzS9RfNYcXN9xjh8Zv9) |
| 26 | [▪ Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities (Xiong et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.02973&sa=D&source=editors&ust=1779048536020931&usg=AOvVaw1i21Ps5mLWaUFmJ2Dmg6ny) |
| 27 | [▪ Vulnerability-Aware Robust Multimodal Adversarial Training (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.18138&sa=D&source=editors&ust=1779048536021026&usg=AOvVaw2KQxdJAXhx13h6Vfw9oF1T) |
| 28 | [▪ Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation (Shang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19257&sa=D&source=editors&ust=1779048536021132&usg=AOvVaw2CGofVl96robY2ltBK6Y5p) |
| 29 | [▪ Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security (Zhao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.16229&sa=D&source=editors&ust=1779048536021230&usg=AOvVaw3TVKNwkhHCO-B1H7pThtdL) |
| 30 | [▪ Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach (Huang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.01042&sa=D&source=editors&ust=1779048536021333&usg=AOvVaw1L4RWu0rY6yjDPuqhZLrjr) |
| 31 | [▪ MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents (Aichberger et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.10809&sa=D&source=editors&ust=1779048536021435&usg=AOvVaw3tJCc_5dE-y9iLyoswDpk2) |
| 32 | [▪ Security Risk of Misalignment between Text and Image in Multi-modal Model (Wang, Ge, and Wang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26105&sa=D&source=editors&ust=1779048536021534&usg=AOvVaw30gjzdNlESszs6sg-LvPPM) |
| 33 | [▪ DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning (Liu, Li, and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18438&sa=D&source=editors&ust=1779048536021656&usg=AOvVaw32TMGSoOxEZqOiuGI0mGFx) |
| 34 | [▪ CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks (Zhang, Li, and Lu, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17687&sa=D&source=editors&ust=1779048536021761&usg=AOvVaw1ohGulDvHtkeoddQrLSLjs) |
| 35 | [▪ Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations (Doppalapudi, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15953&sa=D&source=editors&ust=1779048536021872&usg=AOvVaw1zd3QfgYKmMQpMIvGZ3QwB) |
| 36 | [▪ IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion (Tian et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12395&sa=D&source=editors&ust=1779048536021969&usg=AOvVaw1TXzJXWCju1goxI9CbVfOF) |
| 37 | [▪ Cross-Modal Content Optimization for Steering Web Agent Preferences (Jiang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03612&sa=D&source=editors&ust=1779048536022058&usg=AOvVaw2f6W1ss8Gkf9_bpVtF8iQq) |
| 38 | [▪ Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems (Zhang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.15213&sa=D&source=editors&ust=1779048536022145&usg=AOvVaw2s-2z8GQVWj_YMW908kyEc) |
| 39 | [▪ Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems (Balashov, Ponomarova, and Zhai, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15613&sa=D&source=editors&ust=1779048536022236&usg=AOvVaw2uiS_u5idL4voPQc2oB7bz) |
| 40 | [▪ Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems (Ko et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.23847&sa=D&source=editors&ust=1779048536022328&usg=AOvVaw0k1HJAHJtgOXbTbC25FuuP) |
| 41 | [▪ Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities (Qu, Backes, and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11155&sa=D&source=editors&ust=1779048536022422&usg=AOvVaw0EIYthmzj8kXGcVsptDrJR) |
| 42 | [▪ LLM-Stackelberg Games: Conjectural Reasoning Equilibria and Their Applications to Spearphishing (Zhu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09407&sa=D&source=editors&ust=1779048536022516&usg=AOvVaw01wAF79Er6tW15TCinNW5U) |
| 43 | [▪ The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10016&sa=D&source=editors&ust=1779048536022625&usg=AOvVaw30K4xb3XjBZYXuU6SxdT0G) |
| 44 | [▪ CLIProv: A Contrastive Log-to-Intelligence Multimodal Approach for Threat Detection and Provenance Analysis (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09133&sa=D&source=editors&ust=1779048536022724&usg=AOvVaw3e2I52Ot9csVOmMV4saJL0) |
| 45 | [▪ FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images (Yang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02995&sa=D&source=editors&ust=1779048536022821&usg=AOvVaw37IPrdr_YE5tBtgaKEjjWR) |
| 46 | [▪ JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models\<br>\<br> (Peng et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17568&sa=D&source=editors&ust=1779048536022913&usg=AOvVaw1lX1mrrRuFci1fmeo_sBck) |
| 47 | [▪ Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks (Qraitem et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.00626&sa=D&source=editors&ust=1779048536023005&usg=AOvVaw1TXCBg_WzP6pbu3TPQKtXj) |
| 48 | [▪ Detecting Malicious Concepts Without Image Generation in AIGC (Xu et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08921&sa=D&source=editors&ust=1779048536023092&usg=AOvVaw2zRosa-vbzT4oO-x-U41W-) |
| 49 | [▪ Typographic Attacks in a Multi-Image Setting (Wang, Zhao, and Larson, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08193&sa=D&source=editors&ust=1779048536023186&usg=AOvVaw31P0sIVKYJOk0wvD_wL6nG) |
| 50 | [▪ T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models (Miao et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.05965&sa=D&source=editors&ust=1779048536023282&usg=AOvVaw2BxKcJlhTQXJScstzgpzhc) |

|     |
| --- |
| Multi-Modal Attacks |

**>**

**<**

‍

#### LLM Data Leakage and ML Artifact Collection

- MITRE ATLAS Exfiltration & Collection

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : LLM Data Leakage and ML Artifact Collection

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Identifying AI Web Scrapers Using Canary Tokens (Seiden et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13706&sa=D&source=editors&ust=1779048536758598&usg=AOvVaw34RRoBTD5iln_cq5ggZ4wl) |
| 3 | [▪ LeakDojo: Decoding the Leakage Threats of RAG Systems (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05818&sa=D&source=editors&ust=1779048536758964&usg=AOvVaw1MwtNMJF_hpNedUsTiJviu) |
| 4 | [▪ SecureMCP: A Policy-Enforced LLM Data Access Framework for AIoT Systems via Model Context Protocol (Kim and Yoo, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.05260&sa=D&source=editors&ust=1779048536759194&usg=AOvVaw0wUaKFfFH_SvfZP0i3yPQO) |
| 5 | [▪ Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks (Al-Lawati and Wang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17644&sa=D&source=editors&ust=1779048536759390&usg=AOvVaw34JL0dZ5pm4MZ6-0su5Acx) |
| 6 | [▪ On the Privacy of LLMs: An Ablation Study (Makhlouf et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02255&sa=D&source=editors&ust=1779048536759523&usg=AOvVaw098OVJzIPE3nYx0VcEnL2u) |
| 7 | [▪ Quantamination: Dynamic Quantization Leaks Your Data Across the Batch (Foerster et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26505&sa=D&source=editors&ust=1779048536759653&usg=AOvVaw1xDEdngyLznUCIQU2zeXVv) |
| 8 | [▪ OpenSOC-AI: Democratizing Security Operations with Parameter Efficient LLM Log Analysis (Garware and Zisad, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.26217&sa=D&source=editors&ust=1779048536759810&usg=AOvVaw0J6KamKH28zJwpckydwWWw) |
| 9 | [▪ LLM-CEG: Extending the Classification Error Gauge Framework for Privacy Auditing of Large Language Models (Mivule, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23795&sa=D&source=editors&ust=1779048536759991&usg=AOvVaw10hRgPTc4WyGyn80IoRyDV) |
| 10 | [▪ Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17814&sa=D&source=editors&ust=1779048536760151&usg=AOvVaw3aO-RD23iLcRlzfqh53twH) |
| 11 | [▪ Data Leakage in Automotive Perception: Practitioners' Insights (Babu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06899&sa=D&source=editors&ust=1779048536760308&usg=AOvVaw0wqcYhC7wfSm0jmwPY1iki) |
| 12 | [▪ LLM-Enabled Open-Source Systems in the Wild: An Empirical Study of Vulnerabilities in GitHub Security Advisories (Shifat et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04288&sa=D&source=editors&ust=1779048536760486&usg=AOvVaw0aXv26xorCpcbIRTASZLRO) |
| 13 | [▪ Expert Selections In MoE Models Reveal (Almost) As Much As Text (Nuriyev and Kulp, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04105&sa=D&source=editors&ust=1779048536760647&usg=AOvVaw1xyCBNfBdiFIkaMtkukR2U) |
| 14 | [▪ You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents (Kao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11862&sa=D&source=editors&ust=1779048536760815&usg=AOvVaw33fk3bgMRPKgH1-fTtXl7t) |
| 15 | [▪ Detecting Cryptographically Relevant Software Packages with Collaborative LLMs (Hirsch et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07204&sa=D&source=editors&ust=1779048536760953&usg=AOvVaw2Z--OfR6cn_IvVeKBlGyCF) |
| 16 | [▪ The Silent Spill: Measuring Sensitive Data Leaks Across Public URL Repositories (Ramadan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21826&sa=D&source=editors&ust=1779048536761118&usg=AOvVaw3I-GgHkELZkEHIcwy6OfOi) |
| 17 | [▪ From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models (Soltaniani and Ghafari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22946&sa=D&source=editors&ust=1779048536761251&usg=AOvVaw0xfa386TtYrE8w-32N66-B) |
| 18 | [▪ Can Large Language Models Really Recognize Your Name? (Pham et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.14549&sa=D&source=editors&ust=1779048536761406&usg=AOvVaw1lcLAYqeW_oNbKg_PDaboh) |
| 19 | [▪ A Systemic Evaluation of Multimodal RAG Privacy (Al-Lawati and Wang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17644&sa=D&source=editors&ust=1779048536761589&usg=AOvVaw3rAN10vTHbSn_qsoJBX62W) |
| 20 | [▪ Network-Level Prompt and Trait Leakage in Local Research Agents (Jeong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20282&sa=D&source=editors&ust=1779048536761753&usg=AOvVaw0o-F5LRsNqpq-yA9BwBocR) |
| 21 | [▪ Burn-After-Use for Preventing Data Leakage through a Secure Multi-Tenant Architecture in Enterprise LLM (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06627&sa=D&source=editors&ust=1779048536761898&usg=AOvVaw1gUlItRJ6LZPXpqWDWaJL_) |
| 22 | [▪ SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use (Desai et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06366&sa=D&source=editors&ust=1779048536762037&usg=AOvVaw0cCbxmlskKjFFo7_RhQIIj) |
| 23 | [▪ Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users (Bhardwaj, Diwan, and Wang, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06301&sa=D&source=editors&ust=1779048536762192&usg=AOvVaw1CbDW1OhXGITAr1eUHz4Yp) |
| 24 | [▪ Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation (Qiao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16310&sa=D&source=editors&ust=1779048536762367&usg=AOvVaw38sau7RlA5wHRn1ahik2r2) |
| 25 | [▪ ContextLeak: Auditing Leakage in Private In-Context Learning Methods (Choi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16059&sa=D&source=editors&ust=1779048536762533&usg=AOvVaw0owOoFxIYcTDfO0pMlOwT1) |
| 26 | [▪ ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation (Liu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.21123&sa=D&source=editors&ust=1779048536762703&usg=AOvVaw3fbESHsjszD0bPZUku4bZR) |
| 27 | [▪ Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs (Liu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04668&sa=D&source=editors&ust=1779048536762838&usg=AOvVaw1YKNXUT-sqEW2btHs5nnJo) |
| 28 | [▪ Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models (Dang and Mohaisen, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20799&sa=D&source=editors&ust=1779048536763003&usg=AOvVaw0nYZhqHaAlTOHJ6cYJ4aPC) |
| 29 | [▪ Structured Extraction of Vulnerabilities in OpenVAS and Tenable WAS Reports Using LLMs (Machado et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15745&sa=D&source=editors&ust=1779048536763143&usg=AOvVaw262rtIdYPYar2UeGfqejhk) |
| 30 | [▪ MCP-RiskCue: Can LLM Infer Risk Information From MCP Server System Logs? (Fu and Sun, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05867&sa=D&source=editors&ust=1779048536763322&usg=AOvVaw0FJjb_mnVAOwkWO-hxmDL_) |
| 31 | [▪ MCP-RiskCue: Can LLM infer risk information from MCP server System Logs? (Fu and Sun, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05867&sa=D&source=editors&ust=1779048536763462&usg=AOvVaw2wE5WmCw4QsZRvcI1xBEzF) |
| 32 | [▪ Security Logs to ATT&CK Insights: Leveraging LLMs for High-Level Threat Understanding and Cognitive Trait Inference (Hans et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20930&sa=D&source=editors&ust=1779048536763611&usg=AOvVaw0eO2huUuDHrTiIIfsJon-T) |
| 33 | [▪ Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.24379&sa=D&source=editors&ust=1779048536763777&usg=AOvVaw2_iAOnQvpcmJ10IqA1BLno) |
| 34 | [▪ Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs (Kaneko and Baldwin, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17000&sa=D&source=editors&ust=1779048536764071&usg=AOvVaw2gIqOD3zV1QC6ZRuF626RK) |
| 35 | [▪ Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.07176&sa=D&source=editors&ust=1779048536764256&usg=AOvVaw3zsWrSH8G2z36drNM2WCZG) |
| 36 | [▪ AttackSeqBench: Benchmarking Large Language Models in Analyzing Attack Sequences within Cyber Threat Intelligence (Ma et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.03170&sa=D&source=editors&ust=1779048536764533&usg=AOvVaw1XMoLhSd8sacHfqjQwP1C_) |
| 37 | [▪ You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models (Dubniczky, Borsos, and Norbert, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.03761&sa=D&source=editors&ust=1779048536764771&usg=AOvVaw16t7KBb-z_27Ty0w_cRtXx) |
| 38 | [▪ External Data Extraction Attacks against Retrieval-Augmented Large Language Models (He et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02964&sa=D&source=editors&ust=1779048536764916&usg=AOvVaw2mw23vhV_tj567iwQA2LVL) |
| 39 | [▪ Sentry: Authenticating Machine Learning Artifacts on the Fly (Gan and Ghodsi, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00554&sa=D&source=editors&ust=1779048536765049&usg=AOvVaw3ve_uu8oZyf-f9fpWwjFcR) |
| 40 | [▪ Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence (Meng et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23573&sa=D&source=editors&ust=1779048536765203&usg=AOvVaw1VM1OT67u7maFO9Rca0Nlv) |
| 41 | [▪ LeakyCLIP: Extracting Training Data from CLIP (Chen et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00756&sa=D&source=editors&ust=1779048536765354&usg=AOvVaw0btZiJruyq5XLs1lePjpdx) |
| 42 | [▪ Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data (Croce and South, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19880&sa=D&source=editors&ust=1779048536765480&usg=AOvVaw22s9mQi4-fnYs4B3Y20Qyu) |
| 43 | [▪ LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models (Ran et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18302&sa=D&source=editors&ust=1779048536765587&usg=AOvVaw2E8QbDL5s7e0v9VEWaEoIp) |
| 44 | [▪ Optimizing Canaries for Privacy Auditing with Metagradient Descent (Boglioni et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15836&sa=D&source=editors&ust=1779048536765697&usg=AOvVaw2Fq4SrPW9BMALS1aFoDWzU) |
| 45 | [▪ Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data (Hyrup et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.12216&sa=D&source=editors&ust=1779048536765806&usg=AOvVaw19V1QXxlzJwn-NiLSTGLdZ) |
| 46 | [▪ Training Set Reconstruction from Differentially Private Forests: How Effective is DP? (Gorg\\'e et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05307&sa=D&source=editors&ust=1779048536765924&usg=AOvVaw10rYYKzJfCOtHqmE6txyzx) |
| 47 | [▪ Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs (Ferrand et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16534&sa=D&source=editors&ust=1779048536766026&usg=AOvVaw3BkMl1NtSrprl1GyHhv_Q-) |
| 48 | [▪ From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application? (Pedro et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2308.01990&sa=D&source=editors&ust=1779048536766137&usg=AOvVaw2o5465yVbPMLK9qR9K9vPe) |
| 49 | [▪ RAG-Thief: Scalable Extraction of Private Data from Retrieval-Augmented Generation Applications with Agent-based Attacks (Jiang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14110&sa=D&source=editors&ust=1779048536766257&usg=AOvVaw1boQEf79ug5bQ74yprn0Cn) |
| 50 | [▪ Towards More Realistic Extraction Attacks: An Adversarial Perspective (More, Ganesh, Farnadi, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.02596&sa=D&source=editors&ust=1779048536766377&usg=AOvVaw2Zr-Ezwn4YKXJaNGxqn4Fl) |
| 51 | [▪ Stealing User Prompts from Mixture of Experts (Yona et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22884&sa=D&source=editors&ust=1779048536766486&usg=AOvVaw0mOL03jre78nqCrtzTWY47) |
| 52 | [▪ Breach By A Thousand Leaks: Unsafe Information Leakage in \`Safe' AI Responses (Glukhov et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.02551&sa=D&source=editors&ust=1779048536766590&usg=AOvVaw3AHQ0bEX2w6CkWyNMIppUo) |
| 53 | [▪ CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants (Finkman Noah et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2404.09066&sa=D&source=editors&ust=1779048536766698&usg=AOvVaw3vED_JkEVpR9UhcfRhZeAQ) |
| 54 | [▪ Towards a Theoretical Understanding of Memorization in Diffusion Models (Chen et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02467&sa=D&source=editors&ust=1779048536766811&usg=AOvVaw3xRyDypq1AFyIWgdzjt1TP) |
| 55 | [▪ Extracting Memorized Training Data via Decomposition (Su et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.12367&sa=D&source=editors&ust=1779048536766914&usg=AOvVaw3qGDnZf1aV1t6uS4RDri-S) |
| 56 | [▪ Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships (Gyo Oh et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.12189&sa=D&source=editors&ust=1779048536767013&usg=AOvVaw0CJDr6I36IZj9cl6YlnCDF) |
| 57 | [▪ Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens (Zhang and Wu, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.21248&sa=D&source=editors&ust=1779048536767104&usg=AOvVaw0x418rRxu3n_jxaoabc4u8) |

|     |
| --- |
| LLM Data Leakage and ML Artifact Collection |

**>**

**<**

‍

#### Evade ML Model

Covers:

- MITRE ATLAS Defense Evasion & Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Evade ML Model

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ The Role of Learning in Attacking ML-based Network Intrusion Detection (Domico, Ferrand, and McDaniel, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10299&sa=D&source=editors&ust=1779048536737685&usg=AOvVaw2UNMYg_qwed6cOW6_tWwyS) |
| 3 | [▪ WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08310&sa=D&source=editors&ust=1779048536737864&usg=AOvVaw0BraEyWb0PG1NZeNG88rZQ) |
| 4 | [▪ Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption (Morales, Pastrana, and Tapiador, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06910&sa=D&source=editors&ust=1779048536737977&usg=AOvVaw36J9trFAAKdmIJ9TngyfE4) |
| 5 | [▪ Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00305&sa=D&source=editors&ust=1779048536738072&usg=AOvVaw39FUmnMhI5Xba9sMK8oX2u) |
| 6 | [▪ MalPurifier: Enhancing Android Malware Detection with Adversarial Purification against Evasion Attacks (Zhou et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2312.06423&sa=D&source=editors&ust=1779048536738168&usg=AOvVaw3-qbwsWGpTtpeyhfotyeQ3) |
| 7 | [▪ Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting (Ediga and Chattopadhyay, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01186&sa=D&source=editors&ust=1779048536738260&usg=AOvVaw31dhkROslSYtLKDuagyWf7) |
| 8 | [▪ Trident: Improving Malware Detection with LLMs and Behavioral Features (Saul et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00297&sa=D&source=editors&ust=1779048536738347&usg=AOvVaw06PAX5RNbunSrSIN7S-6rd) |
| 9 | [▪ AsmRAG: LLM-Driven Malware Detection by Retrieving Functionally Similar Assembly Code (Karbab, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23196&sa=D&source=editors&ust=1779048536738434&usg=AOvVaw0VZzKJZ_in-y_MNIN1WGwF) |
| 10 | [▪ Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations (Acharya and Zhang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21310&sa=D&source=editors&ust=1779048536738521&usg=AOvVaw0p19KDjO3dAqZcPIMD-umj) |
| 11 | [▪ Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing (Peng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05719&sa=D&source=editors&ust=1779048536738611&usg=AOvVaw0zW6g7J5xmvXfzn8EsfjQj) |
| 12 | [▪ Evasion Adversarial Attacks Remain Impractical Against ML-based Network Intrusion Detection Systems, Especially Dynamic Ones (elShehaby and Matrawy, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2306.05494&sa=D&source=editors&ust=1779048536738710&usg=AOvVaw1yexZ2Ej5cTlNhPD_D11Eu) |
| 13 | [▪ ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks (Elnawawy et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26093&sa=D&source=editors&ust=1779048536738804&usg=AOvVaw1KFODjNISaP8x3I4-Tjv0C) |
| 14 | [▪ Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25500&sa=D&source=editors&ust=1779048536738893&usg=AOvVaw35n2_HhseqJIV-AB_j5ti_) |
| 15 | [▪ Targeted Adversarial Traffic Generation : Black-box Approach to Evade Intrusion Detection Systems in IoT Networks (Debicha et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23438&sa=D&source=editors&ust=1779048536738981&usg=AOvVaw3N_a1NMJ3poSUOkgpuWE7o) |
| 16 | [▪ Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents (Aonzo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15457&sa=D&source=editors&ust=1779048536739085&usg=AOvVaw13J9YOSWUh6UurWJhx567a) |
| 17 | [▪ FuzzySQL: Uncovering Hidden Vulnerabilities in DBMS Special Features with LLM-Driven Fuzzing (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19490&sa=D&source=editors&ust=1779048536739222&usg=AOvVaw0ppY05-yONCeyYjQoeyhW7) |
| 18 | [▪ Anticipating Adversary Behavior in DevSecOps Scenarios through Large Language Models (Caballero et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14106&sa=D&source=editors&ust=1779048536739365&usg=AOvVaw1oHbu8znYtrYCCs6Zh2ckC) |
| 19 | [▪ SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.16981&sa=D&source=editors&ust=1779048536739469&usg=AOvVaw1NSQf__-epi5blale0PYZD) |
| 20 | [▪ Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach (Ali et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12183&sa=D&source=editors&ust=1779048536739557&usg=AOvVaw1G243QphCQjx6X3yC0geva) |
| 21 | [▪ StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors (Ranganath and Ramesh, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08934&sa=D&source=editors&ust=1779048536739646&usg=AOvVaw0Go1Tq6SIEgepqfeAntPVv) |
| 22 | [▪ DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems (Ou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02569&sa=D&source=editors&ust=1779048536739750&usg=AOvVaw2a4tFaPABhoqZgSN7MF9mW) |
| 23 | [▪ Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks (Zhang and Zhang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23101&sa=D&source=editors&ust=1779048536739844&usg=AOvVaw0nDUWHtjzexWmWZmjVzQK5) |
| 24 | [▪ "Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00364&sa=D&source=editors&ust=1779048536739932&usg=AOvVaw2bS3FWA80iL2weTzQ0QHGW) |
| 25 | [▪ Semantics-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00305&sa=D&source=editors&ust=1779048536740015&usg=AOvVaw0DfYB8eI3IXivAwgrbr4QN) |
| 26 | [▪ In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement (Shetty, Joshi, and Kanhere, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22169&sa=D&source=editors&ust=1779048536740103&usg=AOvVaw0SsbypYLwhVkFTRKB3obNq) |
| 27 | [▪ AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises (Tung et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22720&sa=D&source=editors&ust=1779048536740197&usg=AOvVaw0tWPUFSBtoh2cj-SipQJ7J) |
| 28 | [▪ ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses (He et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21586&sa=D&source=editors&ust=1779048536740286&usg=AOvVaw10Rp9chZCh2z2kscwbARNn) |
| 29 | [▪ CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13112&sa=D&source=editors&ust=1779048536740379&usg=AOvVaw0NMZGy5TBUeBu4zepLStpY) |
| 30 | [▪ A Decompilation-Driven Framework for Malware Detection with Large Language Models (Chawla and Prasad, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09035&sa=D&source=editors&ust=1779048536740464&usg=AOvVaw0Cotsx8WhuZMwUiiuAHNBY) |
| 31 | [▪ MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization (Gu, Li, and Hu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08564&sa=D&source=editors&ust=1779048536740549&usg=AOvVaw13nxylzR1mANHewoyvFd0V) |
| 32 | [▪ VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference (Qi et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06461&sa=D&source=editors&ust=1779048536740642&usg=AOvVaw1HF3GZSGx80eBz8jC1Lw7f) |
| 33 | [▪ Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools? (Quantrill et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00559&sa=D&source=editors&ust=1779048536740734&usg=AOvVaw0wNs0yJlBBiLPHuEV3-Ls5) |
| 34 | [▪ Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks (Xu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21241&sa=D&source=editors&ust=1779048536740825&usg=AOvVaw3AxTmrGIby5_lVYi1E6ZC0) |
| 35 | [▪ Automated Penetration Testing with LLM Agents and Classical Planning (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11143&sa=D&source=editors&ust=1779048536740910&usg=AOvVaw143zrMtkkYdDAI79CgOZai) |
| 36 | [▪ NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00119&sa=D&source=editors&ust=1779048536740995&usg=AOvVaw0F8i8DMP627ez808KFY2Wt) |
| 37 | [▪ ExtendAttack: Attacking Servers of LRMs via Extending Reasoning (Zhu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.13737&sa=D&source=editors&ust=1779048536741077&usg=AOvVaw382YNAdlDmsgolJFhSEXNp) |
| 38 | [▪ Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning (Domico et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.01734&sa=D&source=editors&ust=1779048536741166&usg=AOvVaw0neaw1yQNHqetOFrfPg6HC) |
| 39 | [▪ MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification (Luo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14129&sa=D&source=editors&ust=1779048536741269&usg=AOvVaw2MbekWEBqeeR5-_1KLE63-) |
| 40 | [▪ Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment (Zhang and sun, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06852&sa=D&source=editors&ust=1779048536741354&usg=AOvVaw1lcsI4mVaybTXLKdbknmR0) |
| 41 | [▪ GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors (Meng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.08188&sa=D&source=editors&ust=1779048536741441&usg=AOvVaw2x6KOh0pR-H8f9AKniQURi) |
| 42 | [▪ The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems (Choi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.20995&sa=D&source=editors&ust=1779048536741529&usg=AOvVaw10ryQV_vTlqVgx-trAhz1D) |
| 43 | [▪ Adversarial Pre-Padding: Generating Evasive Network Traffic Against Transformer-Based Classifiers (Jing et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25810&sa=D&source=editors&ust=1779048536741623&usg=AOvVaw0SX9Xc1CO0x3JuM-88lX8W) |
| 44 | [▪ Detecting Various DeFi Price Manipulations with LLM Reasoning (Zhong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.11521&sa=D&source=editors&ust=1779048536741727&usg=AOvVaw1-TLPfitdlx8a4Bmqd7sX6) |
| 45 | [▪ Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks (Feng and Sakurai, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23313&sa=D&source=editors&ust=1779048536741855&usg=AOvVaw0CxnR_N8T_ru-jGu6a7aAP) |
| 46 | [▪ Black-Box Evasion Attacks on Data-Driven Open RAN Apps: Tailored Design and Experimental Evaluation (Gajjar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18160&sa=D&source=editors&ust=1779048536741956&usg=AOvVaw0lI2DNEMngGfsXrqQIlwe7) |
| 47 | [▪ From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15 (Rehman et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17883&sa=D&source=editors&ust=1779048536742058&usg=AOvVaw2v0sLsYJnxtZ_sbkl0EPIT) |
| 48 | [▪ SoK: Adversarial Evasion Attacks Practicality in NIDS Domain and the Impact of Dynamic Learning (elShehaby and Matrawy, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2306.05494&sa=D&source=editors&ust=1779048536742147&usg=AOvVaw1xFXLqy6WHXGUkLyDzAJ2z) |
| 49 | [▪ A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14906&sa=D&source=editors&ust=1779048536742240&usg=AOvVaw3Ro6tjwxdVG2ESsJ_t6pwd) |
| 50 | [▪ VulSolver: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00882&sa=D&source=editors&ust=1779048536742326&usg=AOvVaw27PwafB3LaD-zD7k6iwQud) |
| 51 | [▪ Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework (Han et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.13811&sa=D&source=editors&ust=1779048536742415&usg=AOvVaw2gmuh7_wduEDGpgataWsSW) |
| 52 | [▪ Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks (Engineering et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.11711&sa=D&source=editors&ust=1779048536742508&usg=AOvVaw2k1hlt1KNz9t4k5xmVa0rk) |
| 53 | [▪ Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection (Creo, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01887&sa=D&source=editors&ust=1779048536742598&usg=AOvVaw06GmbCZTG2OLbTm0-qMSGA) |
| 54 | [▪ AdVAR-DNN: Adversarial Misclassification Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01107&sa=D&source=editors&ust=1779048536742713&usg=AOvVaw2dAdRbUgwL9qLn7kD-7z8z) |
| 55 | [▪ ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models (Yook et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21985&sa=D&source=editors&ust=1779048536742855&usg=AOvVaw2_yGxmeqGPKQKgCS6pNkyA) |
| 56 | [▪ Hierarchical Graph Neural Network for Compressed Speech Steganalysis (Hemis et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21591&sa=D&source=editors&ust=1779048536742953&usg=AOvVaw3rvxHPMpIbZusKYNqBIH8q) |
| 57 | [▪ Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms (Sato, Ohki, and Nishigaki, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21538&sa=D&source=editors&ust=1779048536743049&usg=AOvVaw0_G_JWitRLXSMyIq-a_StR) |
| 58 | [▪ GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI (Kalyanapu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17033&sa=D&source=editors&ust=1779048536743165&usg=AOvVaw0qaFG5CfVrMEh1-ujcBgvv) |
| 59 | [▪ BandFuzz: An ML-powered Collaborative Fuzzing Framework (Shi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10845&sa=D&source=editors&ust=1779048536743262&usg=AOvVaw35zIPc_76pWUrNqDXaLZTa) |
| 60 | [▪ PotentRegion4MalDetect: Advanced Features from Potential Malicious Regions for Malware Detection (Koppanati, Santra, and Peddoju, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06723&sa=D&source=editors&ust=1779048536743354&usg=AOvVaw0iL7HCP697tmNFPazSsio3) |
| 61 | [▪ Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors (Wang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19230&sa=D&source=editors&ust=1779048536743437&usg=AOvVaw2Lol43ag1iUMRdaJA-1Xzh) |

|     |
| --- |
| Evade ML Model |

**>**

**<**

‍

#### Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access

Covers:

- OWASP LLM 10: Model Theft
- OWASP ML 05: Model Theft
- MITRE ATLAS Exfiltration and ML Model Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It? (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12827&sa=D&source=editors&ust=1779048535970070&usg=AOvVaw1c94tACkp85hmX8e99KGNZ) |
| 3 | [▪ Protecting the Trace: A Principled Black-Box Approach Against Distillation Attacks (Hartman et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23238&sa=D&source=editors&ust=1779048535970203&usg=AOvVaw2AZMn8BcjioZE_n2Bdrl8J) |
| 4 | [▪ Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model! (Yoon et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.03014&sa=D&source=editors&ust=1779048535970292&usg=AOvVaw0WaTq2TKyE9M8-ArqbWOTS) |
| 5 | [▪ Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21829&sa=D&source=editors&ust=1779048535970391&usg=AOvVaw0VKa4ZT5Iesxy7AuQRPT2r) |
| 6 | [▪ TrEEStealer: Stealing Decision Trees via Enclave Side Channels (Sander et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18716&sa=D&source=editors&ust=1779048535970469&usg=AOvVaw2iGckd1PTZV2vMZLOwk3G4) |
| 7 | [▪ AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection (Zhao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11950&sa=D&source=editors&ust=1779048535970553&usg=AOvVaw3ceSrHkVCHSuq0on1IYtv7) |
| 8 | [▪ Automated Malware Family Classification using Weighted Hierarchical Ensembles of Large Language Models (Bai et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02490&sa=D&source=editors&ust=1779048535970643&usg=AOvVaw3bni719x6m5l7LjOBRgoGI) |
| 9 | [▪ LiteGuard: Efficient Task-Agnostic Model Fingerprinting with Enhanced Generalization (Yang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.24982&sa=D&source=editors&ust=1779048535970720&usg=AOvVaw1OY-DCTEESgst3mxaj4eoP) |
| 10 | [▪ Fingerprinting Deep Neural Networks for Ownership Protection: An Analytical Approach (Yang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.21411&sa=D&source=editors&ust=1779048535970795&usg=AOvVaw3wMoQnp8YMH_ng37cMR8xL) |
| 11 | [▪ Towards Model Extraction Attacks in GAN-Based Image Translation via Domain Shift Mitigation (Mi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2403.07673&sa=D&source=editors&ust=1779048535970874&usg=AOvVaw31xPEevuvO8zwyliGBj3hq) |
| 12 | [▪ Good-Enough LLM Obfuscation (GELO) (Belikov and Fedotov, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.05035&sa=D&source=editors&ust=1779048535970971&usg=AOvVaw3JdPCDj9ZvLuOMFluXdnMo) |
| 13 | [▪ Osmosis Distillation: Model Hijacking with the Fewest Samples (Shi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04859&sa=D&source=editors&ust=1779048535971070&usg=AOvVaw31cngVUYKNNaVXpeboMECc) |
| 14 | [▪ Few-shot Model Extraction Attacks against Sequential Recommender Systems (Zhang and Liu, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.11677&sa=D&source=editors&ust=1779048535971164&usg=AOvVaw33MIzzg0SvgKdCvjQCenrM) |
| 15 | [▪ DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection (Yan et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16530&sa=D&source=editors&ust=1779048535971250&usg=AOvVaw3gGs-U7sWQ5py9ZbnLVjEs) |
| 16 | [▪ Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious Packages (Guo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16458&sa=D&source=editors&ust=1779048535971338&usg=AOvVaw1M0igAyBaXgg3JeXOroNsA) |
| 17 | [▪ KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12986&sa=D&source=editors&ust=1779048535971425&usg=AOvVaw2ZCUhqiDA2mxdVaLQvQfIi) |
| 18 | [▪ Deep Dive into the Abuse of DL APIs To Create Malicious AI Models and How to Detect Them (Nabeel and Starov, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.04553&sa=D&source=editors&ust=1779048535971516&usg=AOvVaw15wZfHYoqWErviVlKBPwCO) |
| 19 | [▪ Aggressive Compression Enables LLM Weight Theft (Brown et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01296&sa=D&source=editors&ust=1779048535971602&usg=AOvVaw2h5CPts_IqGznsyLXo6QpU) |
| 20 | [▪ Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00274&sa=D&source=editors&ust=1779048535971708&usg=AOvVaw0D7_EjfYlQ_6HNOFTHwFeV) |
| 21 | [▪ DivQAT: Enhancing Robustness of Quantized Convolutional Neural Networks against Model Extraction Attacks (Khaled, Magalh\\~aes, and Nicolescu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.23948&sa=D&source=editors&ust=1779048535971808&usg=AOvVaw1hB7xvRLsUfn4TcoGLNmVk) |
| 22 | [▪ To See or Not to See -- Fingerprinting Devices in Adversarial Environments Amid Advanced Machine Learning (Feng, Haddad, and Sehatbakhsh, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.08264&sa=D&source=editors&ust=1779048535971907&usg=AOvVaw2UasmfBR98HT58CM-b2vIk) |
| 23 | [▪ A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16538&sa=D&source=editors&ust=1779048535972003&usg=AOvVaw248YthB0qoSGdQ26MRKoun) |
| 24 | [▪ AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models (Qiu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.20804&sa=D&source=editors&ust=1779048535972123&usg=AOvVaw0NmDlYMk-W8BmQXgRG6A2g) |
| 25 | [▪ On Stealing Graph Neural Network Models (Podhajski et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07170&sa=D&source=editors&ust=1779048535972217&usg=AOvVaw3b6bokgcTZws_3F1TZycut) |
| 26 | [▪ Quantifying the Risk of Transferred Black Box Attacks (Cox and Bunzel, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05102&sa=D&source=editors&ust=1779048535972310&usg=AOvVaw2CCQTTPQcPb4DPt8OquM3b) |
| 27 | [▪ SLIP-SEC: Formalizing Secure Protocols for Model IP Protection (Jain et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24999&sa=D&source=editors&ust=1779048535972430&usg=AOvVaw2AnwYQd2fLIkJSSLcB5FpA) |
| 28 | [▪ $\\delta$-STEAL: LLM Stealing Attack with Local Differential Privacy (Dang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21946&sa=D&source=editors&ust=1779048535972571&usg=AOvVaw0KORGGDc0OUaMF5KUt02xF) |
| 29 | [▪ Black Box Absorption: LLMs Undermining Innovative Ideas (Cao, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20612&sa=D&source=editors&ust=1779048535972682&usg=AOvVaw1DFRSvdA6qiU27G_19stp2) |
| 30 | [▪ When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking (Rehman et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17884&sa=D&source=editors&ust=1779048535972769&usg=AOvVaw3jIRgDeIjBOUocWVIOKm0r) |
| 31 | [▪ DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge (Mohanty et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16716&sa=D&source=editors&ust=1779048535972847&usg=AOvVaw3F9lhdrSCMfjl7MYcAxpnp) |
| 32 | [▪ Rotation, Scale, and Translation Resilient Black-box Fingerprinting for Intellectual Property Protection of EaaS Models (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16706&sa=D&source=editors&ust=1779048535972945&usg=AOvVaw2xT0RAHXH7d-qo_xbO6gAO) |
| 33 | [▪ CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.13903&sa=D&source=editors&ust=1779048535973027&usg=AOvVaw0qyGoL9o4STFMtEbAS-2kR) |
| 34 | [▪ LLMAtKGE: Large Language Models as Explainable Attackers against Knowledge Graph Embeddings (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11584&sa=D&source=editors&ust=1779048535973128&usg=AOvVaw3RU0g0TAFOj_66gJkis0y1) |
| 35 | [▪ Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial? (Ito, Miura, and Todo, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06692&sa=D&source=editors&ust=1779048535973221&usg=AOvVaw08BBKiaqzYYvWMI5pjo9S0) |
| 36 | [▪ Real-VulLLM: An LLM Based Assessment Framework in the Wild (Safdar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.04056&sa=D&source=editors&ust=1779048535973313&usg=AOvVaw0ITQDZZRVL6CPwjEm3kPGm) |
| 37 | [▪ From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization (Xi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02389&sa=D&source=editors&ust=1779048535973399&usg=AOvVaw299oKA66rMWcsVH2pgo719) |
| 38 | [▪ POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment (Tang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01552&sa=D&source=editors&ust=1779048535973488&usg=AOvVaw343qUARn6TOzDWKWKnGDaZ) |
| 39 | [▪ Stealing AI Model Weights Through Covert Communication Channels (Barbaza et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00151&sa=D&source=editors&ust=1779048535973586&usg=AOvVaw3yCVtpFHHu9Ds7XrjxCg-5) |
| 40 | [▪ Model Extraction Attacks Revisited (Liang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.05386&sa=D&source=editors&ust=1779048535973671&usg=AOvVaw0cUw7W1xe29f3fOrXjp6TH) |
| 41 | [▪ Scalable Fingerprinting of Large Language Models (Nasery et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.07760&sa=D&source=editors&ust=1779048535973757&usg=AOvVaw3rjQIDhSqqrZr_GEfRcFmk) |
| 42 | [▪ Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs (Cai et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.04715&sa=D&source=editors&ust=1779048535973847&usg=AOvVaw0P4eERoQ2LKz3OnSpei0Wk) |
| 43 | [▪ StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data (Wang et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23594&sa=D&source=editors&ust=1779048535973938&usg=AOvVaw2Z0d43E3K0qJCIM-P4CQ4V) |
| 44 | [▪ PhishParrot: LLM-Driven Adaptive Crawling to Unveil Cloaked Phishing Sites (Nakano, Koide, and Chiba, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.02035&sa=D&source=editors&ust=1779048535974036&usg=AOvVaw3l-CNRmR-cH9T88gnoetJB) |
| 45 | [▪ RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging (He et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01784&sa=D&source=editors&ust=1779048535974139&usg=AOvVaw0-FkLFb-2BZRcn1VCH0a5e) |
| 46 | [▪ "Energon": Unveiling Transformers from GPU Power and Thermal Side-Channels (Chaudhuri et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01768&sa=D&source=editors&ust=1779048535974236&usg=AOvVaw3r9MTBuEs0JlhoEoOhOxLq) |
| 47 | [▪ Leveraging Machine Learning for Botnet Attack Detection in Edge-Computing Assisted IoT Networks (Rupanetti and Kaabouch, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01542&sa=D&source=editors&ust=1779048535974335&usg=AOvVaw3vqszqbap-XgrOwYJSeqb3) |
| 48 | [▪ Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report (Weerawardhena et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01059&sa=D&source=editors&ust=1779048535974425&usg=AOvVaw2bI0oT1kt23jUwC_MPDMqS) |
| 49 | [▪ Safe machine learning model release from Trusted Research Environments: The SACRO-ML package (Smith et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2212.01233&sa=D&source=editors&ust=1779048535974513&usg=AOvVaw0KxED_LS3HKC3JH7Nl3GJD) |
| 50 | [▪ Theoretically Unmasking Inference Attacks Against LDP-Protected Clients in Federated Vision Models (Nguyen et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17292&sa=D&source=editors&ust=1779048535974604&usg=AOvVaw21zneMcBqlc_-eE3P7OACp) |
| 51 | [▪ Medical Image De-Identification Benchmark Challenge (Pei et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.23608&sa=D&source=editors&ust=1779048535974690&usg=AOvVaw0elqLcdMygX3iwlrddCunQ) |
| 52 | [▪ LLM-Based Identification of Infostealer Infection Vectors from Screenshots: The Case of Aurora (Ruellan, Clay, and Ascoli, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.23611&sa=D&source=editors&ust=1779048535974769&usg=AOvVaw18dllBgesqss_kMcVUC9B-) |
| 53 | [▪ Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation (Chen et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.23229&sa=D&source=editors&ust=1779048535974847&usg=AOvVaw2IYGan5qAvHQFsw461AOqr) |
| 54 | [▪ The Impact of Train-Test Leakage on Machine Learning-based Android Malware Detection (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19364&sa=D&source=editors&ust=1779048535974927&usg=AOvVaw3ymq86ufIM_x58yoMRDXG_) |
| 55 | [▪ Cascading and Proxy Membership Inference Attacks (Du et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21412&sa=D&source=editors&ust=1779048535975001&usg=AOvVaw3xUMcjzu0yzWlmOChbCKTr) |
| 56 | [▪ Radio Adversarial Attacks on EMG-based Gesture Recognition Networks (Xie, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21387&sa=D&source=editors&ust=1779048535975072&usg=AOvVaw0rfTkr_PzASbYGfQDm54h8) |
| 57 | [▪ Learning-based Privacy-Preserving Graph Publishing Against Sensitive Link Inference Attacks (Wu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21139&sa=D&source=editors&ust=1779048535975142&usg=AOvVaw2CCGI8KQNuOdrtPNUv5KyF) |
| 58 | [▪ Privacy-Preserving AI for Encrypted Medical Imaging: A Framework for Secure Diagnosis and Learning (Siam and Shohan, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21060&sa=D&source=editors&ust=1779048535975213&usg=AOvVaw0b5Wc_XXu-1MoZ1915RxOw) |
| 59 | [▪ Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset (Song et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20688&sa=D&source=editors&ust=1779048535975283&usg=AOvVaw085N59Lmtennf0i0ocjaXG) |
| 60 | [▪ Encrypted-State Quantum Compilation Scheme Based on Quantum Circuit Obfuscation (Zhang, Shang, and Guo, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17589&sa=D&source=editors&ust=1779048535975358&usg=AOvVaw2PFCxkv9Qd7RXqkzrrl7Z4) |
| 61 | [▪ CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16872&sa=D&source=editors&ust=1779048535975429&usg=AOvVaw2i2NMAUdoN_BZuxFC1fUDy) |
| 62 | [▪ SVAgent: AI Agent for Hardware Security Verification Assertion (Guo et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16203&sa=D&source=editors&ust=1779048535975509&usg=AOvVaw3gWNbm72DX0CFvo4K4xKMf) |
| 63 | [▪ DP2Guard: A Lightweight and Byzantine-Robust Privacy-Preserving Federated Learning Scheme for Industrial IoT (Han et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16134&sa=D&source=editors&ust=1779048535975592&usg=AOvVaw18I_SwYZ4rg3Rr-2yyC5zk) |
| 64 | [▪ zkFL: Zero-Knowledge Proof-based Gradient Aggregation for Federated Learning (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2310.02554&sa=D&source=editors&ust=1779048535975692&usg=AOvVaw0VaUBQFsM8Ywj5HpEmTCj2) |
| 65 | [▪ Detecting Benchmark Contamination Through Watermarking (Sander et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.17259&sa=D&source=editors&ust=1779048535975781&usg=AOvVaw3GGqeTyXTFLbcGEcnpVuP8) |
| 66 | [▪ Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection (Li, Zhang, and Zhao, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15101&sa=D&source=editors&ust=1779048535975876&usg=AOvVaw0JanwsD0JBbx1Szbxv5JpV) |
| 67 | [▪ VMask: Tunable Label Privacy Protection for Vertical Federated Learning via Layer Masking (Tan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14629&sa=D&source=editors&ust=1779048535975962&usg=AOvVaw21ptRmYqfZXqgeAXdgC2hZ) |
| 68 | [▪ Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives (Zeng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14519&sa=D&source=editors&ust=1779048535976058&usg=AOvVaw3oc5pjH47N7KFkvseTpYA0) |
| 69 | [▪ FuSeFL: Fully Secure and Scalable Cross-Silo Federated Learning (Ghinani and Sadredini, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13591&sa=D&source=editors&ust=1779048535976139&usg=AOvVaw3baLmjqESSVubm5BFbS4Lu) |
| 70 | [▪ A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique (Sueyoshi, Nishikawa, and Kiya, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12730&sa=D&source=editors&ust=1779048535976222&usg=AOvVaw3pO4RgRRLXSVwoMSb1SEru) |
| 71 | [▪ A Crowdsensing Intrusion Detection Dataset For Decentralized Federated Learning Models (Feng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13313&sa=D&source=editors&ust=1779048535976304&usg=AOvVaw35kFH0CxPrZgherMDLxTQp) |
| 72 | [▪ Privacy Against Agnostic Inference Attacks in Vertical Federated Learning (Varasteh, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2302.05545&sa=D&source=editors&ust=1779048535976388&usg=AOvVaw2QhBACX2Ji0AHd_5hsd0Zk) |
| 73 | [▪ FacialMotionID: Identifying Users of Mixed Reality Headsets using Abstract Facial Motion Representations (Castro et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11138&sa=D&source=editors&ust=1779048535976483&usg=AOvVaw1_2K8Dq0crWt_-u3pvHcFO) |
| 74 | [▪ AdRo-FL: Informed and Secure Client Selection for Federated Learning in the Presence of Adversarial Aggregator (Hossain et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17805&sa=D&source=editors&ust=1779048535976566&usg=AOvVaw0UxD36L_Q9OXpPdnOI485r) |
| 75 | [▪ TimberStrike: Dataset Reconstruction Attack Revealing Privacy Leakage in Federated Tree-Based Systems (Gennaro et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.07605&sa=D&source=editors&ust=1779048535976646&usg=AOvVaw282iHmaLac41a0atU6cjDz) |
| 76 | [▪ Split Happens: Combating Advanced Threats with Split Learning and Function Secret Sharing (Khan, Budzys, and Michalas, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10494&sa=D&source=editors&ust=1779048535976727&usg=AOvVaw2R76_QTPJA1vZqOa8KbyFw) |
| 77 | [▪ Secure and Efficient UAV-Based Face Detection via Homomorphic Encryption and Edge Computing (Duc et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09860&sa=D&source=editors&ust=1779048535976805&usg=AOvVaw0At6_8Kuo77y9n72vGNv03) |
| 78 | [▪ Efficient Private Inference Based on Helper-Assisted Malicious Security Dishonest Majority MPC (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09607&sa=D&source=editors&ust=1779048535976881&usg=AOvVaw20ZWNKF6ukZTjfC8kC6M4S) |
| 79 | [▪ Invariant-based Robust Weights Watermark for Large Language Models (Guo et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08288&sa=D&source=editors&ust=1779048535976972&usg=AOvVaw3B0vCx6a98qoVUAmD3hvib) |
| 80 | [▪ Research on Data Right Confirmation Mechanism of Federated Learning based on Blockchain (Cheng and Guo, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.08476&sa=D&source=editors&ust=1779048535977058&usg=AOvVaw39eeIVbvpXSE2vOd2FITtc) |
| 81 | [▪ FedP3E: Privacy-Preserving Prototype Exchange for Non-IID IoT Malware Detection in Cross-Silo Federated Learning (Darwish et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07258&sa=D&source=editors&ust=1779048535977147&usg=AOvVaw1NbN4gGnbpIn69ESu5Sc-U) |
| 82 | [▪ A Blockchain Solution for Collaborative Machine Learning over IoT (Beis-Penedo et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2311.14136&sa=D&source=editors&ust=1779048535977232&usg=AOvVaw3cjn1S85_BUqgP0yJVYK-M) |
| 83 | [▪ ZKTorch: Compiling ML Inference to Zero-Knowledge Proofs via Parallel Proof Accumulation (Chen, Tang, and Kang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07031&sa=D&source=editors&ust=1779048535977318&usg=AOvVaw0-CWZAeNl-e_noExwqh-8X) |
| 84 | [▪ BarkBeetle: Stealing Decision Tree Models with Fault Injection (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06986&sa=D&source=editors&ust=1779048535977401&usg=AOvVaw1wHeFWU33gYEWVhAdgxrV1) |
| 85 | [▪ Fundamental Limits of Hierarchical Secure Aggregation with Cyclic User Association (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.04564&sa=D&source=editors&ust=1779048535977483&usg=AOvVaw003QRQAIvpBqMEImotFmYZ) |
| 86 | [▪ Learning Federated Neural Graph Databases for Answering Complex Queries from Distributed Knowledge Graphs (Hu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.14609&sa=D&source=editors&ust=1779048535977580&usg=AOvVaw1TsZSJ3sH7EpqigXfnxlbC) |
| 87 | [▪ TT-TFHE: a Torus Fully Homomorphic Encryption-Friendly Neural Network Architecture (Benamira et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2302.01584&sa=D&source=editors&ust=1779048535977672&usg=AOvVaw2mYEjcnprL5J9BpUw3StgH) |
| 88 | [▪ A Model Stealing Attack Against Multi-Exit Networks (Pan et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2305.13584&sa=D&source=editors&ust=1779048535977764&usg=AOvVaw2WhojYsHn0mhcBZKmmalPM) |
| 89 | [▪ Model Stealing Attack against Graph Classification with Authenticity, Uncertainty and Diversity (Zhu et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2312.10943&sa=D&source=editors&ust=1779048535977849&usg=AOvVaw215wr1e5Esbb5dOMCRqQBp) |

|     |
| --- |
| Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access |

**>**

**<**

‍

#### Model Inversion Attack

Covers:

- OWASP ML 03: Model Inversion Attack
- MITRE ATLAS Exfiltration

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Inversion Attack

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ FERMI: Exploiting Relations for Membership Inference Against Tabular Diffusion Models (Mahyar et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11527&sa=D&source=editors&ust=1779048535951585&usg=AOvVaw12SXfvuDWmVtgqLiKZYnJ9) |
| 3 | [▪ Auditing Data Membership in Reinforcement Learning With Verifiable Rewards (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14045&sa=D&source=editors&ust=1779048535951748&usg=AOvVaw1AmzcPjGKbCY-hWujHammA) |
| 4 | [▪ Membership Inference Attacks on Vision-Language-Action Models (Peng et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07088&sa=D&source=editors&ust=1779048535951862&usg=AOvVaw0MySaGS4VJ73eOwM-u5PrW) |
| 5 | [▪ SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing (Sun et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01150&sa=D&source=editors&ust=1779048535951969&usg=AOvVaw22LeXB42G0fcrDyLMKY0_3) |
| 6 | [▪ Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models (Chen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06423&sa=D&source=editors&ust=1779048535952082&usg=AOvVaw3vc9U5xXInkEXAhe5BXzqG) |
| 7 | [▪ Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering (Kulkarni, Koskela, and Zumot, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04116&sa=D&source=editors&ust=1779048535952188&usg=AOvVaw2IzrbBxET1mvy9akTLdZvH) |
| 8 | [▪ Membership Inference Attacks Against Video Large Language Models (Song et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27002&sa=D&source=editors&ust=1779048535952319&usg=AOvVaw1Md-xouNKNvjSHKGY7hL1K) |
| 9 | [▪ A Data-Free Membership Inference Attack on Federated Learning in Hardware Assurance (Lee et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19891&sa=D&source=editors&ust=1779048535952424&usg=AOvVaw16j8Z6WTNJ7gH6ve-8Vph1) |
| 10 | [▪ No More Guessing: a Verifiable Gradient Inversion Attack in Federated Learning (Diana et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15063&sa=D&source=editors&ust=1779048535952525&usg=AOvVaw0ZarlLpX4OWX0Sjfv1vEsI) |
| 11 | [▪ Label Leakage Attacks in Machine Unlearning: A Parameter and Inversion-Based Approach (Zheng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07386&sa=D&source=editors&ust=1779048535952646&usg=AOvVaw1xY_1sfoKS7_PQIHvhAfk_) |
| 12 | [▪ FedSpy-LLM: Towards Scalable and Generalizable Data Reconstruction Attacks from Gradients on LLMs (Meerza, Wang, and Liu, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.06297&sa=D&source=editors&ust=1779048535952758&usg=AOvVaw0CFwaKBEztSLI4uIgfaLUq) |
| 13 | [▪ ReproMIA: A Comprehensive Analysis of Model Reprogramming for Proactive Membership Inference Attacks (Huang, Wang, and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28942&sa=D&source=editors&ust=1779048535952865&usg=AOvVaw1-Cu8ZsXhGI0z3T3Z3t5cG) |
| 14 | [▪ A Divide-and-Conquer Strategy for Hard-Label Extraction of Deep Neural Networks via Side-Channel Attacks (Coqueret et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2411.10174&sa=D&source=editors&ust=1779048535952969&usg=AOvVaw2RSl_aqTLGdoZhenQ2oA3T) |
| 15 | [▪ SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models (Dikici et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01147&sa=D&source=editors&ust=1779048535953087&usg=AOvVaw1JVeNnL7Vih7q53r5vaT91) |
| 16 | [▪ \\texttt{ReproMIA}: A Comprehensive Analysis of Model Reprogramming for Proactive Membership Inference Attacks (Huang, Wang, and Wang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.28942&sa=D&source=editors&ust=1779048535953198&usg=AOvVaw04kBNk1kCnmvuFP1FjXKws) |
| 17 | [▪ Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents (Tran, Kotevska, and Xiong, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.19375&sa=D&source=editors&ust=1779048535953306&usg=AOvVaw2MwA38MvCT9JIsondxLAvs) |
| 18 | [▪ ARES: Scalable and Practical Gradient Inversion Attack in Federated Learning through Activation Recovery (Gong et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17623&sa=D&source=editors&ust=1779048535953436&usg=AOvVaw1EGvctM00HMfsOOrmcfCMx) |
| 19 | [▪ Exponential-Family Membership Inference: From LiRA and RMIA to BaVarIA (Br\\"annvall, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11799&sa=D&source=editors&ust=1779048535953556&usg=AOvVaw3zE4hROrtsK7bs4HXJlh__) |
| 20 | [▪ Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions (Jebreel et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07567&sa=D&source=editors&ust=1779048535953672&usg=AOvVaw1nStTexx_TxT5hHdHBzkU5) |
| 21 | [▪ How to Steal Reasoning Without Reasoning Traces (Zhang, Morris, and Shmatikov, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07267&sa=D&source=editors&ust=1779048535953787&usg=AOvVaw3a5euv2F3HA_sORhBrOC5S) |
| 22 | [▪ Protection against Source Inference Attacks in Federated Learning (Athanasiou, Jung, and Palamidessi, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02017&sa=D&source=editors&ust=1779048535953907&usg=AOvVaw0G8aN-dcnGTTe4xmMphnuk) |
| 23 | [▪ No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings (Jeon et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.22689&sa=D&source=editors&ust=1779048535954038&usg=AOvVaw3h3VK8aWWBnxgXktHZRPnA) |
| 24 | [▪ ImpMIA: Leveraging Implicit Bias for Membership Inference Attack (Golbari et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10625&sa=D&source=editors&ust=1779048535954147&usg=AOvVaw2gZCNduee2E9Moo205OY_f) |
| 25 | [▪ LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings (Oksuz, Halimi, and Ayday, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18934&sa=D&source=editors&ust=1779048535954278&usg=AOvVaw1MWNtU8PdnRMggsH6_p2r9) |
| 26 | [▪ Sequential Membership Inference Attacks (Michel, Basu, and Kaufmann, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.16596&sa=D&source=editors&ust=1779048535954393&usg=AOvVaw0gLZtLWZ1AGPwI7LOsYKnk) |
| 27 | [▪ The Role of Learning in Attacking Intrusion Detection Systems (Domico, Ferrand, and McDaniel, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10299&sa=D&source=editors&ust=1779048535954508&usg=AOvVaw1AMj2Kw0xRAen2zsFARNBx) |
| 28 | [▪ Practical Feasibility of Gradient Inversion Attacks in Federated Learning (Valadi et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.19819&sa=D&source=editors&ust=1779048535954633&usg=AOvVaw26GTPBdPANi_u08dwSrBH1) |
| 29 | [▪ Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks (Tsai et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.07090&sa=D&source=editors&ust=1779048535954751&usg=AOvVaw2ukq8xnl2tdk_rt8aRW-_Z) |
| 30 | [▪ Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software (Kordonsky et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04894&sa=D&source=editors&ust=1779048535954870&usg=AOvVaw29_Qqz3vmjJ_k7nbb9dPFF) |
| 31 | [▪ Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models (Wu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12340&sa=D&source=editors&ust=1779048535954971&usg=AOvVaw3o9whDYOaZiQN5u8p_ITp2) |
| 32 | [▪ Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective (Ganev and Cristofaro, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22434&sa=D&source=editors&ust=1779048535955082&usg=AOvVaw1bbLXUwV-bY8j4O4q1hOKU) |
| 33 | [▪ Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise (Lian et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21628&sa=D&source=editors&ust=1779048535955187&usg=AOvVaw3cIM9n8rsslRNhzIgvgIbC) |
| 34 | [▪ What Hard Tokens Reveal: Exploiting Low-confidence Tokens for Membership Inference Attacks against Large Language Models (Jawad, Xiao, and Wu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20885&sa=D&source=editors&ust=1779048535955296&usg=AOvVaw1yyNVP1z8bw1nsV2T7Rtis) |
| 35 | [▪ UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion (Xue et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20325&sa=D&source=editors&ust=1779048535955395&usg=AOvVaw0k0CfqOD72iLQ1A5mtYptn) |
| 36 | [▪ VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks (Tsaprazlis et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.18413&sa=D&source=editors&ust=1779048535955496&usg=AOvVaw2PTC01GJufY_KBguEB1HgI) |
| 37 | [▪ How does Graph Structure Modulate Membership-Inference Risk for Graph Neural Networks? (Khosla, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17130&sa=D&source=editors&ust=1779048535955598&usg=AOvVaw1NCIfJ1w1K5QdvIFmr98hI) |
| 38 | [▪ Reconstructing Training Data from Adapter-based Federated Large Language Models (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17533&sa=D&source=editors&ust=1779048535955698&usg=AOvVaw0f4F2vvtfUV9ImMpNDP8qQ) |
| 39 | [▪ Res-MIA: A Training-Free Resolution-Based Membership Inference Attack on Federated Learning Models (Zare and Shamsinejadbabaki, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.17378&sa=D&source=editors&ust=1779048535955801&usg=AOvVaw3uFiWZ_3CrxCsv8JhOJuc1) |
| 40 | [▪ Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning (Zhou et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.15595&sa=D&source=editors&ust=1779048535955901&usg=AOvVaw2fVXNep7ZU5B7NKsu1-dZm) |
| 41 | [▪ Powerful Training-Free Membership Inference Against Autoregressive Language Models (Ili\\'c, Stanojevi\\'c, and Cvejoski, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.12104&sa=D&source=editors&ust=1779048535956004&usg=AOvVaw0an6gxxXk54qsAjrWC_9SP) |
| 42 | [▪ When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models (Hu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13607&sa=D&source=editors&ust=1779048535956111&usg=AOvVaw2edPVmZgx0SX1u5Gk9EWFO) |
| 43 | [▪ Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks (Guo et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2503.11514&sa=D&source=editors&ust=1779048535956210&usg=AOvVaw3qVkt39zgW5apV9g0rhQRU) |
| 44 | [▪ DeepLeak: Privacy Enhancing Hardening of Model Explanations Against Membership Leakage (Hmida et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03429&sa=D&source=editors&ust=1779048535956308&usg=AOvVaw3sg0YiYlACWCoElLKrQz03) |
| 45 | [▪ Window-based Membership Inference Attacks Against Fine-tuned Large Language Models (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02751&sa=D&source=editors&ust=1779048535956408&usg=AOvVaw0gQWWHOdl0rkaCnA9zndR5) |
| 46 | [▪ Assessing the Effectiveness of Membership Inference on Generative Music (Chow et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21762&sa=D&source=editors&ust=1779048535956506&usg=AOvVaw2rtxxDqHgIaEAW3AULe10K) |
| 47 | [▪ Membership Inference Attack with Partial Features (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.06244&sa=D&source=editors&ust=1779048535956626&usg=AOvVaw0Z_54q5cQjHgRKBo6m7PiI) |
| 48 | [▪ Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks (Fan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18264&sa=D&source=editors&ust=1779048535956769&usg=AOvVaw1O9W6345gXnxLT7IKJbHJ-) |
| 49 | [▪ In-Context Probing for Membership Inference in Fine-Tuned Language Models (Lu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16292&sa=D&source=editors&ust=1779048535956888&usg=AOvVaw0GdXKt4xoKJMP4sUUzNEin) |
| 50 | [▪ How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code? (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15468&sa=D&source=editors&ust=1779048535956997&usg=AOvVaw0UvMrlRY-F-CLzz26p2p5D) |
| 51 | [▪ An Efficient Gradient-Based Inference Attack for Federated Learning (Monta\\~na-Fern\\'andez and Ortega-Fernandez, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15143&sa=D&source=editors&ust=1779048535957106&usg=AOvVaw0AA_MJFB34yDdZ6IRuY5K2) |
| 52 | [▪ IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol (Yao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14166&sa=D&source=editors&ust=1779048535957216&usg=AOvVaw0hUOcXmBVdKtUBO_lrXLO9) |
| 53 | [▪ Non-Linear Trajectory Modeling for Multi-Step Gradient Inversion Attacks in Federated Learning (Xia et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22082&sa=D&source=editors&ust=1779048535957329&usg=AOvVaw0Wa7zxKIFzZUuBgG_1gB6t) |
| 54 | [▪ On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models (Sahili, Chehab, and Tajeddine, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13352&sa=D&source=editors&ust=1779048535957436&usg=AOvVaw1YdoXeX20ZgPvbyzqVYTb-) |
| 55 | [▪ Imitative Membership Inference Attack (Du et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.06796&sa=D&source=editors&ust=1779048535957553&usg=AOvVaw3zSVf8eHIEb16fkPvJu4ys) |
| 56 | [▪ Reference Recommendation based Membership Inference Attack against Hybrid-based Recommender Systems (Chi et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.09442&sa=D&source=editors&ust=1779048535957667&usg=AOvVaw1WHH1q1y1MmRtxlOwgdLfj) |
| 57 | [▪ Unlearning Inversion Attacks for Graph Neural Networks (Zhang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00808&sa=D&source=editors&ust=1779048535957790&usg=AOvVaw1bLDJs9Pm3A3x07NYV-0fS) |
| 58 | [▪ Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models (Tong, Sun, and Nguyen, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03121&sa=D&source=editors&ust=1779048535957905&usg=AOvVaw3YpDglvNPVSbR5Z36jDLPl) |
| 59 | [▪ ICAS: Detecting Training Data from Autoregressive Image Generative Models (Yu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05068&sa=D&source=editors&ust=1779048535958013&usg=AOvVaw31XTO3jmkKATQOasrx4vqE) |
| 60 | [▪ Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filtering (Yu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.05814&sa=D&source=editors&ust=1779048535958147&usg=AOvVaw3wv8eJWv9WZck81XaMVcEU) |
| 61 | [▪ Ghosting Your LLM: Without The Knowledge of Your Gradient and Data (Almalky et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.22700&sa=D&source=editors&ust=1779048535958243&usg=AOvVaw3dbU_O1B43SbUll8kAXUgp) |
| 62 | [▪ Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage? (Amebley and Dibbo, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.20710&sa=D&source=editors&ust=1779048535958340&usg=AOvVaw1i-LzBEdb_UblGOmB8QHs_) |
| 63 | [▪ Do Spikes Protect Privacy? Investigating Black-Box Model Inversion Attacks in Spiking Neural Networks (Poursiami, Moshruba, and Parsa, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05509&sa=D&source=editors&ust=1779048535958439&usg=AOvVaw0Y9BW6XExvBqry9f-ArsVK) |
| 64 | [▪ Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.05034&sa=D&source=editors&ust=1779048535958538&usg=AOvVaw1Ek4bbXgNQL9noEsGdm0Sb) |
| 65 | [▪ GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14045&sa=D&source=editors&ust=1779048535958644&usg=AOvVaw0f28cD_yrtV8PrERuC-DzI) |
| 66 | [▪ GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning (Song and Palanisamy, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10936&sa=D&source=editors&ust=1779048535958740&usg=AOvVaw1GnqJoE6fDTky5jhMF7jfc) |
| 67 | [▪ On the Detectability of Active Gradient Inversion Attacks in Federated Learning (Carletti et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10502&sa=D&source=editors&ust=1779048535958836&usg=AOvVaw04JOl5YYhzwoDsOEefaf01) |
| 68 | [▪ Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models (Meng et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10423&sa=D&source=editors&ust=1779048535958935&usg=AOvVaw1NuJ4yFgXsuShovejeNB4U) |
| 69 | [▪ Safeguarding Graph Neural Networks against Topology Inference Attacks (Fu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05429&sa=D&source=editors&ust=1779048535959062&usg=AOvVaw1cjAUxL4Bji88Aqzlp2QDz) |
| 70 | [▪ Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models (Belfiore, Passerat-Palmbach, and Usynin, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07503&sa=D&source=editors&ust=1779048535959255&usg=AOvVaw35pg3Fbg-UNI3uLUvg2wmn) |
| 71 | [▪ P-MIA: A Profiled-Based Membership Inference Attack on Cognitive Diagnosis Models (Hou et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.04716&sa=D&source=editors&ust=1779048535959361&usg=AOvVaw19B7vSZKICKuGpMgrpq6E7) |
| 72 | [▪ Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing (Yin et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.01952&sa=D&source=editors&ust=1779048535959452&usg=AOvVaw3k8qbyOIaQoffAMO8UhHJ9) |
| 73 | [▪ TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion (Tang, Jiang, and Niu, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.17302&sa=D&source=editors&ust=1779048535959537&usg=AOvVaw0HOBhBdFjIrTGtl_oa21o9) |
| 74 | [▪ Model Inversion Attacks Meet Cryptographic Fuzzy Extractors (Prabhakar, Xu, and Saxena, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25687&sa=D&source=editors&ust=1779048535959618&usg=AOvVaw1vM5x2AWgu42GN_C7nk-UO) |
| 75 | [▪ Practical Bayes-Optimal Membership Inference Attacks (Lassila et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.24089&sa=D&source=editors&ust=1779048535959701&usg=AOvVaw3GqvzZsj7Wc1ZvQy6o2YVY) |
| 76 | [▪ SPEAR++: Scaling Gradient Inversion via Sparsely-Used Dictionary Learning (Bakarsky et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.24200&sa=D&source=editors&ust=1779048535959784&usg=AOvVaw3Ujum39yk3vWA3Qlbkqprq) |
| 77 | [▪ Membership Inference Attacks for Unseen Classes (Thaker et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06488&sa=D&source=editors&ust=1779048535959863&usg=AOvVaw39Iqp-ohR5p73bPgDFqP3S) |
| 78 | [▪ Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models (Li, Yu, and Xu, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21783&sa=D&source=editors&ust=1779048535959945&usg=AOvVaw2xL_OZMfmQclqoNLyzxQyf) |
| 79 | [▪ Fast-MIA: Efficient and Scalable Membership Inference for LLMs (Takahashi and Ishihara, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23074&sa=D&source=editors&ust=1779048535960025&usg=AOvVaw3gnCsI2lFoJbST8nKXIfU8) |
| 80 | [▪ GUIDE: Enhancing Gradient Inversion Attacks in Federated Learning with Denoising Models (Carletti et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17621&sa=D&source=editors&ust=1779048535960107&usg=AOvVaw0W03EU3_Y4tlBWHuWKK7k5) |
| 81 | [▪ Detecting Adversarial Fine-tuning with Auditing Agents (Egler, Schulman, and Carlini, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16255&sa=D&source=editors&ust=1779048535960197&usg=AOvVaw2qwDJc4FN0yBVIaHIAaST9) |
| 82 | [▪ Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts (Bai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13451&sa=D&source=editors&ust=1779048535960282&usg=AOvVaw2j_Q2KzAeibjjKJuslyN61) |
| 83 | [▪ ImpMIA: Leveraging Implicit Bias for Membership Inference Attack under Realistic Scenarios (Golbari et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.10625&sa=D&source=editors&ust=1779048535960363&usg=AOvVaw3aNqsh1D8Jaydqrw3cahdi) |
| 84 | [▪ DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.18015&sa=D&source=editors&ust=1779048535960446&usg=AOvVaw0wrzb2DkeQXUuvxvpzX9BF) |
| 85 | [▪ Diffusion-aided Task-oriented Semantic Communications with Model Inversion Attack (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.19886&sa=D&source=editors&ust=1779048535960546&usg=AOvVaw3mcLasWj2pHPzCJ9KX5IpF) |
| 86 | [▪ Uncovering Privacy Vulnerabilities through Analytical Gradient Inversion Attacks (Eltaras et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.18871&sa=D&source=editors&ust=1779048535960638&usg=AOvVaw0yvPef0ZM7CAIGFbKXrNad) |
| 87 | [▪ Unveiling Impact of Frequency Components on Membership Inference Attacks for Diffusion Models (Lian et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.20955&sa=D&source=editors&ust=1779048535960729&usg=AOvVaw0Y0okI3tpREUmRUivG4Cmh) |
| 88 | [▪ Accurate Latent Inversion for Generative Image Steganography via Rectified Flow (Qian et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00434&sa=D&source=editors&ust=1779048535960820&usg=AOvVaw0Pp6O_YL1rLbgNgSLWwM7O) |
| 89 | [▪ An Inversion-based Measure of Memorization for Diffusion Models (Ma et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.05846&sa=D&source=editors&ust=1779048535960906&usg=AOvVaw1vgYGPy9B_30Ve5vrRrurf) |
| 90 | [▪ MASQUE: A Text-Guided Diffusion-Based Framework for Localized and Customized Adversarial Makeup (Kwon and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.10549&sa=D&source=editors&ust=1779048535960998&usg=AOvVaw0Ai3DeChs0K0PSDQTyouwL) |
| 91 | [▪ Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition (Nguyen, \\v{S}rndi\\'c, and Neth, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2401.10337&sa=D&source=editors&ust=1779048535961094&usg=AOvVaw127v7_5wZGLXmpSSU53Rr_) |
| 92 | [▪ Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based Priors (Ma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17577&sa=D&source=editors&ust=1779048535961196&usg=AOvVaw2cv58EN0Camot_ClgPclVB) |
| 93 | [▪ Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs (German et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17259&sa=D&source=editors&ust=1779048535961278&usg=AOvVaw1bfPcBJk5bECvJX3VWFNX-) |
| 94 | [▪ Depth Gives a False Sense of Privacy: LLM Internal States Inversion (Dong et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16372&sa=D&source=editors&ust=1779048535961355&usg=AOvVaw28zK5XBZWJvcd_PxTQS1wK) |
| 95 | [▪ Blackbox Dataset Inference for LLM (Zhou et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.03619&sa=D&source=editors&ust=1779048535961430&usg=AOvVaw2Bj-l7bYQUtfVjNAnRWBSt) |
| 96 | [▪ Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems (Tomashenko, Vincent, and Tommasi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15214&sa=D&source=editors&ust=1779048535961511&usg=AOvVaw14V0_axaNsUBpFzn5YA2Gn) |
| 97 | [▪ REAL-IoT: Characterizing GNN Intrusion Detection Robustness under Practical Adversarial Attack (Zhan, Zhou, and Haddadi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10836&sa=D&source=editors&ust=1779048535961592&usg=AOvVaw28WG4yKODDuYt1YpNvkTVV) |
| 98 | [▪ Random Erasing vs. Model Inversion: A Promising Defense or a False Hope? (Tran et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.01062&sa=D&source=editors&ust=1779048535961670&usg=AOvVaw1Ima_N9kny-jqdVJoL32DG) |
| 99 | [▪ AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical Perspective (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09857&sa=D&source=editors&ust=1779048535961752&usg=AOvVaw2Dwyv7XWr4BLlfQlSWYL6u) |
| 100 | [▪ Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy (Kulynych et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06969&sa=D&source=editors&ust=1779048535961834&usg=AOvVaw2NNYJpWr6wzc09nAEfCifR) |
| 101 | [▪ Detection of Intelligent Tampering in Wireless Electrocardiogram Signals Using Hybrid Machine Learning (Deshpande, Getnet, and Dargie, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06402&sa=D&source=editors&ust=1779048535961916&usg=AOvVaw0VN1fT1jslIK4_X8hFdEii) |
| 102 | [▪ Deep Learning Model Inversion Attacks and Defenses: A Comprehensive Survey (Yang et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.18934&sa=D&source=editors&ust=1779048535962013&usg=AOvVaw2zoeu_iLpwN7KjZAh2TTpu) |
| 103 | [▪ Trap-MID: Trapdoor-based Defense against Model Inversion Attacks (Liu and Chen, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.08460&sa=D&source=editors&ust=1779048535962101&usg=AOvVaw0mm1J-tPXjCVUd6WzTLUHx) |
| 104 | [▪ Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning (Shan et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14937&sa=D&source=editors&ust=1779048535962195&usg=AOvVaw2hkWvkhvuYX0334zWrR3uO) |
| 105 | [▪ MIBench: A Comprehensive Benchmark for Model Inversion Attack and Defense (Qiu et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.05159&sa=D&source=editors&ust=1779048535962284&usg=AOvVaw2Ju1OYgJIAgtJMKPv_wJ2g) |
| 106 | [▪ Defending against Model Inversion Attacks via Random Erasing (Tran et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.01062&sa=D&source=editors&ust=1779048535962369&usg=AOvVaw2w_q5LuxyEvFu-LxIROBFH) |
| 107 | [▪ Improving Robustness to Model Inversion Attacks via Sparse Coding Architectures (Dibbo et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.14772&sa=D&source=editors&ust=1779048535962456&usg=AOvVaw1kn69XDQOVvWUllJHn4arl) |

|     |
| --- |
| Model Inversion Attack |

**>**

**<**

‍

#### Exfiltration via Cyber Means

Covers:

- MITRE ATLAS Exfiltration

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Exfiltration via Cyber Means

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ VectorSmuggle: Steganographic Exfiltration in Embedding Stores and a Cryptographic Provenance Defense (Wanger, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13764&sa=D&source=editors&ust=1779048536726994&usg=AOvVaw0Z7adDWdxX3MhIYNxbwAGI) |
| 3 | [▪ DECIFR: Domain-Aware Exfiltration of Circuit Information from Federated Gradient Reconstruction (Lee et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19915&sa=D&source=editors&ust=1779048536727401&usg=AOvVaw3tCyh8etLBIhSx9ewMR7wG) |
| 4 | [▪ Improving DNS Exfiltration Detection via Transformer Pretraining (Tomi\\'c, Cvetanovi\\'c, and Tadi\\'c, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09849&sa=D&source=editors&ust=1779048536727539&usg=AOvVaw2kC7vCNefsKQWjTfH1bJ7C) |
| 5 | [▪ FLARE: A Wireless Side-Channel Fingerprinting Attack on Federated Learning (Shuvo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10296&sa=D&source=editors&ust=1779048536727639&usg=AOvVaw2qhKE65qt6GCLzLT8cPoWD) |
| 6 | [▪ Malicious GenAI Chrome Extensions: Unpacking Data Exfiltration and Malicious Behaviours (Seetharam, Nabeel, and Melicher, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10029&sa=D&source=editors&ust=1779048536727734&usg=AOvVaw2s8U3ngX94edfVBvVN0vJ5) |
| 7 | [▪ Data Exfiltration by Compression Attack: Definition and Evaluation on Medical Image Data (Li, Ayache, and Delingette, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21227&sa=D&source=editors&ust=1779048536727832&usg=AOvVaw1ug8GXTYVKNvO4E14FU7Wv) |
| 8 | [▪ Exploiting Web Search Tools of AI Agents for Data Exfiltration (Rall et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09093&sa=D&source=editors&ust=1779048536727958&usg=AOvVaw3YCJXzn7X0vZWNZ0DKJct3) |

|     |
| --- |
| Exfiltration via Cyber Means |

**>**

**<**

‍

#### Model Skewing Attack

Covers:

- OWASP ML 08: Model Skewing

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Model Skewing Attack

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models (Hartman et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23238&sa=D&source=editors&ust=1779048535880989&usg=AOvVaw3Svg_gmMjJlkgT2-vP_rms) |
| 3 | [▪ Conflicts Make Large Reasoning Models Vulnerable to Attacks (Liu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.09750&sa=D&source=editors&ust=1779048535881140&usg=AOvVaw3usJML11LH4LWk6ByNtnsr) |
| 4 | [▪ Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models (Liu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2308.07847&sa=D&source=editors&ust=1779048535881213&usg=AOvVaw2hGKzdp2IeL3jJXvYVsJxI) |
| 5 | [▪ Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks (Struppek, Gleave, and Pelrine, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14689&sa=D&source=editors&ust=1779048535881269&usg=AOvVaw0H7MaDhJtlMt5Wp_8bGoM3) |
| 6 | [▪ Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges (Ding et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13576&sa=D&source=editors&ust=1779048535881323&usg=AOvVaw1F9MzNdm1LDUhONP-FPnQe) |
| 7 | [▪ Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs (Jiang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08621&sa=D&source=editors&ust=1779048535881373&usg=AOvVaw26QhNwcQpvrxw1sVv1NCsI) |
| 8 | [▪ From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.23088&sa=D&source=editors&ust=1779048535881420&usg=AOvVaw2QfJfNBni0qIUi4LY17sMI) |
| 9 | [▪ Beyond Denial-of-Service: The Puppeteer's Attack for Fine-Grained Control in Ranking-Based Federated Learning (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14687&sa=D&source=editors&ust=1779048535881468&usg=AOvVaw3FTvKzVQhVg6ViV0IRLh0W) |
| 10 | [▪ Breaking Diffusion with Cache: Exploiting Approximate Caches in Diffusion Models (Sun, Jie, and Liu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20424&sa=D&source=editors&ust=1779048535881518&usg=AOvVaw0HNzHCZkmXPX8M4iL51OGB) |
| 11 | [▪ SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models (Djuhera et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00062&sa=D&source=editors&ust=1779048535881569&usg=AOvVaw1olF0aDfk865jwmuTcW4NR) |
| 12 | [▪ SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification (Lai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17368&sa=D&source=editors&ust=1779048535881641&usg=AOvVaw1KzX0sqDAx4lH8SiaYh9Rt) |
| 13 | [▪ Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (Gong et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.11611&sa=D&source=editors&ust=1779048535881705&usg=AOvVaw1jH2SEGGO8OxpIziGEvWft) |
| 14 | [▪ Understanding Concept Drift with Deprecated Permissions in Android Malware Detection (Sabbah et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22231&sa=D&source=editors&ust=1779048535881760&usg=AOvVaw1B6pYI7pBwh0ehzlA1sAmO) |
| 15 | [▪ A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census (Abowd et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.11283&sa=D&source=editors&ust=1779048535881810&usg=AOvVaw2HMBbRwwZ3XO2uy1gYsH8V) |
| 16 | [▪ HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning (He and Chang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10162&sa=D&source=editors&ust=1779048535881859&usg=AOvVaw2VVARbWxKLVPjQhYn1prHl) |
| 17 | [▪ SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents (Kutasov et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15740&sa=D&source=editors&ust=1779048535881909&usg=AOvVaw35T8_wIHv3WfFnD2opIzsT) |
| 18 | [▪ False Alarms, Real Damage: Adversarial Attacks Using LLM-based Models on Text-based Cyber Threat Intelligence Systems (Shafee, Bessani, and Ferreira, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06252&sa=D&source=editors&ust=1779048535881956&usg=AOvVaw0GrR09eDK01HMwgQmX_ywJ) |

|     |
| --- |
| Model Skewing Attack |

**>**

**<**

‍

#### Evade ML Model

Covers:

- MITRE ATLAS Initial Access
- MITRE ATLAS Reconnaissance

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Evade ML Model

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ The Role of Learning in Attacking ML-based Network Intrusion Detection (Domico, Ferrand, and McDaniel, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10299&sa=D&source=editors&ust=1779048536989038&usg=AOvVaw3QNP56TzHbXRJWvRrru3dW) |
| 3 | [▪ WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation (Liu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.08310&sa=D&source=editors&ust=1779048536989213&usg=AOvVaw0ImOjWdcCWPom_s7gU8Z4X) |
| 4 | [▪ Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption (Morales, Pastrana, and Tapiador, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06910&sa=D&source=editors&ust=1779048536989308&usg=AOvVaw19yHfe7Oz-_Sckb3I_VEwn) |
| 5 | [▪ Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00305&sa=D&source=editors&ust=1779048536989398&usg=AOvVaw1dQdldcEtjuPFFEyrqjgdT) |
| 6 | [▪ MalPurifier: Enhancing Android Malware Detection with Adversarial Purification against Evasion Attacks (Zhou et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2312.06423&sa=D&source=editors&ust=1779048536989484&usg=AOvVaw14jDdYr8udG6qK_bpdAT1j) |
| 7 | [▪ Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting (Ediga and Chattopadhyay, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01186&sa=D&source=editors&ust=1779048536989564&usg=AOvVaw0IqTs4hgEzzT9uIPKXM0il) |
| 8 | [▪ Trident: Improving Malware Detection with LLMs and Behavioral Features (Saul et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00297&sa=D&source=editors&ust=1779048536989643&usg=AOvVaw3RepR-47P1f3p8pyz1661V) |
| 9 | [▪ AsmRAG: LLM-Driven Malware Detection by Retrieving Functionally Similar Assembly Code (Karbab, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23196&sa=D&source=editors&ust=1779048536989722&usg=AOvVaw1VhamfwSfng8OXBhD4QysP) |
| 10 | [▪ Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations (Acharya and Zhang, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.21310&sa=D&source=editors&ust=1779048536989804&usg=AOvVaw1Qe5Bv9pLOMXH7c12Mnkeo) |
| 11 | [▪ Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing (Peng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05719&sa=D&source=editors&ust=1779048536989895&usg=AOvVaw2bIlr3DXceqGtIR92kIghy) |
| 12 | [▪ Evasion Adversarial Attacks Remain Impractical Against ML-based Network Intrusion Detection Systems, Especially Dynamic Ones (elShehaby and Matrawy, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2306.05494&sa=D&source=editors&ust=1779048536989986&usg=AOvVaw36aK4LM0xP-N05aupqtyfc) |
| 13 | [▪ ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks (Elnawawy et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26093&sa=D&source=editors&ust=1779048536990096&usg=AOvVaw1WHb7CjBrRz3JpdEKrM9lN) |
| 14 | [▪ Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25500&sa=D&source=editors&ust=1779048536990197&usg=AOvVaw0bmCoXa5qB-z-3sQaU7LSW) |
| 15 | [▪ Targeted Adversarial Traffic Generation : Black-box Approach to Evade Intrusion Detection Systems in IoT Networks (Debicha et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23438&sa=D&source=editors&ust=1779048536990281&usg=AOvVaw1G9BtREylEo0-StZsBybpt) |
| 16 | [▪ Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents (Aonzo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.15457&sa=D&source=editors&ust=1779048536990363&usg=AOvVaw3hps70BdPXf7me0etIXgKx) |
| 17 | [▪ FuzzySQL: Uncovering Hidden Vulnerabilities in DBMS Special Features with LLM-Driven Fuzzing (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.19490&sa=D&source=editors&ust=1779048536990449&usg=AOvVaw1huxi8Zilv5g65Y8RNqcsF) |
| 18 | [▪ Anticipating Adversary Behavior in DevSecOps Scenarios through Large Language Models (Caballero et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14106&sa=D&source=editors&ust=1779048536990528&usg=AOvVaw1rFUNeSKXXlrF_gu2sU2bR) |
| 19 | [▪ SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization (Zhang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.16981&sa=D&source=editors&ust=1779048536990609&usg=AOvVaw28UFfELVqFbLuiitRKLz5w) |
| 20 | [▪ Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach (Ali et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.12183&sa=D&source=editors&ust=1779048536990689&usg=AOvVaw3Z0y9usxslR10t9DSJHi98) |
| 21 | [▪ StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors (Ranganath and Ramesh, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08934&sa=D&source=editors&ust=1779048536990777&usg=AOvVaw2AdP1bsynDkM0vhomdLNgG) |
| 22 | [▪ DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems (Ou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02569&sa=D&source=editors&ust=1779048536990869&usg=AOvVaw3p7SPXJB5lAN-nF421WUS1) |
| 23 | [▪ Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks (Zhang and Zhang, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23101&sa=D&source=editors&ust=1779048536990954&usg=AOvVaw02-FWNG1PH1rBWU90iWUnS) |
| 24 | [▪ "Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00364&sa=D&source=editors&ust=1779048536991035&usg=AOvVaw17Bi7udqoBUtKHJFv9sHtH) |
| 25 | [▪ Semantics-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.00305&sa=D&source=editors&ust=1779048536991113&usg=AOvVaw2mrWlkWCO1dfHnqvz4SWOp) |
| 26 | [▪ In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement (Shetty, Joshi, and Kanhere, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22169&sa=D&source=editors&ust=1779048536991193&usg=AOvVaw3-nCzYfzNMZKiDJKTGjWvp) |
| 27 | [▪ AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises (Tung et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22720&sa=D&source=editors&ust=1779048536991276&usg=AOvVaw1C-Q7EcTuXW5FQM2wOn9r8) |
| 28 | [▪ ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses (He et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21586&sa=D&source=editors&ust=1779048536991389&usg=AOvVaw2c60THwbKmHGqwhAunwaDD) |
| 29 | [▪ CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13112&sa=D&source=editors&ust=1779048536991478&usg=AOvVaw0yYWhCDutaVEmfz7GRENHS) |
| 30 | [▪ A Decompilation-Driven Framework for Malware Detection with Large Language Models (Chawla and Prasad, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09035&sa=D&source=editors&ust=1779048536991565&usg=AOvVaw2lpU7D8jS6LfVN2xdObO5c) |
| 31 | [▪ MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization (Gu, Li, and Hu, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.08564&sa=D&source=editors&ust=1779048536991645&usg=AOvVaw14UaeP6ZEELp-v0emm6zaC) |
| 32 | [▪ VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference (Qi et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06461&sa=D&source=editors&ust=1779048536991723&usg=AOvVaw3vvXl6eq6nJqklAPE8t9E1) |
| 33 | [▪ Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools? (Quantrill et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.00559&sa=D&source=editors&ust=1779048536991819&usg=AOvVaw1U9pt08MV9dkFs1X6tjSyN) |
| 34 | [▪ Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks (Xu et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.21241&sa=D&source=editors&ust=1779048536991918&usg=AOvVaw0MCA5nOw5nX9e98hgr6Hu5) |
| 35 | [▪ Automated Penetration Testing with LLM Agents and Classical Planning (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11143&sa=D&source=editors&ust=1779048536992009&usg=AOvVaw3b_EbMl6M6A2USgteyJ7k2) |
| 36 | [▪ NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00119&sa=D&source=editors&ust=1779048536992098&usg=AOvVaw3hEnXtVNuYkJAF6HN0adLo) |
| 37 | [▪ ExtendAttack: Attacking Servers of LRMs via Extending Reasoning (Zhu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.13737&sa=D&source=editors&ust=1779048536992184&usg=AOvVaw1vvim2uXbReKs2a1IPtMYS) |
| 38 | [▪ Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning (Domico et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.01734&sa=D&source=editors&ust=1779048536992271&usg=AOvVaw2118KClYPbmvzi2AxmVFQy) |
| 39 | [▪ MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification (Luo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14129&sa=D&source=editors&ust=1779048536992377&usg=AOvVaw26H_yFebjQQWpYmj0b7-8T) |
| 40 | [▪ Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment (Zhang and sun, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06852&sa=D&source=editors&ust=1779048536992471&usg=AOvVaw16EM5espKPQws6JJxTiA5d) |
| 41 | [▪ GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors (Meng et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.08188&sa=D&source=editors&ust=1779048536992549&usg=AOvVaw0JHHUejB23qNQAp7pN4_S2) |
| 42 | [▪ The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems (Choi et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.20995&sa=D&source=editors&ust=1779048536992629&usg=AOvVaw3oveqlvt32lYs5Ixzgewdl) |
| 43 | [▪ Adversarial Pre-Padding: Generating Evasive Network Traffic Against Transformer-Based Classifiers (Jing et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25810&sa=D&source=editors&ust=1779048536992707&usg=AOvVaw2R_LI1TVoSuse-5xLynJRo) |
| 44 | [▪ Detecting Various DeFi Price Manipulations with LLM Reasoning (Zhong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.11521&sa=D&source=editors&ust=1779048536992784&usg=AOvVaw2BEpYvd9PLH3VaFGoH8yGS) |
| 45 | [▪ Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks (Feng and Sakurai, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23313&sa=D&source=editors&ust=1779048536992863&usg=AOvVaw16WFeHDDq9P2xSfrcYLc9O) |
| 46 | [▪ Black-Box Evasion Attacks on Data-Driven Open RAN Apps: Tailored Design and Experimental Evaluation (Gajjar et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18160&sa=D&source=editors&ust=1779048536992946&usg=AOvVaw1wBYxkxSvruJ8JWaGsmnVZ) |
| 47 | [▪ From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15 (Rehman et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17883&sa=D&source=editors&ust=1779048536993028&usg=AOvVaw1gE3Y609ARDtaG9w6bCEDK) |
| 48 | [▪ SoK: Adversarial Evasion Attacks Practicality in NIDS Domain and the Impact of Dynamic Learning (elShehaby and Matrawy, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2306.05494&sa=D&source=editors&ust=1779048536993127&usg=AOvVaw3FU2bh_Su-odsmfzs71DEh) |
| 49 | [▪ A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14906&sa=D&source=editors&ust=1779048536993252&usg=AOvVaw1mkLi6xPcUGyh5GrJDV3Uc) |
| 50 | [▪ VulSolver: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00882&sa=D&source=editors&ust=1779048536993340&usg=AOvVaw2JHtyqmQ_bMZpDZnfZ_VZ7) |
| 51 | [▪ Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework (Han et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.13811&sa=D&source=editors&ust=1779048536993437&usg=AOvVaw3zRu1a0auo0xVvT2_YiHCI) |
| 52 | [▪ Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks (Engineering et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.11711&sa=D&source=editors&ust=1779048536993531&usg=AOvVaw1cNHVJVuFtcxTawdp4DrNy) |
| 53 | [▪ Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection (Creo, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01887&sa=D&source=editors&ust=1779048536993621&usg=AOvVaw3n3S6owu3gKE1Tuuyxwx76) |
| 54 | [▪ AdVAR-DNN: Adversarial Misclassification Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01107&sa=D&source=editors&ust=1779048536993719&usg=AOvVaw0mshshqjslGr8DrNg5SMZi) |
| 55 | [▪ ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models (Yook et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21985&sa=D&source=editors&ust=1779048536993798&usg=AOvVaw0b_Xdawk3xln7EFOf8Wdh_) |
| 56 | [▪ Hierarchical Graph Neural Network for Compressed Speech Steganalysis (Hemis et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21591&sa=D&source=editors&ust=1779048536993878&usg=AOvVaw1CXQXuiuZBsNw_RgSpRmw2) |
| 57 | [▪ Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms (Sato, Ohki, and Nishigaki, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21538&sa=D&source=editors&ust=1779048536993981&usg=AOvVaw0c9Q3hlfG1rHBMHN8TXFRk) |
| 58 | [▪ GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI (Kalyanapu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17033&sa=D&source=editors&ust=1779048536994061&usg=AOvVaw1fK-kBxjVtQol3vCFZEA3b) |
| 59 | [▪ BandFuzz: An ML-powered Collaborative Fuzzing Framework (Shi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10845&sa=D&source=editors&ust=1779048536994139&usg=AOvVaw2B2KGlhe4AgUDikO10TsVO) |
| 60 | [▪ PotentRegion4MalDetect: Advanced Features from Potential Malicious Regions for Malware Detection (Koppanati, Santra, and Peddoju, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06723&sa=D&source=editors&ust=1779048536994221&usg=AOvVaw0rsp98zm849VRsctT96iHX) |
| 61 | [▪ Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors (Wang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19230&sa=D&source=editors&ust=1779048536994296&usg=AOvVaw1q44qWv2XtedKtXIE94HAC) |

|     |
| --- |
| Evade ML Model |

**>**

**<**

‍

#### Discover ML Artifacts, Data from Information Repositories and Local System, and Acquire Public ML Artifacts

Covers:

- MITRE ATLAS Resource Development, Discovery, and Collection

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Discover ML Model Family and Ontology/Model Extraction

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies (Cotti et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.00081&sa=D&source=editors&ust=1779048536928355&usg=AOvVaw3onj_W_yQfmCXZhcTeY8CZ) |
| 3 | [▪ Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs (Liu, Evans, and Xiong, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.18697&sa=D&source=editors&ust=1779048536928580&usg=AOvVaw0pduBRihDBfu3sQcFIXvgY) |
| 4 | [▪ CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models (Lee, Koo, and Kwak, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.16363&sa=D&source=editors&ust=1779048536928747&usg=AOvVaw0Lg_L05KH4_D0pL3Mfr0dH) |
| 5 | [▪ AttnDiff: Attention-based Differential Fingerprinting for Large Language Models (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05502&sa=D&source=editors&ust=1779048536928879&usg=AOvVaw0H-66EEPWd3Wh9IeT6jVvD) |
| 6 | [▪ Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test (Zhu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06975&sa=D&source=editors&ust=1779048536929005&usg=AOvVaw3yIP_0N8Gs75MUWLH3bIGE) |
| 7 | [▪ Navigating the Deep: End-to-End Extraction on Deep Neural Networks (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17047&sa=D&source=editors&ust=1779048536929139&usg=AOvVaw1EQjYqCIc-3guq-ucAffQm) |
| 8 | [▪ A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors (Xu and Sheng, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09434&sa=D&source=editors&ust=1779048536929250&usg=AOvVaw1lKx29Fl8gkVQVTeliTQxu) |
| 9 | [▪ FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting (Fu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.16029&sa=D&source=editors&ust=1779048536929355&usg=AOvVaw0aoRWB79fpVdN4Xa8nO_Q_) |
| 10 | [▪ Identifying Models Behind Text-to-Image Leaderboards (Naseh et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09647&sa=D&source=editors&ust=1779048536929490&usg=AOvVaw1y2DcirMw26lSxw6muvWfQ) |
| 11 | [▪ Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures (Wang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06390&sa=D&source=editors&ust=1779048536929613&usg=AOvVaw3WXBbHhBk32M2SUGDRHYNT) |
| 12 | [▪ A Fingerprint for Large Language Models (Yang and Wu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.01235&sa=D&source=editors&ust=1779048536929715&usg=AOvVaw3MHk8D7FRLb1hAyqCQziid) |
| 13 | [▪ SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting (Zhang and Zheng, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03620&sa=D&source=editors&ust=1779048536929865&usg=AOvVaw1WLt5QMMwFm-vRlj5zjnTK) |
| 14 | [▪ A Systematic Study of Model Extraction Attacks on Graph Foundation Models (Xu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11912&sa=D&source=editors&ust=1779048536930026&usg=AOvVaw1i-e8Eio7TEPRqMlRDK_UE) |
| 15 | [▪ Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution Approach (Li et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07033&sa=D&source=editors&ust=1779048536930209&usg=AOvVaw0gcBAtxl8qX7wnL8b6ejv4) |
| 16 | [▪ Ghost in the Transformer: Tracing LLM Lineage with SVD-Fingerprint (Wang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.06390&sa=D&source=editors&ust=1779048536930433&usg=AOvVaw24BtpbVNBugZJSqh6uDQ8t) |
| 17 | [▪ Structuring Security: A Survey of Cybersecurity Ontologies, Semantic Log Processing, and LLMs Application (Louren\\c{c}o et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16610&sa=D&source=editors&ust=1779048536930943&usg=AOvVaw1fOKFHCTuCcEY8eY-fmc_T) |
| 18 | [▪ MalCVE: Malware Detection and CVE Association Using Large Language Models (Cristea, Molnes, and Li, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15567&sa=D&source=editors&ust=1779048536931235&usg=AOvVaw3m3iREjjg2DyHUGLQ6pkmJ) |
| 19 | [▪ Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation (Shao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06605&sa=D&source=editors&ust=1779048536931475&usg=AOvVaw3K5j3Sr0-PNQ9G-1-8A3yU) |
| 20 | [▪ SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From (Tong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.26404&sa=D&source=editors&ust=1779048536931785&usg=AOvVaw2fkb0UWlEoP6OqC61PXui7) |
| 21 | [▪ LLM-Assisted Model-Based Fuzzing of Protocol Implementations (Huang, Wang, and Zhou, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01750&sa=D&source=editors&ust=1779048536932053&usg=AOvVaw08mfx-pGvGlifGYr_g0zLE) |
| 22 | [▪ PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models (Gasmi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.19185&sa=D&source=editors&ust=1779048536932287&usg=AOvVaw1sFsxUW291kOTdvbrSKwXK) |
| 23 | [▪ Evaluating Ensemble and Deep Learning Models for Static Malware Detection with Dimensionality Reduction Using the EMBER Dataset (Abedin and Mehrub, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16952&sa=D&source=editors&ust=1779048536932503&usg=AOvVaw1jYLZTcvg2MN6q84EMPxtb) |
| 24 | [▪ Revisiting Pre-trained Language Models for Vulnerability Detection (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16887&sa=D&source=editors&ust=1779048536932656&usg=AOvVaw0zCr_m7rfyKSrt3bmScKPR) |
| 25 | [▪ Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications (H\\"arer, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.10467&sa=D&source=editors&ust=1779048536932841&usg=AOvVaw2R_jHHNnC2-Fpb_amzK-tW) |
| 26 | [▪ TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2304.02838&sa=D&source=editors&ust=1779048536933090&usg=AOvVaw0kH2XbiZWb2KnfYIH3uL9_) |
| 27 | [▪ Toward an Intent-Based and Ontology-Driven Autonomic Security Response in Security Orchestration Automation and Response (Huang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12061&sa=D&source=editors&ust=1779048536933341&usg=AOvVaw20sq0li9r_cpmoKqDF8-tk) |
| 28 | [▪ SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents (Masoor, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10562&sa=D&source=editors&ust=1779048536933633&usg=AOvVaw3zEYBYlsqJlbmh_LsIz8tN) |
| 29 | [▪ UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification (Cai et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.12318&sa=D&source=editors&ust=1779048536933824&usg=AOvVaw1FK1U15ozLVxNIPU0XlHIm) |
| 30 | [▪ AICrypto: A Comprehensive Benchmark For Evaluating Cryptography Capabilities of Large Language Models (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09580&sa=D&source=editors&ust=1779048536934020&usg=AOvVaw2wVT5cROx51CqlEM3Monds) |
| 31 | [▪ BISON: Blind Identification with Stateless scOped pseudoNyms (Heher, More, and Heimberger, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.01518&sa=D&source=editors&ust=1779048536934203&usg=AOvVaw1B1E3DRMkUAsKpooyvilrV) |
| 32 | [▪ Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms (Gasmi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06323&sa=D&source=editors&ust=1779048536934406&usg=AOvVaw13abMWYpBUNlPwIQKME6WK) |
| 33 | [▪ From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks (Khouna, Ferry, and Vidal, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.05325&sa=D&source=editors&ust=1779048536934544&usg=AOvVaw3APEn2yLQwGfczXRBFRttv) |
| 34 | [▪ One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs\<br>\<br> (Li et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.17598&sa=D&source=editors&ust=1779048536934686&usg=AOvVaw1jQiUCFMEFrKuY4Vq5WIIH) |
| 35 | [▪ Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries (Wang et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.15420&sa=D&source=editors&ust=1779048536934879&usg=AOvVaw3o09kjKg-ZorD7F0IbbFMq) |
| 36 | [▪ How to Backdoor the Knowledge Distillation (Wu et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21323&sa=D&source=editors&ust=1779048536935022&usg=AOvVaw11oQ2nso_FaqDok2Lqqg3p) |
| 37 | [▪ Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations (Ezzeddine, Ayoub, and Giordano, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21323&sa=D&source=editors&ust=1779048536935152&usg=AOvVaw3zWh5k44hvQ9Cn-d0X4g7R) |
| 38 | [▪ Efficient and Effective Model Extraction (Zhu et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14122&sa=D&source=editors&ust=1779048536935305&usg=AOvVaw0d24IJU50_4Y1MQ4ov4P1J) |
| 39 | [▪ CaBaGe: Data-Free Model Extraction using ClAss BAlanced Generator Ensemble (Rosenthal et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10643&sa=D&source=editors&ust=1779048536935434&usg=AOvVaw15SlwaTL1xIQw_8d_F_aIp) |
| 40 | [▪ Alignment-Aware Model Extraction Attacks on Large Language Models (Liang et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.02718%23&sa=D&source=editors&ust=1779048536935546&usg=AOvVaw2djZQz9Ac1EeAPvADZjowV) |

|     |
| --- |
| Discover ML Model Family and Ontology/Model Extraction |

**>**

**<**

‍

#### User Execution, Command and Scripting Interpreter

Covers:

- MITRE ATLAS Execution

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : User Execution, Command and Scripting Interpreter

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards (Normann et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17673&sa=D&source=editors&ust=1779048535949526&usg=AOvVaw028LSw5durFVZeWsbnJdR-) |
| 3 | [▪ RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents (Nakatani, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.16730&sa=D&source=editors&ust=1779048535949703&usg=AOvVaw3cx0x0xz6R3pEokGlmcBvc) |
| 4 | [▪ CHAI: Command Hijacking against embodied AI (Burbano et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00181&sa=D&source=editors&ust=1779048535949797&usg=AOvVaw0A-UdkkGgE5fId3FkzJxP8) |

|     |
| --- |
| User Execution, Command and Scripting Interpreter |

**>**

**<**

‍

#### Physical Model Access and Full Model Access

Covers:

- MITRE ATLAS ML Model Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Physical Model Access and Full Model Access

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models (Hadad and Guri, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.25403&sa=D&source=editors&ust=1779048536613753&usg=AOvVaw0hE7xD_hyO5llqKplou_Pf) |
| 3 | [▪ T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving (Ma et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12956&sa=D&source=editors&ust=1779048536613881&usg=AOvVaw1cHtAcyc59pY2X5zrDyS3t) |
| 4 | [▪ SleepWalk: Exploiting Context Switching and Residual Power for Physical Side-Channel Attacks (Sanjaya, Jayasena, and Mishra, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22306&sa=D&source=editors&ust=1779048536614011&usg=AOvVaw0CUx0NWoLS1HSMvtprvWtU) |
| 5 | [▪ Rainbow Artifacts from Electromagnetic Signal Injection Attacks on Image Sensors (Zhang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07773&sa=D&source=editors&ust=1779048536614116&usg=AOvVaw07WARZNZlcApilp3psSYMD) |

|     |
| --- |
| Physical Model Access and Full Model Access |

**>**

**<**

‍

#### Valid Accounts

Covers:

- MITRE ATLAS Initial Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Valid Accounts

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |

|     |
| --- |
| Valid Accounts |

**>**

**<**

‍

#### Exploit Public Facing Application

Covers:

- MITRE ATLAS Initial Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Exploit Public Facing Application

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems (Siddiq et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.14163&sa=D&source=editors&ust=1779048536728626&usg=AOvVaw1crsILNtdULG-hSYX1riS0) |

|     |
| --- |
| Exploit Public Facing Application |

**>**

**<**

‍

### Threats from AI Model

#### Misinformation

Covers:

- MITRE ATLAS Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Misinformation

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Day-to-Day Traffic Network Modeling under Route-Guidance Misinformation: Endogenous Trust and Resilience in CAV Environments (Ka and Ukkusuri, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14204&sa=D&source=editors&ust=1779048536739569&usg=AOvVaw099K71Pde0EmVcEfE_mZa1) |
| 3 | [▪ CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-Bunking of Online Misinformation (Loth, Kappes, and Pahl, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.20856&sa=D&source=editors&ust=1779048536739810&usg=AOvVaw0soq5tGf6lIfq5Zu4j8MSY) |
| 4 | [▪ The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation (Chrysidis, Papadopoulos, and Papadopoulos, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.15372&sa=D&source=editors&ust=1779048536739982&usg=AOvVaw1McumsogNAdXyvol0T_5cz) |
| 5 | [▪ Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection (Guan et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10460&sa=D&source=editors&ust=1779048536740130&usg=AOvVaw2jAp1vImv0-fS4VcF2adhD) |
| 6 | [▪ "That's another doom I haven't thought about": A User Study on AI Labels as a Safeguard Against Image-Based Misinformation (H\\"oltervennhoff et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.22845&sa=D&source=editors&ust=1779048536740314&usg=AOvVaw361NuDaVkqQQBlUzlb69Zk) |
| 7 | [▪ FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X (Essahli et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25932&sa=D&source=editors&ust=1779048536740445&usg=AOvVaw1xAdzsb8VdW_zlKnvZa4QX) |
| 8 | [▪ Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks (Aldahoul and Zaki, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.08605&sa=D&source=editors&ust=1779048536740570&usg=AOvVaw1IPUMOkTN16PXhxf-gNI4q) |
| 9 | [▪ Fake or Real: The Impostor Hunt in Texts for Space Operations (Kaczmarek et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13508&sa=D&source=editors&ust=1779048536740724&usg=AOvVaw3fr_UhVfki1b702dduoaky) |

|     |
| --- |
| Misinformation |

**>**

**<**

‍

#### Over Reliance on LLM Outputs and External (Social) Harms

Covers:

- OWASP LLM 09: Overreliance
- MITRE ATLAS Impact

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Over Reliance on LLM Outputs and External (Social) Harms

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ False Security Confidence in Benign LLM Code Generation (Ren, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.17014&sa=D&source=editors&ust=1779048535997160&usg=AOvVaw1eVa8B58qhlN8yF-0nZPIA) |
| 3 | [▪ Measuring and Exploiting Confirmation Bias in LLM-Assisted Security Code Review (Mitropoulos et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18740&sa=D&source=editors&ust=1779048535997571&usg=AOvVaw2TdTt4ev2n9xpulF9zqeYw) |
| 4 | [▪ Can Developers rely on LLMs for Secure IaC Development? (Firouzi, Bhatt, and Ghafari, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03648&sa=D&source=editors&ust=1779048535997751&usg=AOvVaw3CwSRtRHL3gswuzvZHRYHL) |
| 5 | [▪ Exploring the Secondary Risks of Large Language Models (Chen et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12382&sa=D&source=editors&ust=1779048535997924&usg=AOvVaw2bqkU5DkyB4YdV6vj_lI1j) |
| 6 | [▪ Large Language Models Are Unreliable for Cyber Threat Intelligence (Mezzi, Massacci, and Tuma, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.23175&sa=D&source=editors&ust=1779048535998142&usg=AOvVaw1cyXHBW6uWI4FZ714eeF2J) |
| 7 | [▪ Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions (Prakash et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14571&sa=D&source=editors&ust=1779048535998373&usg=AOvVaw2an6U6EVgKgxnAAXCddG7A) |
| 8 | [▪ Cyber-Zero: Training Cybersecurity Agents without Runtime (Zhuo et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00910&sa=D&source=editors&ust=1779048535998582&usg=AOvVaw0tXzjYQwHIqCehvBjCfyCu) |
| 9 | [▪ Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions (Qu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22617&sa=D&source=editors&ust=1779048535998763&usg=AOvVaw2Dypv_QHjhptLv79kV4Orw) |
| 10 | [▪ Verification Cost Asymmetry in Cognitive Warfare: A Complexity-Theoretic Framework (Luberisse, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21258&sa=D&source=editors&ust=1779048535998909&usg=AOvVaw0mTvFP3oKzc6d9s3fa42qK) |
| 11 | [▪ Decentralized AI-driven IoT Architecture for Privacy-Preserving and Latency-Optimized Healthcare in Pandemic and Critical Care Scenarios (Sammangi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15859&sa=D&source=editors&ust=1779048535999053&usg=AOvVaw1h-pQQMLt1sb7aSMwTOwtG) |
| 12 | [▪ Differential Privacy in Kernelized Contextual Bandits via Random Projections (Pavlovic, Salgia, and Zhao, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13639&sa=D&source=editors&ust=1779048535999219&usg=AOvVaw073EkbV7ukFI5W6dbBvs68) |
| 13 | [▪ Large Language Models are Unreliable for Cyber Threat Intelligence (Mezzi, Massacci, and Tuma, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.23175&sa=D&source=editors&ust=1779048535999381&usg=AOvVaw3AQnjPNhqmBSZ_IhZR_Ab9) |
| 14 | [▪ X Hacking: The Threat of Misguided AutoML (Sharma et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2401.08513&sa=D&source=editors&ust=1779048535999503&usg=AOvVaw1n0qkNPnno307Dk0TVbOI4) |
| 15 | [▪ Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 (Ying et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12430&sa=D&source=editors&ust=1779048535999635&usg=AOvVaw2kODfutViAR_LRXulsbjqb) |
| 16 | [▪ Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (Ren et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.21792&sa=D&source=editors&ust=1779048535999769&usg=AOvVaw0U2Q3WBN2ThdDvK0TtAL-E) |

|     |
| --- |
| Over Reliance on LLM Outputs and External (Social) Harms |

**>**

**<**

‍

#### Fake Resources and Phishing

Covers:

- MITRE ATLAS Initial Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Fake Resources and Phishing

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Phishing the Phishers with SpecularNet: Hierarchical Graph Autoencoding for Reference-Free Web Phishing Detection (Song, Casas, and Meo, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01874&sa=D&source=editors&ust=1779048536613672&usg=AOvVaw2W7RQx9I2FGfR8pRBgjsRJ) |
| 3 | [▪ Assessing Spear-Phishing Website Generation in Large Language Model Coding Agents (Malloy and Bissyande, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13363&sa=D&source=editors&ust=1779048536613838&usg=AOvVaw1U0CuGT0NlqYTEa6GANL3w) |
| 4 | [▪ CIC-Trap4Phish: A Unified Multi-Format Dataset for Phishing and Quishing Attachment Detection (Nejati et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09015&sa=D&source=editors&ust=1779048536613920&usg=AOvVaw1xkSmzuDfaKBsptftdYiOX) |
| 5 | [▪ User-Centric Phishing Detection: A RAG and LLM-Based Approach (Barwani, Korba, and Anwar, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.21261&sa=D&source=editors&ust=1779048536614027&usg=AOvVaw0cwuiVyl2rItjaFjrpaklS) |
| 6 | [▪ PhishLumos: An Adaptive Multi-Agent System for Proactive Phishing Campaign Mitigation (Chiba, Nakano, and Koide, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.21772&sa=D&source=editors&ust=1779048536614101&usg=AOvVaw2wL9OlJ080_7wjPwcVpKIR) |
| 7 | [▪ Phishing Email Detection Using Large Language Models (Hasan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10104&sa=D&source=editors&ust=1779048536614166&usg=AOvVaw1vo6Unx1BVd8dFkNx8MVx6) |
| 8 | [▪ LLM-PEA: Leveraging Large Language Models Against Phishing Email Attacks (Hassan et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10104&sa=D&source=editors&ust=1779048536614233&usg=AOvVaw3cqqZVAKw6qPJghCKkwQ8G) |
| 9 | [▪ Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features (Faisal, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.06925&sa=D&source=editors&ust=1779048536614300&usg=AOvVaw2UPqGilofNHy20aZYHc1G2) |
| 10 | [▪ Constructing and Benchmarking: a Labeled Email Dataset for Text-Based Phishing and Spam Detection Framework (Toth, Bisztray, and Dubniczky, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21448&sa=D&source=editors&ust=1779048536614370&usg=AOvVaw0WDVYg11tYIm4vPpmE3hul) |
| 11 | [▪ Small Language Models for Phishing Website Detection: Cost, Performance, and Privacy Trade-Offs (Goldenits et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15434&sa=D&source=editors&ust=1779048536614439&usg=AOvVaw0SEi88ILTPizlJWU3pdFlZ) |
| 12 | [▪ Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments (Zhou, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.15165&sa=D&source=editors&ust=1779048536614515&usg=AOvVaw0OOfbOvSV4h2aONmVVmCu_) |
| 13 | [▪ How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models (Ji and Kim, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.09606&sa=D&source=editors&ust=1779048536614591&usg=AOvVaw2pgRlywSEPAa-rSEOa3wAE) |
| 14 | [▪ MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection ((GECAD et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17978&sa=D&source=editors&ust=1779048536614663&usg=AOvVaw0o4Fg7MhlTbLN3Hj3Mq2jV) |
| 15 | [▪ A Login Page Transparency and Visual Similarity Based Zero Day Phishing Defense Protocol (Varshney et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09564&sa=D&source=editors&ust=1779048536614743&usg=AOvVaw2RBClbpvlotX_PrWXvOVCt) |
| 16 | [▪ From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks (Kulkarni et al, Jul 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2407.20361&sa=D&source=editors&ust=1779048536614816&usg=AOvVaw1nyI5OyQpSGSL5Vt1n5EvC) |

|     |
| --- |
| Fake Resources and Phishing |

**>**

**<**

‍

#### Social Manipulation

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Social Manipulation

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09497&sa=D&source=editors&ust=1779048536809569&usg=AOvVaw30S8yZiQQLFUbD6HpU1zAx) |
| 3 | [▪ A Synthetic Conversational Smishing Dataset for Social Engineering Detection (Lochstampfor and Roy, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11752&sa=D&source=editors&ust=1779048536809860&usg=AOvVaw18VrPAHJWisH_SyODl_rj7) |
| 4 | [▪ Synthetic Trust Attacks: Modeling How Generative AI Manipulates Human Decisions in Social Engineering Fraud (Ashraf, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.04951&sa=D&source=editors&ust=1779048536810028&usg=AOvVaw0MvUr_Et9scOI5u_To9BLM) |
| 5 | [▪ Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems (Huang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07473&sa=D&source=editors&ust=1779048536810170&usg=AOvVaw2zKB_mHVIkUUNC7U6jP65l) |
| 6 | [▪ When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent (Wu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.07263&sa=D&source=editors&ust=1779048536810281&usg=AOvVaw2xGdo44TI5qK7F69N-fuCt) |
| 7 | [▪ Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams (Gressel et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16280&sa=D&source=editors&ust=1779048536810383&usg=AOvVaw0_SPZsVJGgGcWPuJuu0ZoJ) |
| 8 | [▪ AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning (Hossain et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05362&sa=D&source=editors&ust=1779048536810535&usg=AOvVaw32sdUmLcf6YGNJo-Ere4Jk) |
| 9 | [▪ MURMUR: Using cross-user chatter to break collaborative language agents in groups (Patlan et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.17671&sa=D&source=editors&ust=1779048536810655&usg=AOvVaw3GrHi-97_NOVa4_NR-YIxc) |
| 10 | [▪ Investigating the Impact of Dark Patterns on LLM-Based Web Agents (Ersoy et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18113&sa=D&source=editors&ust=1779048536810754&usg=AOvVaw2uJ0Nw_DeLg5Cs2qR6EhFR) |
| 11 | [▪ Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental Health (Kwesi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10695&sa=D&source=editors&ust=1779048536810871&usg=AOvVaw0F0gQpsy-4QTy11c-iwior) |
| 12 | [▪ Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents (Patlan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.16248&sa=D&source=editors&ust=1779048536810990&usg=AOvVaw320RwKG_YQZgTu5cQoqnUl) |
| 13 | [▪ PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety (Zhang et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2401.11880&sa=D&source=editors&ust=1779048536811106&usg=AOvVaw1yVQsKBgOPifh55dr4zK7Y) |

|     |
| --- |
| Social Manipulation |

**>**

**<**

‍

#### Deep Fakes, Content Provenance, and Watermarking

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Deep Fakes, Content Provenance, and Watermarking

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks (Huang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.20924&sa=D&source=editors&ust=1779048535941700&usg=AOvVaw15jo4pSAzYyUpdCgNjlGJ6) |
| 3 | [▪ Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games (Kim, Fang, and Sandholm, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.14283&sa=D&source=editors&ust=1779048535941882&usg=AOvVaw2pJ9C09ZG724H2L0ec2rla) |
| 4 | [▪ DeePen: Penetration Testing for Audio Deepfake Detection (M\\"uller et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.20427&sa=D&source=editors&ust=1779048535941984&usg=AOvVaw2davTfWLDbE3LquqZ6xcDv) |
| 5 | [▪ Watermarking Should Be Treated as a Monitoring Primitive (Aremu, Lukas, and Zhang, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.13095&sa=D&source=editors&ust=1779048535942073&usg=AOvVaw2SHiMiO2YThYjlllbchQKD) |
| 6 | [▪ TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection (Sander et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12456&sa=D&source=editors&ust=1779048535942184&usg=AOvVaw3Y6JGNAOz1QXRVmpmvX0Y7) |
| 7 | [▪ The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive (Raza, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.12075&sa=D&source=editors&ust=1779048535942282&usg=AOvVaw1T501tIGc8iD9_m5wv4eJb) |
| 8 | [▪ Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark (Gloaguen et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11653&sa=D&source=editors&ust=1779048535942375&usg=AOvVaw0iP1UCsPjXWb2Hrz8sMas2) |
| 9 | [▪ Sequential Behavioral Watermarking for LLM Agents (An et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.11036&sa=D&source=editors&ust=1779048535942465&usg=AOvVaw1C5Bwu7fte_qTWmRUKBBK-) |
| 10 | [▪ PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks (Ai and He, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.10977&sa=D&source=editors&ust=1779048535942564&usg=AOvVaw0ixpy8x6b8iQbiAtku3IrH) |
| 11 | [▪ Majority Bit-Aware Watermarking For Large Language Models (Xu et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.03829&sa=D&source=editors&ust=1779048535942663&usg=AOvVaw3PEs-ujkH6whidOc3B-pUj) |
| 12 | [▪ Robust Spectral Watermark for Synthetic Tabular Data (Zhao et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21600&sa=D&source=editors&ust=1779048535942765&usg=AOvVaw0JvEPLgzDzf9U1eaW3BDvQ) |
| 13 | [▪ Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning (Jiang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09935&sa=D&source=editors&ust=1779048535942849&usg=AOvVaw0NrRM8d-fqThLe76UNLYyZ) |
| 14 | [▪ "Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking (Zhang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09646&sa=D&source=editors&ust=1779048535942939&usg=AOvVaw2dGR31sCoUgx1cwq4kuxq8) |
| 15 | [▪ Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal (Goonatilake and Ateniese, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.09203&sa=D&source=editors&ust=1779048535943024&usg=AOvVaw35GK-nWh0e5oz8ZEXhBBup) |
| 16 | [▪ MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking (Jiang et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22246&sa=D&source=editors&ust=1779048535943106&usg=AOvVaw0sFhDOgl6t3-kiSH5OYMCO) |
| 17 | [▪ Vaporizer: Breaking Watermarking Schemes for Large Language Model Outputs (Ng, Ngo, and Chattopadhyay, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.07481&sa=D&source=editors&ust=1779048535943196&usg=AOvVaw19oudqC9CJ-DyJTO5nIWvV) |
| 18 | [▪ Guidance Watermarking for Diffusion Models (Gesny et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.22126&sa=D&source=editors&ust=1779048535943279&usg=AOvVaw0xHmVoyQrn6ZTQ5SDLt_1o) |
| 19 | [▪ Secure Seed-Based Multi-bit Watermarking for Diffusion Models from First Principles (Gesny and Giboulot, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.06153&sa=D&source=editors&ust=1779048535943367&usg=AOvVaw3iFR4rS4uZgBgnZ_DryDcs) |
| 20 | [▪ SWAN: Semantic Watermarking with Abstract Meaning Representation (Ye et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.04305&sa=D&source=editors&ust=1779048535943448&usg=AOvVaw3tzjiuFgapbwl4azHBjqqV) |
| 21 | [▪ SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction (Yao and Juarez, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2508.05691&sa=D&source=editors&ust=1779048535943531&usg=AOvVaw0anS-wm6yNyh1JhvcB1pPn) |
| 22 | [▪ MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech (Jin et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.01515&sa=D&source=editors&ust=1779048535943618&usg=AOvVaw1usAUY6G51H_i9fXhO3Qxs) |
| 23 | [▪ VertMark: A Unified Training-Free Robust Watermarking Framework for Vertical Domain Pre-trained Language Models (Kong et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.02557&sa=D&source=editors&ust=1779048535943736&usg=AOvVaw2Ee-gYD6c2DqWcxOxZipu8) |
| 24 | [▪ Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification (Rantahalvari et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.00218&sa=D&source=editors&ust=1779048535943827&usg=AOvVaw2f9MU3hi0WU1cIfw42qrFi) |
| 25 | [▪ Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics (Chang and Echizen, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05753&sa=D&source=editors&ust=1779048535943915&usg=AOvVaw06VxM0YuAPJqftr4zQzKnD) |
| 26 | [▪ VOW: Verifiable and Oblivious Watermark Detection for Large Language Models (Luan et al., May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.27666&sa=D&source=editors&ust=1779048535944005&usg=AOvVaw3UFzSjDyEEFP2nQdh1Cdfh) |
| 27 | [▪ R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.25247&sa=D&source=editors&ust=1779048535944089&usg=AOvVaw1SiE5TB4VqVjc1RfcLUB0n) |
| 28 | [▪ DeepSignature: Digitally Signed, Content-Encoding Watermarks for Robust and Transparent Image Authentication (Graf et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.23016&sa=D&source=editors&ust=1779048535944178&usg=AOvVaw1gw2PbxonprCoOYCvbFUbY) |
| 29 | [▪ PoLO: Proof-of-Learning and Proof-of-Ownership at Once with Chained Watermarking (Deng et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.12296&sa=D&source=editors&ust=1779048535944263&usg=AOvVaw14EYo40V39M3hEtc0nFN6G) |
| 30 | [▪ ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders (Jiang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22550&sa=D&source=editors&ust=1779048535944348&usg=AOvVaw3fp-gFxZbLC3xcPk0bdq9r) |
| 31 | [▪ SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking (Gu, Du, and Grundy, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.22438&sa=D&source=editors&ust=1779048535944431&usg=AOvVaw1M7O90zJR5Q5UnBBOq3Qbr) |
| 32 | [▪ Dual-Guard: Dual-Channel Latent Watermarking for Provenance and Tamper Localization in Diffusion Images (Xie et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19090&sa=D&source=editors&ust=1779048535944517&usg=AOvVaw3s1UYQAjaEbArE2IS0-gz6) |
| 33 | [▪ CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2510.11251&sa=D&source=editors&ust=1779048535944601&usg=AOvVaw2SXUCsP6bThNh1hPmnRM8O) |
| 34 | [▪ Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking (Nemecek et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.13776&sa=D&source=editors&ust=1779048535944702&usg=AOvVaw0PFBghOLTzfMIdmtElgnoy) |
| 35 | [▪ TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC (Che, Du, and Gao, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12216&sa=D&source=editors&ust=1779048535944793&usg=AOvVaw3gVfYLJFQ3LZuaapy_5c-q) |
| 36 | [▪ Can we Watermark Low-Entropy LLM Outputs? (Mazor, Morgan, and Pass, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.12051&sa=D&source=editors&ust=1779048535944876&usg=AOvVaw2-9j0bZaMs6UC9srM0yBm-) |
| 37 | [▪ On the Robustness of Watermarking for Autoregressive Image Generation (M\\"uller et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11720&sa=D&source=editors&ust=1779048535944961&usg=AOvVaw22nx0klU5vpVCMIJ4zEsLZ) |
| 38 | [▪ RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience (Huang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11546&sa=D&source=editors&ust=1779048535945043&usg=AOvVaw2F25mIo8md5vfECqkhVXV4) |
| 39 | [▪ Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models (Zhang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10893&sa=D&source=editors&ust=1779048535945125&usg=AOvVaw0XNbPJr9aBFyEhYptQeqRW) |
| 40 | [▪ SEED: A Large-Scale Benchmark for Provenance Tracing in Sequential Deepfake Facial Edits (Hoi et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.10522&sa=D&source=editors&ust=1779048535945217&usg=AOvVaw1iZCtgTNQOtNp4cfAg6VUC) |
| 41 | [▪ Towards Better Statistical Understanding of Watermarking LLMs (Cai et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2403.13027&sa=D&source=editors&ust=1779048535945304&usg=AOvVaw2C6f79ThV7XZY8fpsQQMso) |
| 42 | [▪ XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts (Xu et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.05242&sa=D&source=editors&ust=1779048535945385&usg=AOvVaw0UYb2ZhTxzqGM7-8yvGi0y) |
| 43 | [▪ An End-to-End Model for Logits-Based Large Language Models Watermarking (Wong et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.02344&sa=D&source=editors&ust=1779048535945465&usg=AOvVaw0ResyPI3vgkeuWQ-xGKVVv) |
| 44 | [▪ Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors (Stan\\v{e}k et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01330&sa=D&source=editors&ust=1779048535945557&usg=AOvVaw3vQ6kpxrx_EF_-N8k97nlp) |
| 45 | [▪ Refined Detection for Gumbel Watermarking (Lattimore, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.30017&sa=D&source=editors&ust=1779048535945636&usg=AOvVaw20PJd385NyCq7xg1LtKjoc) |
| 46 | [▪ SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermark (Bao et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.29742&sa=D&source=editors&ust=1779048535945727&usg=AOvVaw0b11P25jrWwFYNSTEOq0K-) |
| 47 | [▪ Gaussian Shannon: High-Precision Diffusion Model Watermarking Based on Communication (Zhang, Huang, and Zhang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.26167&sa=D&source=editors&ust=1779048535945812&usg=AOvVaw2MHc5_jrDybQMjp5oJxwhP) |
| 48 | [▪ NOWA: Null-space Optical Watermark for Invisible Capture Fingerprinting and Tamper Localization (Vargas et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2512.22501&sa=D&source=editors&ust=1779048535945897&usg=AOvVaw1sK30TNT1eLujEjF6Njiw9) |
| 49 | [▪ Robust Safety Monitoring of Language Models via Activation Watermarking (Aremu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.23171&sa=D&source=editors&ust=1779048535945981&usg=AOvVaw2y8FFMUEPZQaijs1mtMXan) |
| 50 | [▪ Functional Subspace Watermarking for Large Language Models (Ding et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.18793&sa=D&source=editors&ust=1779048535946063&usg=AOvVaw17c64qqH8gVEzVZP7I6-Fp) |
| 51 | [▪ Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.17531&sa=D&source=editors&ust=1779048535946151&usg=AOvVaw2gBCadNpwM27YXyjA7HUcn) |
| 52 | [▪ Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework (Wang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.14968&sa=D&source=editors&ust=1779048535946239&usg=AOvVaw2w_OT4OiAS4kRFazXjKo5h) |
| 53 | [▪ TableMark: A Multi-bit Watermark for Synthetic Tabular Data (Xia et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.13722&sa=D&source=editors&ust=1779048535946354&usg=AOvVaw1u6s1TRcPaCBq9oPxUj3uh) |
| 54 | [▪ Editing Away the Evidence: Diffusion-Based Image Manipulation and the Failure Modes of Robust Watermarking (Qi et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12949&sa=D&source=editors&ust=1779048535946451&usg=AOvVaw2M2TOe6Eyc2m_uG5QMjReu) |
| 55 | [▪ SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking (Gao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12749&sa=D&source=editors&ust=1779048535946544&usg=AOvVaw0ZdJbIYI3102_wN-Wjc8Uf) |
| 56 | [▪ EmbTracker: Traceable Black-box Watermarking for Federated Language Models (Zhao et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12089&sa=D&source=editors&ust=1779048535946636&usg=AOvVaw0nlGNB_CiiVeAjcfJ5zb2C) |
| 57 | [▪ Cluster-Aware Attacks on Graph Watermarks (Nemecek, Yilmaz, and Ayday, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.17971&sa=D&source=editors&ust=1779048535946754&usg=AOvVaw23Od8Osso5Gd7ZKM03I16b) |
| 58 | [▪ Na\\"ive Exposure of Generative AI Capabilities Undermines Deepfake Detection (Kim et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10504&sa=D&source=editors&ust=1779048535946854&usg=AOvVaw0jQUGwRQoz4zJlhcb-69DV) |
| 59 | [▪ The Orthogonal Vulnerabilities of Generative AI Watermarks: A Comparative Empirical Benchmark of Spatial and Latent Provenance (Yu and Wei, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.10323&sa=D&source=editors&ust=1779048535946959&usg=AOvVaw3lymfe9mPIAFllyYjkVnDu) |
| 60 | [▪ ShapeMark: Robust and Diversity-Preserving Watermarking for Diffusion Models (Qian et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09454&sa=D&source=editors&ust=1779048535947044&usg=AOvVaw3H1CESouyaP2ZZkjZPAMMH) |
| 61 | [▪ mAVE: A Watermark for Joint Audio-Visual Generation Models (Si, Pan, and Wen, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.07090&sa=D&source=editors&ust=1779048535947156&usg=AOvVaw2kVI4Qj1Ay8jUfCnozs3_O) |
| 62 | [▪ When Denoising Becomes Unsigning: Theoretical and Empirical Analysis of Watermark Fragility Under Diffusion-Based Image Editing (Gu et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04696&sa=D&source=editors&ust=1779048535947257&usg=AOvVaw0OFAOg3Xt1vZWhF2pu2iA9) |
| 63 | [▪ How Effective Are Publicly Accessible Deepfake Detection Tools? A Comparative Evaluation of Open-Source and Free-to-Use Platforms (Rettinger et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.04456&sa=D&source=editors&ust=1779048535947347&usg=AOvVaw1i5ayfzrWoy7010OaUUDdA) |
| 64 | [▪ On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation (Omidi, Dong, and Wang, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.03410&sa=D&source=editors&ust=1779048535947433&usg=AOvVaw1ymf8itQyDlRH491EApX6F) |
| 65 | [▪ Watermarking Without Standards Is Not AI Governance (Nemecek, Jiang, and Ayday, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.23814&sa=D&source=editors&ust=1779048535947534&usg=AOvVaw0sad0CHO8031UFem1UIR0H) |
| 66 | [▪ Topic-Based Watermarks for Large Language Models (Nemecek, Jiang, and Ayday, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2404.02138&sa=D&source=editors&ust=1779048535947626&usg=AOvVaw1LxBYFNPtKZYCmuK-AJbUf) |
| 67 | [▪ Scores Know Bobs Voice: Speaker Impersonation Attack (Hwang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02781&sa=D&source=editors&ust=1779048535947741&usg=AOvVaw2tMcrl97e3ApeGhZ2s1Fjv) |
| 68 | [▪ Authenticated Contradictions from Desynchronized Provenance and Watermarking (Nemecek et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.02378&sa=D&source=editors&ust=1779048535947865&usg=AOvVaw1cK5nTRI8WJXdeEnk6UYuI) |
| 69 | [▪ PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints (Huo et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.21057&sa=D&source=editors&ust=1779048535948013&usg=AOvVaw1G07FTwAzSNrWbNK-DVoWL) |
| 70 | [▪ SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models (Yang et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00194&sa=D&source=editors&ust=1779048535948143&usg=AOvVaw3D7y-QxREx-rRNUr_rWOAX) |
| 71 | [▪ Hide&Seek: Remove Image Watermarks with Negligible Cost via Pixel-wise Reconstruction (Chen et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.01067&sa=D&source=editors&ust=1779048535948243&usg=AOvVaw1aofrPeNI6NzJVCzUm92bF) |
| 72 | [▪ LLM-Text Watermarking based on Lagrange Interpolation (Janas, Morawiecki, and Pieprzyk, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2505.05712&sa=D&source=editors&ust=1779048535948332&usg=AOvVaw0oa4rmToZrdQ2ZWkqXGQ5w) |
| 73 | [▪ WaterVIB: Learning Minimal Sufficient Watermark Representations via Variational Information Bottleneck (He et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.21508&sa=D&source=editors&ust=1779048535948417&usg=AOvVaw35U50RI-iWBeAP7H9Y3Nnu) |
| 74 | [▪ Vanishing Watermarks: Diffusion-Based Image Editing Undermines Robust Invisible Watermarking (Guo et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20680&sa=D&source=editors&ust=1779048535948507&usg=AOvVaw1KJNA3EOv7YYggXKA4lJux) |
| 75 | [▪ Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models (He et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01428&sa=D&source=editors&ust=1779048535948591&usg=AOvVaw020hMAwvrT1rUOUKBbExYX) |
| 76 | [▪ Can You Tell It's AI? Human Perception of Synthetic Voices in Vishing Scenarios (Bhatti et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.20061&sa=D&source=editors&ust=1779048535948681&usg=AOvVaw21xWePHW4U4hQ5cQwVLjxa) |
| 77 | [▪ Watermarking LLM Agent Trajectories (Meng et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.18700&sa=D&source=editors&ust=1779048535948764&usg=AOvVaw1xJcw_BSUiP0ZwbARaG8Jb) |
| 78 | [▪ MarkSweep: A No-box Removal Attack on AI-Generated Image Watermarking via Noise Intensification and Frequency-aware Denoising (Cao et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15364&sa=D&source=editors&ust=1779048535948850&usg=AOvVaw1jeP_HdNnl7j2MhTW97QVs) |
| 79 | [▪ Unforgeable Watermarks for Language Models via Robust Signatures (Lin, Shahabi, and Song, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.15323&sa=D&source=editors&ust=1779048535948932&usg=AOvVaw27Xo5zBKLd6p0qzgzq7LBz) |
| 80 | [▪ TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.20532&sa=D&source=editors&ust=1779048535949013&usg=AOvVaw27vrGnf6TsTQ7pOjuPY7kc) |
| 81 | [▪ MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages (Cui et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.14030&sa=D&source=editors&ust=1779048535949093&usg=AOvVaw2l8ThPtKNSljXJplPA1a8a) |
| 82 | [▪ MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks (Zhou et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2509.10766&sa=D&source=editors&ust=1779048535949183&usg=AOvVaw2JUPW-CKSt1PPonUjijwVP) |
| 83 | [▪ More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles (Chen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.11793&sa=D&source=editors&ust=1779048535949270&usg=AOvVaw0HfczQZ0XYjvTyV0cIMAcd) |
| 84 | [▪ MerkleSpeech: Public-Key Verifiable, Chunk-Localised Speech Provenance via Perceptual Fingerprints and Merkle Commitments (Ono, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.10166&sa=D&source=editors&ust=1779048535949386&usg=AOvVaw1YhtBTXWEp6BUMtfxaABPl) |
| 85 | [▪ AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.09611&sa=D&source=editors&ust=1779048535949492&usg=AOvVaw0UU0rYiraM_soq6JS-8wbT) |
| 86 | [▪ On Protecting Agentic Systems' Intellectual Property via Watermarking (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.08401&sa=D&source=editors&ust=1779048535949580&usg=AOvVaw1hJmSYmJKbHEV9cztvaRc5) |
| 87 | [▪ A Unified Framework for LLM Watermarks (Gloaguen et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.06754&sa=D&source=editors&ust=1779048535949679&usg=AOvVaw1FkOyo88SxSNQLd073T7EY) |
| 88 | [▪ SynthForensics: A Multi-Generator Benchmark for Detecting Synthetic Video Deepfakes (Leotta et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.04939&sa=D&source=editors&ust=1779048535949773&usg=AOvVaw0ty6LL7ys0F0WoLFx_ipru) |
| 89 | [▪ Origin Lens: A Privacy-First Mobile Framework for Cryptographic Image Provenance and AI Detection (Loth et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.03423&sa=D&source=editors&ust=1779048535949858&usg=AOvVaw2QcHFEFZ-N0dVHRsSvsKbR) |
| 90 | [▪ Position: 3D Gaussian Splatting Watermarking Should Be Scenario-Driven and Threat-Model Explicit (Deng, Nakra, and Wu, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02602&sa=D&source=editors&ust=1779048535949944&usg=AOvVaw28qne0p5-wqIk_-zeDZYo3) |
| 91 | [▪ MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection (Liu et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02222&sa=D&source=editors&ust=1779048535950062&usg=AOvVaw17f-Jsz8oYB5BRwjMfu9Tn) |
| 92 | [▪ WorldCup Sampling for Multi-bit LLM Watermarking (Wang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01752&sa=D&source=editors&ust=1779048535950152&usg=AOvVaw1mPW7eK0UyIc9sjURwhDBd) |
| 93 | [▪ MarkCleaner: High-Fidelity Watermark Removal via Imperceptible Micro-Geometric Perturbation (Kong et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.01513&sa=D&source=editors&ust=1779048535950261&usg=AOvVaw07883s0Vbjb49ihLG1RXpi) |
| 94 | [▪ Provenance Verification of AI-Generated Images via a Perceptual Hash Registry Anchored on Blockchain (Mohit, Aggarwal, and Gondhalekar, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.02412&sa=D&source=editors&ust=1779048535950350&usg=AOvVaw3beJaK6AGjv2jdfpS6GEuf) |
| 95 | [▪ Color Matters: Demosaicing-Guided Color Correlation Training for Generalizable AI-Generated Image Detection (Zhong, Xu, and Zou, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22778&sa=D&source=editors&ust=1779048535950435&usg=AOvVaw1v_4mVkcz5_pocgqI-P-lA) |
| 96 | [▪ VocBulwark: Towards Practical Generative Speech Watermarking via Additional-Parameter Injection (Liu, Li, and Yin, February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22556&sa=D&source=editors&ust=1779048535950552&usg=AOvVaw0VlNDUQpYFZUV5muWjYgo_) |
| 97 | [▪ MirrorMark: A Distortion-Free Multi-Bit Watermark for Large Language Models (Jiang et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.22246&sa=D&source=editors&ust=1779048535950675&usg=AOvVaw0v9BqTKgK_3sKdDy8fjwze) |
| 98 | [▪ SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.20310&sa=D&source=editors&ust=1779048535950778&usg=AOvVaw00AUkR0xKsbmtBHLBC2Kjm) |
| 99 | [▪ SWA-LDM: Toward Stealthy Watermarks for Latent Diffusion Models (Yang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.10495&sa=D&source=editors&ust=1779048535950864&usg=AOvVaw03KqPgVUNUe4AYw9M-aTON) |
| 100 | [▪ Watermark-based Attribution of AI-Generated Content (Jiang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2404.04254&sa=D&source=editors&ust=1779048535950949&usg=AOvVaw2O2GXA3VMHEO3DWwBe6Xro) |
| 101 | [▪ DeMark: A Query-Free Black-Box Attack on Deepfake Watermarking Defenses (Song et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16473&sa=D&source=editors&ust=1779048535951031&usg=AOvVaw20GnxQLB2H5BuIlSXMk6dv) |
| 102 | [▪ Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing (Zhang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.00035&sa=D&source=editors&ust=1779048535951114&usg=AOvVaw2UxLjXQqAqrZhOnYAE7wX9) |
| 103 | [▪ Learning to Watermark in the Latent Space of Generative Models (Rebuffi et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.16140&sa=D&source=editors&ust=1779048535951214&usg=AOvVaw1uGGOpxA_CBb8Enh5fX93k) |
| 104 | [▪ DRGW: Learning Disentangled Representations for Robust Graph Watermarking (Li et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.13569&sa=D&source=editors&ust=1779048535951300&usg=AOvVaw1IOxobmEMb7iRLAxwaHj6_) |
| 105 | [▪ GenPTW: Latent Image Watermarking for Provenance Tracing and Tamper Localization (Gan et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2504.19567&sa=D&source=editors&ust=1779048535951402&usg=AOvVaw0-4SP3UcczY_mhivaX9srO) |
| 106 | [▪ Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks (Xu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2501.09328&sa=D&source=editors&ust=1779048535951532&usg=AOvVaw00irXju7nHmsaW0Izs2jS3) |
| 107 | [▪ Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs (Le, Ritter, and Goyal, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.11629&sa=D&source=editors&ust=1779048535951637&usg=AOvVaw1b2Ww1tBgiyymoQqFTDxJc) |
| 108 | [▪ Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment (Wang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2502.10803&sa=D&source=editors&ust=1779048535951732&usg=AOvVaw2A0WATGkF5oLqp8GUy2CBl) |
| 109 | [▪ Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.06639&sa=D&source=editors&ust=1779048535951832&usg=AOvVaw3k1xsdmHufzeDGMMOOMLR7) |
| 110 | [▪ Deepfake detectors are DUMB: A benchmark to assess adversarial training robustness under transferability constraints (Serrano, Umlil, and Thomas, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05986&sa=D&source=editors&ust=1779048535951935&usg=AOvVaw2Wcd2f6CnuhrYvEPYXtzNZ) |
| 111 | [▪ AgentMark: Utility-Preserving Behavioral Watermarking for Agents (Huang et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.03294&sa=D&source=editors&ust=1779048535952027&usg=AOvVaw3oC1BCQvGeKwYELWg-cmwC) |
| 112 | [▪ Vulnerabilities of Audio-Based Biometric Authentication Systems Against Deepfake Speech Synthesis (Hong et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.02914&sa=D&source=editors&ust=1779048535952130&usg=AOvVaw263yjDxZqTxOxwV5sUzovX) |
| 113 | [▪ SoK: Are Watermarks in LLMs Ready for Deployment? (Dang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.05594&sa=D&source=editors&ust=1779048535952245&usg=AOvVaw0m1YJP8Rk8LVr4kV6tUyI0) |
| 114 | [▪ Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform (Zhang, Li, and Gu, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18791&sa=D&source=editors&ust=1779048535952365&usg=AOvVaw1eluUDJT3V378CSQIPvvMW) |
| 115 | [▪ Pixel Seal: Adversarial-only training for invisible image and video watermarking (Sou\\v{c}ek et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16874&sa=D&source=editors&ust=1779048535952467&usg=AOvVaw2GOO4mFQn3hl80hWm2flk5) |
| 116 | [▪ How Good is Post-Hoc Watermarking With Language Model Rephrasing? (Fernandez et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16904&sa=D&source=editors&ust=1779048535952608&usg=AOvVaw3-F7tLCkn_awRMiUz2NQRc) |
| 117 | [▪ Protecting Deep Neural Network Intellectual Property with Chaos-Based White-Box Watermarking (B et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16658&sa=D&source=editors&ust=1779048535952772&usg=AOvVaw0kf3tOsaCF6-J9xjZ6fi4q) |
| 118 | [▪ DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.16182&sa=D&source=editors&ust=1779048535952917&usg=AOvVaw3K9HxmkV2AUQR8nF8fhlqt) |
| 119 | [▪ Remotely Detectable Robot Policy Watermarking (Amir, Flageat, and Prorok, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15379&sa=D&source=editors&ust=1779048535953081&usg=AOvVaw2es_KUYYYo8U_UFCeVDtNP) |
| 120 | [▪ ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples (Yang et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.15641&sa=D&source=editors&ust=1779048535953225&usg=AOvVaw0-JoZtCEhi4ZAYtRtzWXF_) |
| 121 | [▪ CODE ACROSTIC: Robust Watermarking for Code Generation (Lin et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.14753&sa=D&source=editors&ust=1779048535953361&usg=AOvVaw0nUD4b8S2T5OLReLD6i2_N) |
| 122 | [▪ SPDMark: Selective Parameter Displacement for Robust Video Watermarking (Fares, Tastan, and Nandakumar, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12090&sa=D&source=editors&ust=1779048535953486&usg=AOvVaw1-1yKa4RAMRxI3lE_1G5MY) |
| 123 | [▪ Security and Detectability Analysis of Unicode Text Watermarking Methods Against Large Language Models (Hellmeier, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.13325&sa=D&source=editors&ust=1779048535953595&usg=AOvVaw3tyJs8Wc7sGh2E3EXainIW) |
| 124 | [▪ UniMark: Artificial Intelligence Generated Content Identification Toolkit (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.12324&sa=D&source=editors&ust=1779048535953722&usg=AOvVaw2c5I-xELqqnbWr0Lk0GA0Z) |
| 125 | [▪ Lightweight Model Attribution and Detection of Synthetic Speech via Audio Residual Fingerprints (Pizarro et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.14013&sa=D&source=editors&ust=1779048535953875&usg=AOvVaw1q6PM85UL-SiYdC9W4_sb8) |
| 126 | [▪ TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection (Jiang-Lin et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10652&sa=D&source=editors&ust=1779048535954011&usg=AOvVaw3rBVh0k4csrT7Uyt_AgLMK) |
| 127 | [▪ Watermarks for Language Models via Probabilistic Automata (Wang and Shang, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.10185&sa=D&source=editors&ust=1779048535954122&usg=AOvVaw1-bYaRTYLMOgm4fhDAltD-) |
| 128 | [▪ Towards Robust Protective Perturbation against DeepFake Face Swapping (Yao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07228&sa=D&source=editors&ust=1779048535954224&usg=AOvVaw1VcRWY4klkN4Wx1HqvxACL) |
| 129 | [▪ Ideal Attribution and Faithful Watermarks for Language Models (Song and Shahabi, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.07038&sa=D&source=editors&ust=1779048535954308&usg=AOvVaw03DR3ofOC-tAJH8qI7oT0t) |
| 130 | [▪ Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking (Yao et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.05835&sa=D&source=editors&ust=1779048535954407&usg=AOvVaw2EB1jSWa0QMhRtuIsXqO1d) |
| 131 | [▪ Detection of AI Deepfake and Fraud in Online Payments Using GAN-Based Models (Ke et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.07033&sa=D&source=editors&ust=1779048535954504&usg=AOvVaw1GmQXAcFEE3GgJHovee1Ek) |
| 132 | [▪ MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking (Zhao, Wu, and Block, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.04044&sa=D&source=editors&ust=1779048535954606&usg=AOvVaw1aStHYd5UvV79Sg2Fmr2mK) |
| 133 | [▪ Watermarks for Embeddings-as-a-Service Large Language Models (Shetty, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03079&sa=D&source=editors&ust=1779048535954741&usg=AOvVaw17MJNn_MyeK-ICBfL6SOcC) |
| 134 | [▪ HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions (Tsur et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.06409&sa=D&source=editors&ust=1779048535954878&usg=AOvVaw0ju3aKSMHubBpHK-DwQvTK) |
| 135 | [▪ JPEGs Just Got Snipped: Croppable Signatures Against Deepfake Images (Perazzo et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.01845&sa=D&source=editors&ust=1779048535955017&usg=AOvVaw08wxgqbdeueOdr9PKKcse-) |
| 136 | [▪ HMARK: Radioactive Multi-Bit Semantic-Latent Watermarking for Diffusion Models (Li et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.00094&sa=D&source=editors&ust=1779048535955149&usg=AOvVaw2pz-G9-cxWPUiuozPh76No) |
| 137 | [▪ TAB-DRW: A DFT-based Robust Watermark for Generative Tabular Data (Zhao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.21600&sa=D&source=editors&ust=1779048535955264&usg=AOvVaw1S0vTdH1RrNjmkXocOVbOH) |
| 138 | [▪ Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection (Liu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.19886&sa=D&source=editors&ust=1779048535955388&usg=AOvVaw1LacFr3cLhsPqokdGs_QJC) |
| 139 | [▪ The Coding Limits of Robust Watermarking for Generative Models (Francati et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.10577&sa=D&source=editors&ust=1779048535955511&usg=AOvVaw0_-zYIvKVqgWChHBOMEAFK) |
| 140 | [▪ VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models (Teymoorianfard et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.01406&sa=D&source=editors&ust=1779048535955624&usg=AOvVaw0eIbUbL84_fm1KAghmU6d-) |
| 141 | [▪ ForensicFlow: A Tri-Modal Adaptive Network for Robust Deepfake Detection (Romani, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14554&sa=D&source=editors&ust=1779048535955764&usg=AOvVaw2i0dJBMEA3H0lgbKuGNeBs) |
| 142 | [▪ Sigil: Server-Enforced Watermarking in U-Shaped Split Federated Learning via Gradient Injection (Dai et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.14422&sa=D&source=editors&ust=1779048535955902&usg=AOvVaw115k3ky0k-Rws7RUWRZySS) |
| 143 | [▪ Video Signature: Implicit Watermarking for Video Diffusion Models (Huang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.00652&sa=D&source=editors&ust=1779048535956028&usg=AOvVaw2gJIzWlIlp7tpzD5QKfnbK) |
| 144 | [▪ LLM-driven Provenance Forensics for Threat Investigation and Detection (Mukherjee and Kantarcioglu, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.21323&sa=D&source=editors&ust=1779048535956160&usg=AOvVaw28zOsBfyB3Jc_kd7ismfe0) |
| 145 | [▪ VideoMark: A Distortion-Free Robust Watermarking Framework for Video Diffusion Models (Hu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16359&sa=D&source=editors&ust=1779048535956299&usg=AOvVaw24SCnRHNLBqQcJAQuwb9lL) |
| 146 | [▪ FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning (Gu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12663&sa=D&source=editors&ust=1779048535956412&usg=AOvVaw1aMIwKWQkCejUM0O_ZIGS-) |
| 147 | [▪ DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training (Kumar et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.12048&sa=D&source=editors&ust=1779048535956522&usg=AOvVaw3dvW9ul1v9t1qlBqmoIe6I) |
| 148 | [▪ Robust Client-Server Watermarking for Split Federated Learning (Tang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.13598&sa=D&source=editors&ust=1779048535956633&usg=AOvVaw2KhHmU6Ny_WWGUUCzW5s0s) |
| 149 | [▪ Adaptive and Robust Watermark for Generative Tabular Data (Ngo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.14700&sa=D&source=editors&ust=1779048535956783&usg=AOvVaw3FZf0EB_iZ3VXdk_qgClDX) |
| 150 | [▪ Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio (Chen et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10913&sa=D&source=editors&ust=1779048535956926&usg=AOvVaw3jZ4RxfrWU4bdzjrkb5wxV) |
| 151 | [▪ SEAL: Subspace-Anchored Watermarks for LLM Ownership (Dai et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.11356&sa=D&source=editors&ust=1779048535957064&usg=AOvVaw1tKzVRsojGf_Pfn52RzVhU) |
| 152 | [▪ On the Information-Theoretic Fragility of Robust Watermarking under Diffusion Editing (Ni et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.10933&sa=D&source=editors&ust=1779048535957174&usg=AOvVaw31ZmuLoZD3xywqtHRxGfED) |
| 153 | [▪ Removal Attack and Defense on AI-generated Content Latent-based Watermarking (Lee et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.11745&sa=D&source=editors&ust=1779048535957264&usg=AOvVaw3W3kw9VHi2zllD741SCO_k) |
| 154 | [▪ DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks (Yang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08985&sa=D&source=editors&ust=1779048535957355&usg=AOvVaw2RFR-qUWGEVZzkzoZ1v5ZO) |
| 155 | [▪ Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks (Xiao et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07947&sa=D&source=editors&ust=1779048535957449&usg=AOvVaw1Q34IS0hVHbiQ5yMqIlU6-) |
| 156 | [▪ Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation (Zhou et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.07051&sa=D&source=editors&ust=1779048535957544&usg=AOvVaw2L2qlpHrpFRSP6WLQglsbJ) |
| 157 | [▪ Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks (Fu et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.05598&sa=D&source=editors&ust=1779048535957628&usg=AOvVaw1PrQdyx23ym4_bHotnH6_s) |
| 158 | [▪ Shallow Diffuse: Robust and Invisible Watermarking through Low-Dimensional Subspaces in Diffusion Models (Li, Zhang, and Qu, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.21088&sa=D&source=editors&ust=1779048535957721&usg=AOvVaw10LVwPc8fzopIFGMWy-jqC) |
| 159 | [▪ Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology (Souverain, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.03641&sa=D&source=editors&ust=1779048535957807&usg=AOvVaw3TIH7tCNhFsYC2nCOMoWcN) |
| 160 | [▪ Watermarking Discrete Diffusion Language Models (Bagchi et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.02083&sa=D&source=editors&ust=1779048535957888&usg=AOvVaw2Ga1tU7vi_n964bDgnNbHu) |
| 161 | [▪ Optimizing Token Choice for Code Watermarking: An RL Approach (Guo et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.11925&sa=D&source=editors&ust=1779048535957969&usg=AOvVaw0TBLqahYqdQZRirNYL31kl) |
| 162 | [▪ From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection (Liang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.00181&sa=D&source=editors&ust=1779048535958053&usg=AOvVaw05VjcahDbr_-EJNtuim6c-) |
| 163 | [▪ Robust GNN Watermarking via Implicit Perception of Topological Invariants (Li and Shen, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.25934&sa=D&source=editors&ust=1779048535958135&usg=AOvVaw0ViCU74apP1o_yz6AS0nyc) |
| 164 | [▪ PVMark: Enabling Public Verifiability for LLM Watermarking Schemes (Duan, Xiang, and Zhang, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26274&sa=D&source=editors&ust=1779048535958244&usg=AOvVaw00BCyuxCXvYhp14VZ7OCiv) |
| 165 | [▪ PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs (Xue et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.23891&sa=D&source=editors&ust=1779048535958347&usg=AOvVaw2XY_OuCH54rTv6wyU6bvi-) |
| 166 | [▪ Optimal Detection for Language Watermarks with Pseudorandom Collision (Cai et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22007&sa=D&source=editors&ust=1779048535958442&usg=AOvVaw2g-vKljK3eEaleUYFznc46) |
| 167 | [▪ DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection (Zhao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.22622&sa=D&source=editors&ust=1779048535958534&usg=AOvVaw1b4M1l6JO6QEu9jk-zn5lA) |
| 168 | [▪ WMCopier: Forging Invisible Image Watermarks on Arbitrary Images (Dong et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.22330&sa=D&source=editors&ust=1779048535958626&usg=AOvVaw0BXq0xt0CwZBdjuQWY5L_G) |
| 169 | [▪ A Reinforcement Learning Framework for Robust and Secure LLM Watermarking (An et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21053&sa=D&source=editors&ust=1779048535958730&usg=AOvVaw3CKvHKk5PPajwQGpv4EZ0C) |
| 170 | [▪ Can Current Detectors Catch Face-to-Voice Deepfake Attacks? (Nguyen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21004&sa=D&source=editors&ust=1779048535958824&usg=AOvVaw3QK4E6tiEqC1nhozdgo_Xb) |
| 171 | [▪ Watermarking Autoregressive Image Generation (Jovanovi\\'c et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.16349&sa=D&source=editors&ust=1779048535958925&usg=AOvVaw2MpNqI6ck1RP1ARSUDZjzJ) |
| 172 | [▪ Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models (Sou\\v{c}ek et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.20468&sa=D&source=editors&ust=1779048535959008&usg=AOvVaw3E9W265llqOWY4PO3InXlb) |
| 173 | [▪ Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID (Han et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.20228&sa=D&source=editors&ust=1779048535959108&usg=AOvVaw1OorEXb7u7SahAzmLrOePx) |
| 174 | [▪ Provenance of AI-Generated Images: A Vector Similarity and Blockchain-based Approach (Sharma, Carvalho, and Bhunia, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17854&sa=D&source=editors&ust=1779048535959216&usg=AOvVaw2Bpje0SJi_HwwzbKJpweSq) |
| 175 | [▪ Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption (Liu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.18333&sa=D&source=editors&ust=1779048535959312&usg=AOvVaw0Se68XSvh3BdsxSqAwePs4) |
| 176 | [▪ EditMark: Watermarking Large Language Models based on Model Editing (Li et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.16367&sa=D&source=editors&ust=1779048535959406&usg=AOvVaw0k3Eer4ApDRSoGh4KY2S40) |
| 177 | [▪ Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization (Wang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15976&sa=D&source=editors&ust=1779048535959496&usg=AOvVaw2Dzd6OQThkU5dptQqPIsg1) |
| 178 | [▪ MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models (Pan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.10569&sa=D&source=editors&ust=1779048535959581&usg=AOvVaw3ym7Ohyt09Ai0hj2XPN6OX) |
| 179 | [▪ Every Language Model Has a Forgery-Resistant Signature (Finlayson, Ren, and Swayamdipta, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.14086&sa=D&source=editors&ust=1779048535959672&usg=AOvVaw1FNZXn00jGJJGMZQ157gGO) |
| 180 | [▪ NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models (Goren et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.13793&sa=D&source=editors&ust=1779048535959757&usg=AOvVaw3jg0cceIwxrFBihvfXznkX) |
| 181 | [▪ SimKey: A Semantically Aware Key Module for Watermarking Language Models (Kodama et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12828&sa=D&source=editors&ust=1779048535959839&usg=AOvVaw2SzlOUBK_cM1elKEhP_CfI) |
| 182 | [▪ We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice (Petrov et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.12812&sa=D&source=editors&ust=1779048535959923&usg=AOvVaw04R1orelQOVPAnRiMIS6Hf) |
| 183 | [▪ SWIFT: Semantic Watermarking for Image Forgery Thwarting (Evennou et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2407.18995&sa=D&source=editors&ust=1779048535960005&usg=AOvVaw2qR13J4gl9bTKlQFYcgOCa) |
| 184 | [▪ SynthID-Image: Image watermarking at internet scale (Gowal et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.09263&sa=D&source=editors&ust=1779048535960087&usg=AOvVaw3VFjbM9X8d9PL1Yn_-aUyW) |
| 185 | [▪ STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution (Firc et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.19644&sa=D&source=editors&ust=1779048535960178&usg=AOvVaw2Nr96MfWJuILBB51YdUacs) |
| 186 | [▪ LLM Fingerprinting via Semantically Conditioned Watermarks (Gloaguen et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.16723&sa=D&source=editors&ust=1779048535960282&usg=AOvVaw36mv6XRDAH7xNyKoufowTh) |
| 187 | [▪ Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race (Mao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06544&sa=D&source=editors&ust=1779048535960371&usg=AOvVaw0N5byiqfP6MdsO2YjXfO82) |
| 188 | [▪ Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security (Naseh et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.06525&sa=D&source=editors&ust=1779048535960457&usg=AOvVaw1o0Z2iUzPB_fBH6okrUvyC) |
| 189 | [▪ Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code (Kim, Park, and Han, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.18851&sa=D&source=editors&ust=1779048535960542&usg=AOvVaw0XpAKOSpUoqnHd09zrLVrZ) |
| 190 | [▪ LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data (German et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.14474&sa=D&source=editors&ust=1779048535960629&usg=AOvVaw2h9MhFtqBM2n_DcZHHevQK) |
| 191 | [▪ DMark: Order-Agnostic Watermarking for Diffusion Large Language Models (Wu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02902&sa=D&source=editors&ust=1779048535960719&usg=AOvVaw2Y1KtnoR8v2In5py32eONp) |
| 192 | [▪ Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey (Cao et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02384&sa=D&source=editors&ust=1779048535960802&usg=AOvVaw38GFWbty31bHt6KU51tabK) |
| 193 | [▪ CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models (Zhang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.02342&sa=D&source=editors&ust=1779048535960886&usg=AOvVaw2vceBdq4bzsXIjIEri02lu) |
| 194 | [▪ ZK-WAGON: Imperceptible Watermark for Image Generation Models using ZK-SNARKs (Ramakrishnan et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01967&sa=D&source=editors&ust=1779048535960969&usg=AOvVaw1EK5kdT0Mz5Age_FR_2xKK) |
| 195 | [▪ EditTrack: Detecting and Attributing AI-assisted Image Editing (Jiang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.01173&sa=D&source=editors&ust=1779048535961065&usg=AOvVaw1cmTRjr5wWG0UilXFVD2dN) |
| 196 | [▪ Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors (Evennou, Chappelier, and Kijak, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.00799&sa=D&source=editors&ust=1779048535961152&usg=AOvVaw3Z1lTV4TMGunqt-mJzhS-j) |
| 197 | [▪ Watermark under Fire: A Robustness Evaluation of LLM Watermarking (Liang et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.13425&sa=D&source=editors&ust=1779048535961237&usg=AOvVaw2jfz3PTu-dqSfQK48q9SrX) |
| 198 | [▪ Mitigating Watermark Forgery in Generative Models via Randomized Key Selection (Aremu et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07871&sa=D&source=editors&ust=1779048535961319&usg=AOvVaw3_EIFeaIlEDrtj4fkSZiFZ) |
| 199 | [▪ Watermarking Diffusion Language Models (Gloaguen et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24368&sa=D&source=editors&ust=1779048535961399&usg=AOvVaw0AWzyhpEAojx1S8P5u78qy) |
| 200 | [▪ Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size (Tondi, Costanzo, and Barni, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24823&sa=D&source=editors&ust=1779048535961486&usg=AOvVaw2-nlOF9AiQ8v3mLT2iEhSc) |
| 201 | [▪ PRIVMARK: Private Large Language Models Watermarking with MPC (Fargues et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24624&sa=D&source=editors&ust=1779048535961566&usg=AOvVaw3gqWu2Vs5We0J3XFIf9mev) |
| 202 | [▪ Analyzing and Evaluating Unbiased Language Model Watermark (Wu et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24048&sa=D&source=editors&ust=1779048535961646&usg=AOvVaw1KM3UCfbuyNJ0XyOjJMbrf) |
| 203 | [▪ An Ensemble Framework for Unbiased Language Model Watermarking (Wu et al., September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.24043&sa=D&source=editors&ust=1779048535961735&usg=AOvVaw0vBQMKYNmFzRY3WMe2SZ3K) |
| 204 | [▪ LLM Watermark Evasion via Bias Inversion (Hwang, Park, and Ok, September 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2509.23019&sa=D&source=editors&ust=1779048535961818&usg=AOvVaw3k8YOGATZPSASm6XWUgFmO) |
| 205 | [▪ FakeIDet: Exploring Patches for Privacy-Preserving Fake ID Detection (Mu\\~noz-Haro et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.07761&sa=D&source=editors&ust=1779048535961900&usg=AOvVaw08HWiT885eVgnnJg0YtaLA) |
| 206 | [▪ Efficient and Universal Watermarking for LLM-Generated Code Detection (Li et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.07518&sa=D&source=editors&ust=1779048535961982&usg=AOvVaw0VgWZ75_KmA9zAmZOZvBr8) |
| 207 | [▪ Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos (Pedrouzo-Rodriguez et al., August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.00748&sa=D&source=editors&ust=1779048535962067&usg=AOvVaw0UJNG90--EiJbXDEvcINg3) |
| 208 | [▪ Towards Privacy-preserving Photorealistic Self-avatars in Mixed Reality (Wilson et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22153&sa=D&source=editors&ust=1779048535962147&usg=AOvVaw03zpQ3ctvCAIhwNPnW_vQ5) |
| 209 | [▪ MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models (Mao, Tsai, and Lu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21195&sa=D&source=editors&ust=1779048535962237&usg=AOvVaw12MoNfWMuFqdbD7AdPhKMD) |
| 210 | [▪ Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems (Khan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21157&sa=D&source=editors&ust=1779048535962335&usg=AOvVaw0CPzIE7C3fK_jMDQuyIUIP) |
| 211 | [▪ WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes (Pujari and Rattani, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21150&sa=D&source=editors&ust=1779048535962428&usg=AOvVaw32zV06jW1mpYj35tbkaQup) |
| 212 | [▪ Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge (Cui et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.04036&sa=D&source=editors&ust=1779048535962512&usg=AOvVaw06yxJwrzbQ6QU-npInCCge) |
| 213 | [▪ Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection (Kheir et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.20417&sa=D&source=editors&ust=1779048535962596&usg=AOvVaw0Kgz6PXNE9PYnb7UqOCGfH) |
| 214 | [▪ VENENA: A Deceptive Visual Encryption Framework for Wireless Semantic Secrecy (Han, Yuan, and Schotten, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.10699&sa=D&source=editors&ust=1779048535962685&usg=AOvVaw0ab9EKHjEiJY1CKtcGEnJl) |
| 215 | [▪ AEDR: Training-Free AI-Generated Image Attribution via Autoencoder Double-Reconstruction (Wang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18988&sa=D&source=editors&ust=1779048535962771&usg=AOvVaw1FBkCDSi8LS77KvBj4K9j9) |
| 216 | [▪ NWaaS: Nonintrusive Watermarking as a Service for X-to-Image DNN (An et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18036&sa=D&source=editors&ust=1779048535962852&usg=AOvVaw0gnbthC6fEXYaAUGvKcwks) |
| 217 | [▪ Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering (An et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18034&sa=D&source=editors&ust=1779048535962935&usg=AOvVaw2vZRiKmjKURfaw9gePlzMg) |
| 218 | [▪ Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs (Islam, Vo, and Rane, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.17010&sa=D&source=editors&ust=1779048535963044&usg=AOvVaw1SEG18TuqAxwGAU2TT3r0Q) |
| 219 | [▪ Watermark Anything with Localized Messages (Sander et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.07231&sa=D&source=editors&ust=1779048535963144&usg=AOvVaw3CHQZp0XB9tqMzcw-6mnM6) |
| 220 | [▪ LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech (Liu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16220&sa=D&source=editors&ust=1779048535963235&usg=AOvVaw0FHYDPgd-W64BK6LK5ki7S) |
| 221 | [▪ Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and Profit (Cuevas, Ribeiro, and Christin, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16045&sa=D&source=editors&ust=1779048535963319&usg=AOvVaw2sRN-kRRvELXWvWYX7XVOl) |
| 222 | [▪ PhishIntentionLLM: Uncovering Phishing Website Intentions through Multi-Agent Retrieval-Augmented Generation (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15419&sa=D&source=editors&ust=1779048535963420&usg=AOvVaw1gqtjm0R7eopH6q493o3zh) |
| 223 | [▪ IConMark: Robust Interpretable Concept-Based Watermark For AI Images (Sadasivan, Saberi, and Feizi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.13407&sa=D&source=editors&ust=1779048535963522&usg=AOvVaw11lSW1Q3BOz2xWXlyoNUNu) |
| 224 | [▪ How does Watermarking Affect Visual Language Models in Document Understanding? (Xu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.01048&sa=D&source=editors&ust=1779048535963614&usg=AOvVaw2x-ZwSmi80V970HC3tz6bP) |
| 225 | [▪ Dynamic Risk Assessments for Offensive Cybersecurity Agents (Wei et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.18384&sa=D&source=editors&ust=1779048535963712&usg=AOvVaw3ueqyA-m18br2ax6B8jpOF) |
| 226 | [▪ A Survey on Speech Deepfake Detection (Li, Ahmadiadli, and Zhang, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2404.13914&sa=D&source=editors&ust=1779048535963803&usg=AOvVaw0vM8qzCAJsJzqe_znjAtVD) |
| 227 | [▪ Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking (Yao, Song, and Jin, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.11137&sa=D&source=editors&ust=1779048535963900&usg=AOvVaw1ltr1A2-AEaj7_rTHeZA1P) |
| 228 | [▪ Watermarking Degrades Alignment in Language Models: Analysis and Mitigation (Verma, Phan, and Trivedi, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.04462&sa=D&source=editors&ust=1779048535964012&usg=AOvVaw1boQW2Qf_Hl9GDbTSNwfVu) |
| 229 | [▪ Mitigating Watermark Stealing Attacks in Generative Models via Multi-Key Watermarking (Aremu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07871&sa=D&source=editors&ust=1779048535964130&usg=AOvVaw2vRWt8S-PQJ7eAx3oCKfUd) |
| 230 | [▪ Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models (Thapa et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.07406&sa=D&source=editors&ust=1779048535964219&usg=AOvVaw077d2JGOP81UCVOUBGVFxG) |
| 231 | [▪ Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks (Shen, Huang, and Wan, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06274&sa=D&source=editors&ust=1779048535964302&usg=AOvVaw0I2RykSG1ZfoxfRBFjMHNu) |
| 232 | [https://arxiv.org/abs/2411.05091](https://www.google.com/url?q=https://arxiv.org/abs/2411.05091&sa=D&source=editors&ust=1779048535964378&usg=AOvVaw3yMmXho8OdfIOB5Bdudloi) |
| 233 | [▪ Removing Watermarks with Partial Regeneration using Semantic Information\<br>\<br> (Tallam et al., May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.08234&sa=D&source=editors&ust=1779048535964506&usg=AOvVaw3vKqWMs70yLu-aqSJWvKxB) |
| 234 | [▪ LLM-Text Watermarking based on Lagrange Interpolation\<br>\<br> (Janas, Morawiecki, and Pieprzyk, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.05712&sa=D&source=editors&ust=1779048535964604&usg=AOvVaw0yPoEBCS2YztRnWUpAtwcI) |
| 235 | [▪ VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning (Hu et al, May 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2505.12332&sa=D&source=editors&ust=1779048535964762&usg=AOvVaw0xTiKkI7YbLcgve11rUDnb) |
| 236 | [▪ AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection (Gao et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.21044&sa=D&source=editors&ust=1779048535964889&usg=AOvVaw0vXEtjPhVkCXdgAwVh6fRF) |
| 237 | [▪ Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning (An et al, Apr 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.06575%23&sa=D&source=editors&ust=1779048535964991&usg=AOvVaw1J6bE7jBeIA8mYZtEhFJ76) |
| 238 | [▪ Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking (Xu et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2503.04636&sa=D&source=editors&ust=1779048535965114&usg=AOvVaw1PrSU2BElVIWvp2cHW_X3S) |
| 239 | [▪ Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances (Lu et al, Mar 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.18775&sa=D&source=editors&ust=1779048535965219&usg=AOvVaw1KHOPNDaEurwRZYmhn9JZC) |
| 240 | [▪ Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approach (He et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02890&sa=D&source=editors&ust=1779048535965326&usg=AOvVaw3-yNdG4FlxC_Ld2lYPNmk9) |
| 241 | [▪ ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation (Teng et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2502.08097&sa=D&source=editors&ust=1779048535965464&usg=AOvVaw0QgW1a1qNPgub5Tk7D_iO8) |
| 242 | [▪ Watermarking Language Models with Error Correcting Codes (Chao et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2406.10281&sa=D&source=editors&ust=1779048535965562&usg=AOvVaw1x-tLJCNkElKmTJevPXGzS) |
| 243 | [▪ Is The Watermarking Of LLM-Generated Code Robust? (Suresh et al, Feb 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2403.17983&sa=D&source=editors&ust=1779048535965759&usg=AOvVaw3oTLC4cpnD3rZixyONYOvl) |
| 244 | [▪ Provably Robust Multi-bit Watermarking for AI-generated Text (Qu et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2401.16820&sa=D&source=editors&ust=1779048535965926&usg=AOvVaw32UwhC7Qh4VBI_f_jBYMYt) |
| 245 | [▪ Audio-Visual Deepfake Detection With Local Temporal Inconsistencies (Astrid, Ghorbel, and Aouada, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.08137&sa=D&source=editors&ust=1779048535966040&usg=AOvVaw0M5qRXYlN5bntu9aYOaVRO) |
| 246 | [▪ GaussMark: A Practical Approach for Structural Watermarking of Language Models (Block, Sekhari, and Rakhlin, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.13941&sa=D&source=editors&ust=1779048535966147&usg=AOvVaw3GB-BtKVVTEKXL6HUav1PZ) |
| 247 | [▪ Neural Honeytrace: A Robust Plug-and-Play Watermarking Framework against Model Extraction Attacks (Xu et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.09328&sa=D&source=editors&ust=1779048535966287&usg=AOvVaw1FjchSd5JuUSckAUqoAHkZ) |
| 248 | [▪ ModelShield: Adaptive and Robust Watermark against Model Extraction Attack (Pang et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2405.02365&sa=D&source=editors&ust=1779048535966393&usg=AOvVaw3xsdl6B1w2laf115SSigPa) |
| 249 | [▪ Can Watermarked LLMs be Identified by Users via Crafted Prompts? (Liu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.03168&sa=D&source=editors&ust=1779048535966503&usg=AOvVaw0qYwxmiMnUIbVx8oQlbWZF) |
| 250 | [▪ Watermarking Graph Neural Networks via Explanations for Ownership Protection (Downer et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.05614&sa=D&source=editors&ust=1779048535966607&usg=AOvVaw0Wkh0MHq_AiEm5AzjupAiu) |
| 251 | [▪ AI-generated Image Detection: Passive or Watermark? (Guo et al, Jan 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2411.13553&sa=D&source=editors&ust=1779048535966711&usg=AOvVaw2hfmKWQIm6T-AOG82mq11r) |
| 252 | [▪ Mesh Watermark Removal Attack and Mitigation: A Novel Perspective of Function Space (Zhu et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2311.12059&sa=D&source=editors&ust=1779048535966811&usg=AOvVaw0qFoTFPDI9d9la2OJHosxi) |
| 253 | [▪ PersonaMark: Personalized LLM watermarking for model protection and user attribution (Zhang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.09739&sa=D&source=editors&ust=1779048535966908&usg=AOvVaw0v2m3gNVRCguDcZ68W8jnB) |
| 254 | [▪ WaterPark: A Robustness Assessment of Language Model Watermarking (Liang et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.13425&sa=D&source=editors&ust=1779048535967003&usg=AOvVaw1cgWxIYLjqMe_EIe1mp2pb) |
| 255 | [▪ BlockDoor: Blocking Backdoor Based Watermarks in Deep Neural Networks (Puah et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.12194&sa=D&source=editors&ust=1779048535967098&usg=AOvVaw07MqU4s0pMrat8kGYKVe5q) |
| 256 | [▪ Towards Effective User Attribution for Latent Diffusion Models via Watermark-Informed Blending (Pan et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10958&sa=D&source=editors&ust=1779048535967198&usg=AOvVaw0EONMPTVUiz41cVgtVZ5jf) |
| 257 | [▪ GENIE: Watermarking Graph Neural Networks for Link Prediction (Bachina et al, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.04805&sa=D&source=editors&ust=1779048535967293&usg=AOvVaw2KsS0uKI29X3UIRA0mNdev) |
| 258 | [▪ The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis (Wu and Chandrasekaran, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.02576&sa=D&source=editors&ust=1779048535967393&usg=AOvVaw1d15tXSQ8LjgOAkWT2nJLU) |
| 259 | [▪ SoK: Watermarking for AI-Generated Content (Zhao et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.18479&sa=D&source=editors&ust=1779048535967506&usg=AOvVaw0gwvSVnzxUHmfdewklqH03) |
| 260 | [▪ Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey (Nguyen-Le et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.17911&sa=D&source=editors&ust=1779048535967607&usg=AOvVaw3k1TXnThQZMEwTnRZUuUD_) |
| 261 | [▪ CLUE-MARK: Watermarking Diffusion Models using CLWE (Shehata, Kolluri, and Saxena, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.11434&sa=D&source=editors&ust=1779048535967708&usg=AOvVaw2tEKBZcPq1jyziJ8aYT4af) |
| 262 | [▪ UnMarker: A Universal Attack on Defensive Image Watermarking (Kasiss and Hengartner, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2405.08363&sa=D&source=editors&ust=1779048535967807&usg=AOvVaw3gYQkTFT_GwuEjF9dF-ix7) |
| 263 | [▪ SoK: On the Role and Future of AIGC Watermarking in the Era of Gen-AI (Ren et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.11478&sa=D&source=editors&ust=1779048535967905&usg=AOvVaw2g16n71HTXNuU8OY1WLhJT) |
| 264 | [▪ Conceptwm: A Diffusion Model Watermark for Concept Protection (Lei et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.11688&sa=D&source=editors&ust=1779048535968001&usg=AOvVaw0G7_0rAjP2i4zVETnx8LeX) |
| 265 | [▪ Watermark-based Detection and Attribution of AI-Generated Content (Jiang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/html/2404.04254v1%23:~:text%3Dthe%2520watermark%2520encoder.-,A%2520content%2520is%2520detected%2520as%2520generated%2520by%2520the%2520GenAI%2520service,similar%2520to%2520the%2520decoded%2520one.%26text%3DTheory.&sa=D&source=editors&ust=1779048535968096&usg=AOvVaw1w9TTyXD6Tl4yH-jV3Au-U) |
| 266 | [▪ An undetectable watermark for generative image models (Gunn, Zhao, and Song, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07369&sa=D&source=editors&ust=1779048535968218&usg=AOvVaw3dKNjnf02sCEf_2miVmBB1) |
| 267 | [▪ InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance (Xu et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2411.07795&sa=D&source=editors&ust=1779048535968318&usg=AOvVaw1elpgCZP6pbczJydnzZCzo) |
| 268 | [▪ FoldMark: Protecting Protein Generative Models with Watermarking (Zhang et al, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.20354&sa=D&source=editors&ust=1779048535968413&usg=AOvVaw1x3Z-9uNlITQFciIZPe5C1) |
| 269 | [▪ Invisible Image Watermarks Are Provably Removable Using Generative AI (Zhao et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2306.01953&sa=D&source=editors&ust=1779048535968509&usg=AOvVaw2y35VuKR3SlrzOGwI5ZBkA) |
| 270 | [▪ Embedding Watermarks in Diffusion Process for Model Intellectual Property Protection (Yang, Peng, and Xia, Nov 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.22445&sa=D&source=editors&ust=1779048535968608&usg=AOvVaw2a_xuo7LkEh_hHjK2J-dgz) |
| 271 | [▪ Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature (Zhou et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2406.01946&sa=D&source=editors&ust=1779048535968718&usg=AOvVaw16vHn3BHrPeazrpr1e50_T) |
| 272 | [▪ Watermarking Large Language Models and the Generated Content: Opportunities and Challenges (Zhou and Koushanfar, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.19096&sa=D&source=editors&ust=1779048535968826&usg=AOvVaw0rE6royT9y0ubV4Z1q7Tom) |
| 273 | [▪ An undetectable watermark for generative image models (Gunn, Zhao, and Song, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07369&sa=D&source=editors&ust=1779048535968950&usg=AOvVaw2a-dwGsfzvu7LNJDxEgr_b) |
| 274 | [▪ Deepfake detection in videos with multiple faces using geometric-fakeness features (Vyshegorodtsev et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.07888&sa=D&source=editors&ust=1779048535969070&usg=AOvVaw3HmGcqzrjlya7R8vBoa0HD) |
| 275 | [▪ Universally Optimal Watermarking Schemes for LLMs: from Theory to Practice (He et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02890&sa=D&source=editors&ust=1779048535969199&usg=AOvVaw0n4RpWBlReeV3bJvwcuiLW) |
| 276 | [▪ Discovering Clues of Spoofed LM Watermarks (Gloaguen et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02693&sa=D&source=editors&ust=1779048535969299&usg=AOvVaw2X9pc3nPLLkus0mX3L-DUt) |
| 277 | [▪ Multi-Designated Detector Watermarking for Language Models (Huang et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.17518&sa=D&source=editors&ust=1779048535969423&usg=AOvVaw2hlUVI1cR7GcSvVY5i2Lfc) |
| 278 | [▪ Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection (PN et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.06543&sa=D&source=editors&ust=1779048535969542&usg=AOvVaw2T26xY-rYzrM-djY1SMrki) |
| 279 | [▪ Signal Watermark on Large Language Models (Zu and Sheng, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.06545&sa=D&source=editors&ust=1779048535969644&usg=AOvVaw1Ot-CbXylgM37u6N-ym0JG) |
| 280 | [▪ Diffuse or Confuse: A Diffusion Deepfake Speech Dataset (Firc, Malinka, and Hanáček, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.06796&sa=D&source=editors&ust=1779048535969800&usg=AOvVaw24yoj7DGovYD-xnrkNDTUL) |
| 281 | [▪ A Watermark for Black-Box Language Models (Bahri et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02099&sa=D&source=editors&ust=1779048535969931&usg=AOvVaw091FA9aYoM4XkNl0ZiA4nA) |
| 282 | [▪ Optimizing Adaptive Attacks against Content Watermarks for Language Models (Diaa, Aremu, and Lukas, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.02440&sa=D&source=editors&ust=1779048535970071&usg=AOvVaw2ZLBsGSFiaLWM_1_sjEWMG) |
| 283 | [▪ Social Media Authentication and Combating Deepfakes using Semi-fragile Invisible Image Watermarking (Nadimpalli and Rattani, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.01906&sa=D&source=editors&ust=1779048535970183&usg=AOvVaw2CUNj4Fg987qVO3YzMKEha) |
| 284 | [▪ PITCH: AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response (Mittal et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2402.18085&sa=D&source=editors&ust=1779048535970300&usg=AOvVaw3vj1R6kf_uMFfM7S3cezv3) |
| 285 | [▪ Shaking the Fake: Detecting Deepfake Videos in Real Time via Active Probes (Xie and Luo, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10889&sa=D&source=editors&ust=1779048535970430&usg=AOvVaw158jYBMNdjh2Sh2NUCB_l3) |
| 286 | [▪ XAI-Based Detection of Adversarial Attacks on Deepfake Detectors (Pinhasov et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2403.02955%23&sa=D&source=editors&ust=1779048535970541&usg=AOvVaw0-4fE5-iAodr-8LqEMSzJb) |

|     |
| --- |
| Deep Fakes, Content Provenance, and Watermarking |

**>**

**<**

‍

#### Shallow Fakes

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Shallow Fakes

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |

|     |
| --- |
| Shallow Fakes |

**>**

**<**

‍

#### Misidentification

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Misidentification

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |

|     |
| --- |
| Misidentification |

**>**

**<**

‍

#### Private Information Used in Training

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Private Information Used in Training

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Efficient and High-Accuracy Private CNN Inference with Helper-Assisted Malicious Security (Wang et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09607&sa=D&source=editors&ust=1779048535980109&usg=AOvVaw0xhY33iu1ZvKEIxX6lGw4J) |
| 3 | [▪ Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation (Ma and Rajtmajer, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.07486&sa=D&source=editors&ust=1779048535980325&usg=AOvVaw3WQeNIQVcqx6xJrAI8O7WR) |
| 4 | [▪ A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset (Hong et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.17185&sa=D&source=editors&ust=1779048535980467&usg=AOvVaw3Sw_g2hg0JuiQ5rkPOh4rs) |
| 5 | [▪ Opal: Private Memory for Personal AI (Kaviani et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02522&sa=D&source=editors&ust=1779048535980569&usg=AOvVaw3Mw3AbsGJoboiNWEfiC4q0) |
| 6 | [▪ Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning (Rinberg et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2506.12553&sa=D&source=editors&ust=1779048535980668&usg=AOvVaw3eyjF3S89oEkGKWgA5D52x) |
| 7 | [▪ Combating Data Laundering in LLM Training (Li et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.01904&sa=D&source=editors&ust=1779048535980755&usg=AOvVaw3vxe-Lnfn-rdk2hUcy4kHG) |
| 8 | [▪ Quantifying Memorization and Privacy Risks in Genomic Language Models (Nemecek et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.08913&sa=D&source=editors&ust=1779048535980843&usg=AOvVaw03xSGpUtaNJXryuUpB4rYy) |
| 9 | [▪ The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage (Choudhari and Singh, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.00061&sa=D&source=editors&ust=1779048535980935&usg=AOvVaw322wnrl_jrLL7vZ99V5L-k) |
| 10 | [▪ Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.05635&sa=D&source=editors&ust=1779048535981031&usg=AOvVaw0Brxr9ZiGTJ3dFtuw2Ru3z) |
| 11 | [▪ UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk (Jeon, Kwon, and Koo, January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.01786&sa=D&source=editors&ust=1779048535981121&usg=AOvVaw1bT-B6OPqT-Zf6HDV_kfnV) |
| 12 | [▪ Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models (Qian et al., December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.18035&sa=D&source=editors&ust=1779048535981210&usg=AOvVaw2hp7RDi9XYN02SrNyuhBge) |
| 13 | [▪ DeepShare: Sharing ReLU Across Channels and Layers for Efficient Private Inference (Bornfeld and Avidan, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.17398&sa=D&source=editors&ust=1779048535981304&usg=AOvVaw0r3XInrcBSihZo7LdpDK5y) |
| 14 | [▪ Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models (Catal, Rani, and Gall, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.11482&sa=D&source=editors&ust=1779048535981413&usg=AOvVaw3fxxtUaHvZR8uiMll8ySJc) |
| 15 | [▪ Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs (Joshi and Smith, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.03310&sa=D&source=editors&ust=1779048535981522&usg=AOvVaw0vvpXMU4nM3cXo4yUhzVOH) |
| 16 | [▪ Multi-PA: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models (Zhang et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2412.19496&sa=D&source=editors&ust=1779048535981621&usg=AOvVaw1WoYJWuF7IewP9u_qYu91e) |
| 17 | [▪ How do data owners say no? A case study of data consent mechanisms in web-scraped vision-language AI training datasets (Lee et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2511.08637&sa=D&source=editors&ust=1779048535981718&usg=AOvVaw2vdildXBz5w2rNbSNUJ_rO) |
| 18 | [▪ Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion (Ke et al., November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2402.18905&sa=D&source=editors&ust=1779048535981815&usg=AOvVaw1K5YdxH5Z_y3vEvK_b3N_2) |
| 19 | [▪ AERO: Entropy-Guided Framework for Private LLM Inference (Jha and Reagen, November 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2410.13060&sa=D&source=editors&ust=1779048535981913&usg=AOvVaw165IQZ43qD4nFbjhH0sCO0) |
| 20 | [▪ Toward provably private analytics and insights into GenAI use (Cheu et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.21684&sa=D&source=editors&ust=1779048535982007&usg=AOvVaw3BhkGNnwzfVjsUBFD_lLZX) |
| 21 | [▪ SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority Oversampling (Ganev et al., October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.15083&sa=D&source=editors&ust=1779048535982111&usg=AOvVaw3ffg-qCzoW7cdAVZJM3CXq) |
| 22 | [▪ T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.12612&sa=D&source=editors&ust=1779048535982195&usg=AOvVaw3_acFETx3WGX97KtltmtED) |
| 23 | [▪ A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census: Full Technical Report (Abowd et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2312.11283&sa=D&source=editors&ust=1779048535982284&usg=AOvVaw2tjjhzDYrC4pKcpmFAwH3i) |
| 24 | [▪ IDFace: Face Template Protection for Efficient and Secure Identification (Kim et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12050&sa=D&source=editors&ust=1779048535982369&usg=AOvVaw20AGhHC4nQedSX_CmA2tsw) |
| 25 | [▪ Pantomime: Motion Data Anonymization using Foundation Motion Models (Hanisch, Todt, and Strufe, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2501.07149&sa=D&source=editors&ust=1779048535982460&usg=AOvVaw2SGxca8PKETiuX-5UBHsHl) |
| 26 | [▪ Predicting memorization within Large Language Models fine-tuned for classification (Dentan et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2409.18858&sa=D&source=editors&ust=1779048535982545&usg=AOvVaw32UAhRfuQKB_OaMSgRyAKp) |
| 27 | [▪ "Is it always watching? Is it always listening?" Exploring Contextual Privacy and Security Concerns Toward Domestic Social Robots (Bell et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10786&sa=D&source=editors&ust=1779048535982633&usg=AOvVaw145yhhyLWMZMpaFIV1yFje) |
| 28 | [▪ Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models (Huang et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.08878&sa=D&source=editors&ust=1779048535982718&usg=AOvVaw3hNLOQ-gLcR_48OMWN5YlC) |
| 29 | [▪ RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service (Cheng, Chow and Li, Dec 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2412.12775&sa=D&source=editors&ust=1779048535982800&usg=AOvVaw1rABUruufkYAH0ko5y64pt) |
| 30 | [▪ Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers (Burgess et al, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.16705&sa=D&source=editors&ust=1779048535982886&usg=AOvVaw2dG_2JwZqWrEE5xhwS8awF) |
| 31 | [▪ Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data (Akkus et al, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.11423&sa=D&source=editors&ust=1779048535982976&usg=AOvVaw1oQOpvg0eFlWXFm5DScZcA) |
| 32 | [▪ Catch Me if You Can: Detecting Unauthorized Data Use in Deep Learning Models (Chen and Pattabiraman, Sep 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.06280&sa=D&source=editors&ust=1779048535983081&usg=AOvVaw1rElSRPN8qDoDCAs030JYA) |
| 33 | [▪ Ethical Challenges in Computer Vision: Ensuring Privacy and Mitigating Bias in Publicly Available Datasets (Tahir, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2409.10533&sa=D&source=editors&ust=1779048535983177&usg=AOvVaw1QC17k4Z8tTqOr9Ipy7kGC) |
| 34 | [▪ Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions (Liu and Yang, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.10468&sa=D&source=editors&ust=1779048535983271&usg=AOvVaw2IjcYsYBtZAvEUPm6DbO_4) |

|     |
| --- |
| Private Information Used in Training |

**>**

**<**

‍

#### Unsecured Credentials

Covers:

- MITRE ATLAS Credential Access

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : Unsecured Credentials

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study (Chen et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.03070&sa=D&source=editors&ust=1779048536736887&usg=AOvVaw00aAr9vUbUxH7zWM70v2-R) |
| 3 | [▪ Keys on Doormats: Exposed API Credentials on the Web (Demir et al., March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.12498&sa=D&source=editors&ust=1779048536737101&usg=AOvVaw1lraUlqt4GMO1fT7wGXrqS) |

|     |
| --- |
| Unsecured Credentials |

**>**

**<**

‍

#### AI-Generated/Augmented Exploits

Added this category to cover instances where generative AI systems are used to generate cybersecurity exploits.

Research:

cybersecurity\_tracker - Google Drive

cybersecurity\_tracker : AI-Generated/Augmented Exploits

cybersecurity\_tracker - Google Drive

|  |  |
| --- | --- |
| 2 | [▪ The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code (Hortea and Tapiador, May 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2605.03619&sa=D&source=editors&ust=1779048536613803&usg=AOvVaw0ID8KkmgPki_PTfRGwk7Bl) |
| 3 | [▪ Evaluating LLM-Generated Obfuscated XSS Payloads for Machine Learning-Based Detection (Gabbireddy and Saha, April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.19526&sa=D&source=editors&ust=1779048536613914&usg=AOvVaw21p1MCu0q3yOFQBMpvljCb) |
| 4 | [▪ RedShell: A Generative AI-Based Approach to Ethical Hacking (Bessa et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.11506&sa=D&source=editors&ust=1779048536614002&usg=AOvVaw3iqIFToIB5h10CYwTh9Spb) |
| 5 | [▪ From Theory to Practice: Code Generation Using LLMs for CAPEC and CWE Frameworks (Shahzad et al., April 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2604.02548&sa=D&source=editors&ust=1779048536614077&usg=AOvVaw2LV1oIMNN_H3BfXigC-2nP) |
| 6 | [▪ Automatic Attack Script Generation: a MDA Approach (Goux and Lammari, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.11861&sa=D&source=editors&ust=1779048536614147&usg=AOvVaw2cvlxBHY2-tfFogxaaCEYq) |
| 7 | [▪ Synergistic Directed Execution and LLM-Driven Analysis for Zero-Day AI-Generated Malware Detection (Edwards and Eslamimehr, March 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2603.09044&sa=D&source=editors&ust=1779048536614220&usg=AOvVaw3VK3aH6s9bFMcumzH7-KFU) |
| 8 | [▪ Execution-State-Aware LLM Reasoning for Automated Proof-of-Vulnerability Generation (Li et al., February 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2602.13574&sa=D&source=editors&ust=1779048536614290&usg=AOvVaw1CMo6Gq34YpoOqQ5kGq49Q) |
| 9 | [▪ KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation (Liu et al., January 2026)](https://www.google.com/url?q=https://arxiv.org/abs/2601.09129&sa=D&source=editors&ust=1779048536614368&usg=AOvVaw30N5jZC4OvyMF3AV3B9Un6) |
| 10 | [▪ LLM-based Vulnerable Code Augmentation: Generate or Refactor? (Ouchebara and Dupont, December 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2512.08493&sa=D&source=editors&ust=1779048536614461&usg=AOvVaw3X3s9E8-pHXmnEBVQLQlh-) |
| 11 | [▪ Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositories (Schreiber and Tippe, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.26103&sa=D&source=editors&ust=1779048536614536&usg=AOvVaw2KTdrD3HKLiOpQdeKBBO6l) |
| 12 | [▪ deepSURF: Detecting Memory Safety Vulnerabilities in Rust Through Fuzzing LLM-Augmented Harnesses (Androutsopoulos and Bianchi, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.15648&sa=D&source=editors&ust=1779048536614614&usg=AOvVaw2Yv-JlcLu9rYK1KF5KG9a-) |
| 13 | [▪ PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits (Bhuiya, Aggarwal, and Purwar, October 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2510.17947&sa=D&source=editors&ust=1779048536614698&usg=AOvVaw0zDrr_w296ozVTL841FiRl) |
| 14 | [▪ NATLM: Detecting Defects in NFT Smart Contracts Leveraging LLM (Niu, Li, and Li, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01351&sa=D&source=editors&ust=1779048536614783&usg=AOvVaw3SIA-y3DLkBIqI-c8DBcWp) |
| 15 | [▪ Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs (Bakker and Hastings, August 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2508.01054&sa=D&source=editors&ust=1779048536614870&usg=AOvVaw1Wo7GRSmKBUPr7Evq29cWU) |
| 16 | [▪ Fuzzing: Randomness? Reasoning! Efficient Directed Fuzzing via Large Language Models (Feng et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22065&sa=D&source=editors&ust=1779048536614971&usg=AOvVaw2FdkTmJoUTNXwT3sokam9K) |
| 17 | [▪ SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection (Yu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.22371&sa=D&source=editors&ust=1779048536615052&usg=AOvVaw1xzyGcsJ4QQOFbprWoTda2) |
| 18 | [▪ Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing (Abdulzada, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21113&sa=D&source=editors&ust=1779048536615129&usg=AOvVaw13km4SsmJ5kRa37fNEKVjN) |
| 19 | [▪ Intelligent ARP Spoofing Detection using Multi-layered Machine Learning (ML) Techniques for IoT Networks (Ali, Husain, and Hans, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.21087&sa=D&source=editors&ust=1779048536615230&usg=AOvVaw0JOvWBeursWI3igTHJlYI_) |
| 20 | [▪ Automated Static Vulnerability Detection via a Holistic Neuro-symbolic Approach (Li et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2504.16057&sa=D&source=editors&ust=1779048536615377&usg=AOvVaw3Wqk2S48wTYSCVON3usgHu) |
| 21 | [▪ Are AI-Generated Fixes Secure? Analyzing LLM and Agent Patches on SWE-bench (Sajadi, Damevski, and Chatterjee, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.02976&sa=D&source=editors&ust=1779048536615504&usg=AOvVaw3SbfmSxLJof3T5I10jNfP5) |
| 22 | [▪ Auto-SGCR: Automated Generation of Smart Grid Cyber Range Using IEC 61850 Standard Models (Roomi et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.18249&sa=D&source=editors&ust=1779048536615609&usg=AOvVaw0A1uAC1apbd1KutlBfk13y) |
| 23 | [▪ SynthCTI: LLM-Driven Synthetic CTI Generation to enhance MITRE Technique Mapping (Ruiz-R\\'odenas et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16852&sa=D&source=editors&ust=1779048536615715&usg=AOvVaw2EBS7_rq15-Ar6BbWs3wGE) |
| 24 | [▪ From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction (Lekssays, Sencar, and Yu, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.16576&sa=D&source=editors&ust=1779048536615831&usg=AOvVaw2yLc0P-QPf2pLcKBeEP18b) |
| 25 | [▪ LibLMFuzz: LLM-Augmented Fuzz Target Generation for Black-box Libraries (Hardgrove and Hastings, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.15058&sa=D&source=editors&ust=1779048536615940&usg=AOvVaw0R-949c5uZEa_OrAUCQwal) |
| 26 | [▪ Using Modular Arithmetic Optimized Neural Networks To Crack Affine Cryptographic Schemes Efficiently (Stojanovi\\'c, Lesar, and Bohak, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.14229&sa=D&source=editors&ust=1779048536616066&usg=AOvVaw2pgjjVLF_rSxTkFUiLC_Aw) |
| 27 | [▪ LLAMA: Multi-Feedback Smart Contract Fuzzing Framework with LLM-Guided Seed Generation (Gai et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.12084&sa=D&source=editors&ust=1779048536616187&usg=AOvVaw2z1sylUr01xUEc4EJSsCRe) |
| 28 | [▪ QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration (Hu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2506.23644&sa=D&source=editors&ust=1779048536616308&usg=AOvVaw2IbWCQfLLEw9NnS3qUYcXm) |
| 29 | [▪ MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning (Gajjar, Subramaniakuppusamy, and Kachach, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.10898&sa=D&source=editors&ust=1779048536616430&usg=AOvVaw10cXIz5HqXo88N8Xzv-r4B) |
| 30 | [▪ A Mixture of Linear Corrections Generates Secure Code (Yu et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09508&sa=D&source=editors&ust=1779048536616556&usg=AOvVaw0fiRyV3mKPbC5Z4-0BGdPt) |
| 31 | [▪ LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models (Akil et al., July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.09411&sa=D&source=editors&ust=1779048536616692&usg=AOvVaw1tVXU8iw5O2fgPTzQmnWih) |
| 32 | [▪ PenTest2.0: Towards Autonomous Privilege Escalation Using GenAI (Al-Sinani and Mitchell, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.06742&sa=D&source=editors&ust=1779048536616822&usg=AOvVaw1O-IJcy0jOGKpiGKKhO9xP) |
| 33 | [▪ AI Agent Smart Contract Exploit Generation (Gervais and Zhou, July 2025)](https://www.google.com/url?q=https://arxiv.org/abs/2507.05558&sa=D&source=editors&ust=1779048536616961&usg=AOvVaw07KvWqH_feoyMhfDYRCcc2) |
| 34 | [▪ Metamorphic Malware Evolution: The Potential and Peril of Large Language Models (Madani, Oct 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2410.23894&sa=D&source=editors&ust=1779048536617081&usg=AOvVaw2ng2_hefK7b2TNd0tMCs5m) |
| 35 | [▪ Exploring RAG-based Vulnerability Augmentation with LLMs (Daneshvar et al, Aug 2024)](https://www.google.com/url?q=https://arxiv.org/abs/2408.04125%23&sa=D&source=editors&ust=1779048536617183&usg=AOvVaw1zXa4wwWNVXm5v0yqxZfbg) |

|     |
| --- |
| AI-Generated/Augmented Exploits |

**>**

**<**

‍

‍

‍
