AI Safety+Cybersecurity R&D Tracker

January 1, 2026

‍

Subscribe to our monthly AI Safety and Cybersecurity R&D Tracker updates!

Threats using AI models

Prompt Injection and Input Manipulation (Direct and Indirect)

Covers:

  • OWASP LLM 01: Prompt Injection
  • OWASP ML 01: Input Manipulation Attack
  • MITRE ATLAS Initial Access, Privilege Escalation, and Defense Evasion

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Prompt Injection and Input Manipulation (Direct and Indirect)

cybersecurity_tracker - Google Drive

1 Reference
2 ▪ WARD: Adversarially Robust Defense of Web Agents Against Prompt Injections (Cao et al., May 2026)
3 ▪ Phantom Force: Injecting Adversarial Tactile Perceptions into Embodied Intelligence via EMI (Kong, Zhang, and Chau, May 2026)
4 ▪ Sleeper Channels and Provenance Gates: Persistent Prompt Injection in Always-on Autonomous AI Agents (Maloyan and Namiot, May 2026)
5 ▪ Persona-Conditioned Adversarial Prompting (PCAP): Multi-Identity Red-Teaming for Enhanced Adversarial Prompt Discovery (Morasso et al., May 2026)
6 ▪ Persona-Conditioned Adversarial Prompting: Multi-Identity Red-Teaming for Adversarial Discovery and Mitigation (Morasso et al., May 2026)
7 ▪ IPI-proxy: An Intercepting Proxy for Red-Teaming Web-Browsing AI Agents Against Indirect Prompt Injection (Chia-Pei et al., May 2026)
8 ▪ Adversarial SQL Injection Generation with LLM-Based Architectures (Karakoc and Yilmaz, May 2026)
9 ▪ MCPShield: Content-Aware Attack Detection for LLM Agent Tool-Call Traffic (Zavrak, May 2026)
10 ▪ SAID: Safety-Aware Intent Defense via Prefix Probing for Large Language Models (Chen et al., May 2026)
11 ▪ Preventing Prompt Injection with Type-Directed Privilege Separation (Jacob et al., May 2026)
12 ▪ When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications (Motlagh et al., May 2026)
13 ▪ Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements (Li et al., May 2026)
14 ▪ Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs (Maiorano, May 2026)
15 ▪ One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue (Shen et al., May 2026)
16 ▪ PragLocker: Protecting Agent Intellectual Property in Untrusted Deployments via Non-Portable Prompts (Li et al., May 2026)
17 ▪ Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs (Fei et al., May 2026)
18 ▪ Tailored Prompts, Targeted Protection: Vulnerability-Specific LLM Analysis for Smart Contracts (Zhang et al., May 2026)
19 ▪ Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis (Zhang, Zandsalimy, and Sushmita, May 2026)
20 ▪ A Validated Prompt Bank for Malicious Code Generation: Separating Executable Weapons from Security Knowledge in 1,554 Consensus-Labeled Prompts (Young and Moody, May 2026)
21 ▪ LocalAlign: Enabling Generalizable Prompt Injection Defense via Generation of Near-Target Adversarial Examples for Alignment Training (Gong et al., May 2026)
22 ▪ A Sentence Relation-Based Approach to Sanitizing Malicious Instructions (Datta et al., May 2026)
23 ▪ Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI (Chang et al., May 2026)
24 ▪ FlashRT: Towards Computationally and Memory Efficient Red-Teaming for Prompt Injection and Knowledge Corruption (Wang et al., May 2026)
25 ▪ Indirect Prompt Injection in the Wild: An Empirical Study of Prevalence, Techniques, and Objectives (Khodayari et al., May 2026)
26 ▪ SafeReview: Defending LLM-based Review Systems Against Adversarial Hidden Prompts (Xin et al., April 2026)
27 ▪ "Your AI, My Shell": Demystifying Prompt Injection Attacks on Agentic AI Coding Editors (Liu et al., April 2026)
28 ▪ SnapGuard: Lightweight Prompt Injection Detection for Screenshot-Based Web Agents (Du et al., April 2026)
29 ▪ PARASITE: Conditional System Prompt Poisoning to Hijack LLMs (Pham and Le, April 2026)
30 ▪ Measuring and Exploiting Contextual Bias in LLM-Assisted Security Code Review (Mitropoulos et al., April 2026)
31 ▪ Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models (Rayhan and Jahan, April 2026)
32 ▪ Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation (Zhang et al., April 2026)
33 ▪ XOXO: Stealthy Cross-Origin Context Poisoning Attacks against AI Coding Assistants (\v{S}torek et al., April 2026)
34 ▪ Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection (Munirathinam, April 2026)
35 ▪ CASCADE: A Cascaded Hybrid Defense Architecture for Prompt Injection Detection in MCP-Based Systems (Turgut and G\"um\"u\c{s}, April 2026)
36 ▪ LogJack: Indirect Prompt Injection Through Cloud Logs Against LLM Debugging Agents (Shah, April 2026)
37 ▪ Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection (Chen et al., April 2026)
38 ▪ Random Walk Learning and the Pac-Man Attack (Chen et al., April 2026)
39 ▪ Towards Personalizing Secure Programming Education with LLM-Injected Vulnerabilities (Frazier and Damevski, April 2026)
40 ▪ LLM-Guided Prompt Evolution for Password Guessing (Mazin et al., April 2026)
41 ▪ DeepSeek Robustness Against Semantic-Character Dual-Space Mutated Prompt Injection (Ren et al., April 2026)
42 ▪ WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents (Chen et al., April 2026)
43 ▪ AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption (Wang et al., April 2026)
44 ▪ ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection (Zhao et al., April 2026)
45 ▪ Leave My Images Alone: Preventing Multi-Modal Large Language Models from Analyzing Images via Visual Prompt Injection (Shao et al., April 2026)
46 ▪ ACIArena: Toward Unified Evaluation for Agent Cascading Injection (An et al., April 2026)
47 ▪ PIArena: A Platform for Prompt Injection Evaluation (Geng et al., April 2026)
48 ▪ Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection (Yang et al., April 2026)
49 ▪ Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (Zhao et al., April 2026)
50 ▪ AttackEval: A Systematic Empirical Study of Prompt Injection Attack Effectiveness Against Large Language Models (Wang, April 2026)
51 ▪ AgentWatcher: A Rule-based Prompt Injection Monitor (Wang et al., April 2026)
52 ▪ Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks (Collu et al., March 2026)
53 ▪ Kill-Chain Canaries: Stage-Level Tracking of Prompt Injection Across Attack Surfaces and Model Safety Tiers (Wang, March 2026)
54 ▪ Epistemic Bias Injection: Biasing LLMs via Selective Context Retrieval (Wu and Saxena, March 2026)
55 ▪ PIDP-Attack: Combining Prompt Injection with Database Poisoning Attacks on Retrieval-Augmented Generation Systems (Wang et al., March 2026)
56 ▪ SUAD: Solid-Channel Ultrasound Injection Attack and Defense to Voice Assistants (Liu et al., March 2026)
57 ▪ Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search (Shen et al., March 2026)
58 ▪ The Cognitive Firewall:Securing Browser Based AI Agents Against Indirect Prompt Injection Via Hybrid Edge Cloud Defense (Lan and Kaul, March 2026)
59 ▪ Model Context Protocol Threat Modeling and Analyzing Vulnerabilities to Prompt Injection with Tool Poisoning (Huang et al., March 2026)
60 ▪ Are AI-assisted Development Tools Immune to Prompt Injection? (Huang, Huang, and Fard, March 2026)
61 ▪ Cross-site scripting adversarial attacks based on deep reinforcement learning: Evaluation and extension study (Pasini et al., March 2026)
62 ▪ Prompt Control-Flow Integrity: A Priority-Aware Runtime Defense Against Prompt Injection in LLM Systems (Alam et al., March 2026)
63 ▪ Detecting Sentiment Steering Attacks on RAG-enabled Large Language Models (Andrade et al., March 2026)
64 ▪ How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition (Dziemian et al., March 2026)
65 ▪ SKILLS: Structured Knowledge Injection for LLM-Driven Telecommunications Operations (Brett, March 2026)
66 ▪ Agent Privilege Separation in OpenClaw: A Structural Defense Against Prompt Injection (Cheng and Tsao, March 2026)
67 ▪ PILOT: Command-line Interface Fuzzing via Path-Guided, Iterative Large Language Model Prompting (Shiraishi, Cao, and Shinagawa, March 2026)
68 ▪ PISmith: Reinforcement Learning-based Red Teaming for Prompt Injection Defenses (Yin et al., March 2026)
69 ▪ Prompt Injection as Role Confusion (Ye, Cui, and Hadfield-Menell, March 2026)
70 ▪ The Mirror Design Pattern: Strict Data Geometry over Model Scale for Prompt Injection Detection (Corll, March 2026)
71 ▪ AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations (He et al., March 2026)
72 ▪ Image-based Prompt Injection: Hijacking Multimodal LLMs through Visually Embedded Adversarial Instructions (Nagaraja et al., March 2026)
73 ▪ VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents (Cao et al., March 2026)
74 ▪ Reverse CAPTCHA: Evaluating LLM Susceptibility to Invisible Unicode Instruction Injection (Graves, March 2026)
75 ▪ AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purification (Zhang et al., February 2026)
76 ▪ Silent Egress: When Implicit Prompt Injection Makes LLM Agents Leak Without a Trace (Lan et al., February 2026)
77 ▪ ICON: Indirect Prompt Injection Defense for Agents based on Inference-Time Correction (Wang et al., February 2026)
78 ▪ AdapTools: Adaptive Tool-based Indirect Prompt Injection Attacks on Agentic LLMs (Wang et al., February 2026)
79 ▪ ICSSPulse: A Modular LLM-Assisted Platform for Industrial Control System Penetration Testing (Takaronis et al., February 2026)
80 ▪ Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks (Schmotz et al., February 2026)
81 ▪ Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Models (Wirth, February 2026)
82 ▪ The Vulnerability of LLM Rankers to Prompt Injection Attacks (Yin et al., February 2026)
83 ▪ Can Adversarial Code Comments Fool AI Security Reviewers -- Large-Scale Empirical Study of Comment-Based Attacks and Defenses Against LLM Code Analysis (Thornton, February 2026)
84 ▪ SkillJect: Automating Stealthy Skill-Based Prompt Injection for Coding Agents with Trace-Driven Closed-Loop Refinement (Jia et al., February 2026)
85 ▪ AlignSentinel: Alignment-Aware Detection of Prompt Injection Attacks (Jia et al., February 2026)
86 ▪ SAFuzz: Semantic-Guided Adaptive Fuzzing for LLM-Generated Code (Yang et al., February 2026)
87 ▪ When Skills Lie: Hidden-Comment Injection in LLM Agents (Wang et al., February 2026)
88 ▪ Protecting Context and Prompts: Deterministic Security for Non-Deterministic AI (Rajagopalan and Rao, February 2026)
89 ▪ The Landscape of Prompt Injection Threats in LLM Agents: From Taxonomy to Analysis (Wang et al., February 2026)
90 ▪ The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multistep Malware Delivery Mechanism (Brodt et al., February 2026)
91 ▪ MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks (Syros et al., February 2026)
92 ▪ Efficient and Adaptable Detection of Malicious LLM Prompts via Bootstrap Aggregation (Hassan et al., February 2026)
93 ▪ CausalArmor: Efficient Indirect Prompt Injection Guardrails via Causal Attribution (Kim et al., February 2026)
94 ▪ Clouding the Mirror: Stealthy Prompt Injection Attacks Targeting LLM-based Phishing Detection (Koide, Nakano, and Chiba, February 2026)
95 ▪ WebSentinel: Detecting and Localizing Prompt Injection Attacks for Web Agents (Wang et al., February 2026)
96 ▪ AgentDyn: A Dynamic Open-Ended Benchmark for Evaluating Prompt Injection Attacks of Real-World Agent Security System (Li et al., February 2026)
97 ▪ Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments (Zhang et al., February 2026)
98 ▪ RedVisor: Reasoning-Aware Prompt Injection Defense via Zero-Copy KV Cache Reuse (Liu et al., February 2026)
99 ▪ Bypassing Prompt Injection Detectors through Evasive Injections (Rahman and Alouani, February 2026)
100 ▪ zkCraft: Prompt-Guided LLM as a Zero-Shot Mutation Pattern Oracle for TCCT-Powered ZK Fuzzing (Fu et al., February 2026)
101 ▪ Whispers of Wealth: Red-Teaming Google's Agent Payments Protocol via Prompt Injection (Debi and Zhu, February 2026)
102 ▪ WADBERT: Dual-channel Web Attack Detection Based on BERT Models (Luo et al., January 2026)
103 ▪ Analysis of LLM Vulnerability to GPU Soft Errors: An Instruction-Level Fault Injection Study (Chai et al., January 2026)
104 ▪ MalURLBench: A Benchmark Evaluating Agents' Vulnerabilities When Processing Web URLs (Kong et al., January 2026)
105 ▪ Prompt Injection Evaluations: Refusal Boundary Instability and Artifact-Dependent Compliance in GPT-4-Series Models (Heverin, January 2026)
106 ▪ Breaking the Protocol: Security Analysis of the Model Context Protocol Specification and Prompt Injection Vulnerabilities in Tool-Integrated LLM Agents (Maloyan and Namiot, January 2026)
107 ▪ Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems (Maloyan and Namiot, January 2026)
108 ▪ Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art (Trinh et al., January 2026)
109 ▪ On the Insecurity of Keystroke-Based AI Authorship Detection: Timing-Forgery Attacks Against Motor-Signal Verification (Condrey, January 2026)
110 ▪ RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models (Chugh, January 2026)
111 ▪ SPECTRE: Conditional System Prompt Poisoning to Hijack LLMs (Pham and Le, January 2026)
112 ▪ LLM Security and Safety: Insights from Homotopy-Inspired Prompt Obfuscation (Lazo, Jelodar, and Razavi-Far, January 2026)
113 ▪ Sockpuppetting: Jailbreaking LLMs Without Optimization Through Output Prefix Injection (Dotsinski and Eustratiadis, January 2026)
114 ▪ PINA: Prompt Injection Attack against Navigation Agents (Liu et al., January 2026)
115 ▪ Adversarial News and Lost Profits: Manipulating Headlines in LLM-Driven Algorithmic Trading (Rizvani, Apruzzese, and Laskov, January 2026)
116 ▪ Zero-Shot Embedding Drift Detection: A Lightweight Defense Against Prompt Injections in LLMs (Sekar et al., January 2026)
117 ▪ SD-RAG: A Prompt-Injection-Resilient Framework for Selective Disclosure in Retrieval-Augmented Generation (Masoud, Arazzi, and Nocera, January 2026)
118 ▪ Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG (Guo and Wei, January 2026)
119 ▪ Reasoning Hijacking: Subverting LLM Classification via Decision-Criteria Injection (Liu, Tang, and Tun, January 2026)
120 ▪ ReasAlign: Reasoning Enhanced Safety Alignment against Prompt Injection Attack (Li et al., January 2026)
121 ▪ The Promptware Kill Chain: How Prompt Injections Gradually Evolved Into a Multi-Step Malware (Nassi, Schneier, and Brodt, January 2026)
122 ▪ Baiting AI: Deceptive Adversary Against AI-Protected Industrial Infrastructures (Pasikhani et al., January 2026)
123 ▪ Small Symbols, Big Risks: Exploring Emoticon Semantic Confusion in Large Language Models (Jiang et al., January 2026)
124 ▪ Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning (Nguyen et al., January 2026)
125 ▪ SecureCAI: Injection-Resilient LLM Assistants for Cybersecurity Operations (Ali et al., January 2026)
126 ▪ How Secure is Secure Code Generation? Adversarial Prompts Put LLM Defenses to the Test (Tessa et al., January 2026)
127 ▪ Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM Systems (Chang et al., January 2026)
128 ▪ Defense Against Indirect Prompt Injection via Tool Result Parsing (Yu, Cheng, and Liu, January 2026)
129 ▪ Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning (Chang et al., January 2026)
130 ▪ Topic-FlipRAG: Topic-Orientated Adversarial Opinion Manipulation Attacks to Retrieval-Augmented Generation Models (Gong et al., December 2025)
131 ▪ Prompt Injection attack against LLM-integrated Applications (Liu et al., December 2025)
132 ▪ Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks (Syed, Almutairi, and Moaty, December 2025)
133 ▪ ChatGPT: Excellent Paper! Accept It. Editor: Imposter Found! Review Rejected (Gharami et al., December 2025)
134 ▪ AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens (Li, Wu, and Liu, December 2025)
135 ▪ From Prompt Injections to Protocol Exploits: Threats in LLM-Powered AI Agents Workflows (Ferrag et al., December 2025)
136 ▪ Detecting Prompt Injection Attacks Against Application Using Classifiers (Shaheer et al., December 2025)
137 ▪ CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks (Wang et al., December 2025)
138 ▪ When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection (Sahoo et al., December 2025)
139 ▪ Llama-based source code vulnerability detection: Prompt engineering vs Fine tuning (Ouchebara and Dupont, December 2025)
140 ▪ ObliInjection: Order-Oblivious Prompt Injection Attack to LLM Agents with Multi-source Data (Wang, Jia, and Gong, December 2025)
141 ▪ Attention is All You Need to Defend Against Indirect Prompt Injection Attacks in LLMs (Zhong et al., December 2025)
142 ▪ Degrading Voice: A Comprehensive Overview of Robust Voice Conversion Through Input Manipulation (Song et al., December 2025)
143 ▪ In-Context Representation Hijacking (Yona et al., December 2025)
144 ▪ HarnessAgent: Scaling Automatic Fuzzing Harness Construction with Tool-Augmented LLM Pipelines (Yang et al., December 2025)
145 ▪ Rethinking Security in Semantic Communication: Latent Manipulation as a New Threat (Xi and Zhu, December 2025)
146 ▪ Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs (Mathew et al., December 2025)
147 ▪ Securing Large Language Models (LLMs) from Prompt Injection Attacks (Suri and McCrae, December 2025)
148 ▪ Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis (Kang et al., December 2025)
149 ▪ BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents (Zhang et al., November 2025)
150 ▪ On the Feasibility of Hijacking MLLMs' Decision Chain via One Perturbation (Li et al., November 2025)
151 ▪ Effective Command-line Interface Fuzzing with Path-Aware Large Language Model Orchestration (Shiraishi, Cao, and Shinagawa, November 2025)
152 ▪ Prompt Fencing: A Cryptographic Approach to Establishing Security Boundaries in Large Language Model Prompts (Peh, November 2025)
153 ▪ RoguePrompt: Dual-Layer Ciphering for Self-Reconstruction to Circumvent LLM Moderation (Tafreshian, November 2025)
154 ▪ PSM: Prompt Sensitivity Minimization via LLM-Guided Black-Box Optimization (Jawad and Brunel, November 2025)
155 ▪ Securing AI Agents Against Prompt Injection Attacks (Ramakrishnan and Balaji, November 2025)
156 ▪ On-Premise SLMs vs. Commercial LLMs: Prompt Engineering and Incident Classification in SOCs and CSIRTs (Almeida et al., November 2025)
157 ▪ Cybersecurity AI: Hacking the AI Hackers via Prompt Injection (Mayoral-Vilches and Rynning, November 2025)
158 ▪ Whose Narrative is it Anyway? A KV Cache Manipulation Attack (Ganesh, Iyer, and Ananthan, November 2025)
159 ▪ GRAPHTEXTACK: A Realistic Black-Box Node Injection Attack on LLM-Enhanced GNNs (Ma, Trivedi, and Koutra, November 2025)
160 ▪ Privacy-Preserving Prompt Injection Detection for LLMs Using Federated Learning and Embedding-Based NLP Classification (Jayathilaka, November 2025)
161 ▪ Prompt Engineering vs. Fine-Tuning for LLM-Based Vulnerability Detection in Solana and Algorand Smart Contracts (Boi and Esposito, November 2025)
162 ▪ PISanitizer: Preventing Prompt Injection to Long-Context LLMs via Prompt Sanitization (Geng et al., November 2025)
163 ▪ BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models (Liu et al., November 2025)
164 ▪ DataSentinel: A Game-Theoretic Detection of Prompt Injection Attacks (Liu et al., November 2025)
165 ▪ MPMA: Preference Manipulation Attack Against Model Context Protocol (Wang et al., November 2025)
166 ▪ Prompt Injection Vulnerability of Consensus Generating Applications in Digital Democracy (Gudi\~no-Rosero et al., November 2025)
167 ▪ Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs (Fastowski et al., November 2025)
168 ▪ Black-Box Guardrail Reverse-engineering Attack (Yao et al., November 2025)
169 ▪ Hybrid Fuzzing with LLM-Guided Input Mutation and Semantic Feedback (Lin, November 2025)
170 ▪ Death by a Thousand Prompts: Open Model Vulnerability Analysis (Chang et al., November 2025)
171 ▪ "Give a Positive Review Only": An Early Investigation Into In-Paper Prompt Injection Attacks and Defenses for AI Reviewers (Zhou et al., November 2025)
172 ▪ Prompt Injection as an Emerging Threat: Evaluating the Resilience of Large Language Models (Ganiuly and Smaiyl, November 2025)
173 ▪ Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization (Verma, November 2025)
174 ▪ Measuring the Security of Mobile LLM Agents under Adversarial Prompts from Untrusted Third-Party Channels (Du et al., November 2025)
175 ▪ Broken-Token: Filtering Obfuscated Prompts by Counting Characters-Per-Token (Zychlinski and Kainan, November 2025)
176 ▪ QueryIPI: Query-agnostic Indirect Prompt Injection on Coding Agents (Xie et al., October 2025)
177 ▪ CompressionAttack: Exploiting Prompt Compression as a New Attack Surface in LLM-Powered Agents (Liu et al., October 2025)
178 ▪ Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies (Wang et al., October 2025)
179 ▪ DRIFT: Dynamic Rule-Based Defense with Injection Isolation for Securing LLM Agents (Li et al., October 2025)
180 ▪ PhantomLint: Principled Detection of Hidden LLM Prompts in Structured Documents (Murray, October 2025)
181 ▪ GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments? (Chen et al., October 2025)
182 ▪ CourtGuard: A Local, Multiagent Prompt Injection Classifier (Wu and Maslowski, October 2025)
183 ▪ When AI Takes the Wheel: Security Analysis of Framework-Constrained Program Generation (Liu et al., October 2025)
184 ▪ RoBCtrl: Attacking GNN-Based Social Bot Detectors via Reinforced Manipulation of Bots Control Interaction (Yang et al., October 2025)
185 ▪ Black-box Optimization of LLM Outputs by Asking for Directions (Zhang et al., October 2025)
186 ▪ Prompt injections as a tool for preserving identity in GAI image descriptions (Glazko and Mankoff, October 2025)
187 ▪ Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses (Yang et al., October 2025)
188 ▪ Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers (Zhao et al., October 2025)
189 ▪ RHINO: Guided Reasoning for Mapping Network Logs to Adversarial Tactics and Techniques with Large Language Models (Meng et al., October 2025)
190 ▪ PIShield: Detecting Prompt Injection Attacks via Intrinsic LLM Features (Zou et al., October 2025)
191 ▪ Adversarial Distilled Retrieval-Augmented Guarding Model for Online Malicious Intent Detection (Guo et al., October 2025)
192 ▪ In-Browser LLM-Guided Fuzzing for Real-Time Prompt Injection Testing in Agentic AI Browsers (Cohen, October 2025)
193 ▪ PromptLocate: Localizing Prompt Injection Attacks (Jia et al., October 2025)
194 ▪ LineBreaker: Finding Token-Inconsistency Bugs with Large Language Models (Chen et al., October 2025)
195 ▪ CoSPED: Consistent Soft Prompt Targeted Data Extraction and Defense (Zhuochen, Wai, and Vrizlynn, October 2025)
196 ▪ VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands (Liu and Tang, October 2025)
197 ▪ CommandSans: Securing AI Agents with Surgical Precision Prompt Sanitization (Das et al., October 2025)
198 ▪ AEGIS : Automated Co-Evolutionary Framework for Guarding Prompt Injections Schema (Liu et al., October 2025)
199 ▪ Indirect Prompt Injections: Are Firewalls All You Need, or Stronger Benchmarks? (Bhagwatkar et al., October 2025)
200 ▪ System Prompt Poisoning: Persistent Attacks on Large Language Models Beyond User Injection (Li, Guo, and Cai, October 2025)
201 ▪ RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection (Wen et al., October 2025)
202 ▪ Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers (KumarRavindran, October 2025)
203 ▪ VortexPIA: Indirect Prompt Injection Attack against LLMs for Efficient Extraction of User Privacy (Cui et al., October 2025)
204 ▪ AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents (Li et al., October 2025)
205 ▪ Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs (Chen et al., October 2025)
206 ▪ Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations (Teja et al., October 2025)
207 ▪ Bypassing Prompt Guards in Production with Controlled-Release Prompting (Fairoze et al., October 2025)
208 ▪ In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b (Durner, October 2025)
209 ▪ WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents (Liu et al., October 2025)
210 ▪ Fingerprinting LLMs via Prompt Injection (Hu et al., October 2025)
211 ▪ SecInfer: Preventing Prompt Injection via Inference-time Scaling (Liu et al., September 2025)
212 ▪ Prompt to Pwn: Automated Exploit Generation for Smart Contracts (Xiao et al., August 2025)
213 ▪ AgentArmor: Enforcing Program Analysis on Agent Runtime Trace to Defend Against Prompt Injection (Wang et al., August 2025)
214 ▪ LeakSealer: A Semisupervised Defense for LLMs Against Prompt Injection and Leakage Attacks (Panebianco et al., August 2025)
215 ▪ Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding (Pathade, July 2025)
216 ▪ Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities (Samancioglu, July 2025)
217 ▪ Text2VLM: Adapting Text-Only Datasets to Evaluate Alignment Training in Visual Language Models (Downer et al., July 2025)
218 ▪ MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts? (Wahed et al., July 2025)
219 ▪ Manipulating LLM Web Agents with Indirect Prompt Injection Attack via HTML Accessibility Tree (Johnson, Pham, and Le, July 2025)
220 ▪ Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design (Charles, Curry, and Charles, July 2025)
221 ▪ Can Indirect Prompt Injection Attacks Be Detected and Removed? (Chen et al., July 2025)
222 ▪ TopicAttack: An Indirect Prompt Injection Attack via Topic Transition (Chen et al., July 2025)
223 ▪ Prompt Injection 2.0: Hybrid AI Threats (McHugh, \v{S}ekrst, and Cefalu, July 2025)
224 ▪ MAD-Spear: A Conformity-Driven Prompt Injection Attack on Multi-Agent Debate Systems (Cui and Du, July 2025)
225 ▪ Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems (Hackett et al., July 2025)
226 ▪ Logic layer Prompt Control Injection (LPCI): A Novel Security Vulnerability Class in Agentic Systems (Atta et al., July 2025)
227 ▪ May I have your Attention? Breaking Fine-Tuning based Prompt Injection Defenses using Architecture-Aware Attacks (Pandya et al., July 2025)
228 ▪ How Not to Detect Prompt Injections with an LLM (Choudhary et al., July 2025)
229 ▪ FrameShift: Learning to Resize Fuzzer Inputs Without Breaking Them (Green, Goues, and Brown, July 2025)
230 ▪ Attention Tracker: Detecting Prompt Injection Attacks in LLMs (Hung et al, Apr 2025)
231 ▪ SINCon: Mitigate LLM-Generated Malicious Message Injection Attack for Rumor Detection (Zhang et al, Apr 2025)
232 ▪ RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage (Zhong et al, Feb 2025)
233 ▪ EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Liao et al, Feb 2025)
234 ▪ Self-interpreting Adversarial Images (Zhang et al, Jan 2025)
235 ▪ SecAlign: Defending Against Prompt Injection with Preference Optimization (Chen et al, Jan 2025)
236 ▪ Technical Report for ICML 2024 TiFA Workshop MLLM Attack Challenge: Suffix Injection and Projected Gradient Descent Can Easily Fool An MLLM (Guo et al, Dec 2024)
237 ▪ Defending LVLMs Against Vision Attacks through Partial-Perception Supervision (Zhou et al, Dec 2024)
238 ▪ PROSAC: Provably Safe Certification for Machine Learning Models under Adversarial Attacks (Feng et al, Dec 2024)
239 ▪ Finding a Wolf in Sheep's Clothing: Combating Adversarial Text-To-Image Prompts with Text Summarization (Cooper, Narnoli, and Surdeanu, Dec 2024)
240 ▪ PGD-Imp: Rethinking and Unleashing Potential of Classic PGD with Dual Strategies for Imperceptible Adversarial Attacks (Li et al, Dec 2024)
241 ▪ Red Teaming GPT-4V: Are GPT-4V Safe Against Uni/Multi-Modal Jailbreak Attacks? (Chen et al, Dec 2024)
242 ▪ Failures to Find Transferable Image Jailbreaks Between Vision-Language Models (Shaeffer et al, Dec 2024)
243 ▪ HTS-Attack: Heuristic Token Search for Jailbreaking Text-to-Image Models (Gao et al, Dec 2024)
244 ▪ Comprehensive Assessment of Jailbreak Attacks Against LLMs (Chu et al, Dec 2024)
245 ▪ From Allies to Adversaries: Manipulating LLM Tool-Calling through Adversarial Injection (Wang et al, Dec 2024)
246 ▪ FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks (Wang et al, Nov 2024)
247 ▪ Formalizing and Benchmarking Prompt Injection Attacks and Defenses (Liu et al, Nov 2024)
248 ▪ Universal and Context-Independent Triggers for Precise Control of LLM Outputs (Liang, Li, and Yu, Nov 2024)
249 ▪ SecONN: An Optical Neural Network Framework with Concurrent Detection of Thermal Fault Injection Attacks (Nishida et al, Nov 2024)
250 ▪ Hacking Back the AI-Hacker: Prompt Injection as a Defense Against LLM-driven Cyberattacks (Pasquini et al, Nov 2024)
251 ▪ Optimization-based Prompt Injection Attack to LLM-as-a-Judge (Shi et al, Nov 2024)
252 ▪ Goal-guided Generative Prompt Injection Attack on Large Language Models (Zhang et al, Nov 2024)
253 ▪ Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures (Benjamin et al, Oct 2024)
254 ▪ InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models (Li, Liu, and Xiao, Oct 2024)
255 ▪ FATH: Authentication-based Test-time Defense against Indirect Prompt Injection Attacks (Wang et al, Oct 2024)
256 ▪ Embedding-based classifiers can detect prompt injection attacks (Ayub and Majumdar, Oct 2024)
257 ▪ Imprompter: Tricking LLM Agents into Improper Tool Use (Fu et al, Oct 2024)
258 ▪ System-Level Defense against Indirect Prompt Injection Attacks: An Information Flow Control Perspective (Wu, Cecchetti, and Xiao, Oct 2024)
259 ▪ EIA: Environmental Injection Attack on Generalist Web Agents for Privacy Leakage (Liao et al, Sep 2024)
260 ▪ Efficient Detection of Toxic Prompts in Large Language Models (Liu et al, Sep 2024)
261 ▪ Goal-guided Generative Prompt Injection Attack on Large Language Models (Zhang et al, Sep 2024)
262 ▪ Soft Prompts Go Hard: Steering Visual Language Models with Hidden Meta-Instructions (Zhang et al, Sep 2024)
263 ▪ Optimization-based Prompt Injection Attack to LLM-as-a-Judge (Shi et al, Aug 2024)
264 ▪ Rag and Roll: An End-to-End Evaluation of Indirect Prompt Manipulations in LLM-based Application Frameworks (De Stefano, Schönherr, Pellegrino, Aug 2024)
265 ▪ InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents (Zhang et al, Aug 2024)
266 ▪ LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks (Happe, Kaplan, and Cito, Aug 2024)
267 ▪ Securing the Diagnosis of Medical Imaging: An In-depth Analysis of AI-Resistant Attacks (Biswas et al, Aug 2024)
268 ▪ On Feasibility of Intent Obfuscating Attacks (Li and Shafto, Jul 2024)
Prompt Injection and Input Manipulation (Direct and Indirect)

>

<

‍

System and Meta Prompt Extraction

Covers:

  • MITRE ATLAS Discovery and Exfiltration

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : System and Meta Prompt Extraction

cybersecurity_tracker - Google Drive

2 ▪ ProxyPrompt: Securing System Prompts against Prompt Extraction Attacks (Zhuang et al., April 2026)
3 ▪ A First Look at the Security Issues in the Model Context Protocol Ecosystem (Li and Gao, April 2026)
4 ▪ Peering Behind the Shield: Guardrail Identification in Large Language Models (Yang et al., April 2026)
5 ▪ The System Prompt Is the Attack Surface: How LLM Agent Configuration Shapes Security and Creates Exploitable Vulnerabilities (Litvak, March 2026)
6 ▪ Arbiter: Detecting Interference in LLM Agent System Prompts (Mason, March 2026)
7 ▪ OptiLeak: Efficient Prompt Reconstruction via Reinforcement Learning in Multi-tenant LLM Services (Wang et al., February 2026)
8 ▪ Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models (Zhao et al., January 2026)
9 ▪ Running in CIRCLE? A Simple Benchmark for LLM Code Interpreter Security (Chua, July 2025)
10 ▪ When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMs (Li et al., July 2025)
11 ▪ Slot: Provenance-Driven APT Detection through Graph Reinforcement Learning (Qiao et al., July 2025)
12 ▪ ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space (Chen, Dolan-Gavitt, and Lin, July 2025)
13 ▪ LoRAShield: Data-Free Editing Alignment for Secure Personalized LoRA Sharing (Chen et al., July 2025)
14 ▪ LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users (Hilel et al., July 2025)
15 ▪ Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models (Liang et al, Feb 2025)
16 ▪ Prompt Obfuscation for Large Language Models (Pape, Eisenhofer, and Schönherr, Sep 2024)
17 ▪ Prompt Leakage effect and defense strategies for multi-turn LLM interactions (Agarwal et al, Jul 2024)
System and Meta Prompt Extraction

>

<

‍

Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts

Covers:

  • MITRE ATLAS Resource Development

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts

cybersecurity_tracker - Google Drive

2 ▪ ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks? (Wang et al., May 2026)
3 ▪ Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents (Probst, Happe, and Cito, May 2026)
4 ▪ xOffense: An Autonomous Multi-Agent Framework for Penetration Testing with Domain-Adapted Large Language Models (Luong et al., April 2026)
5 ▪ Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks (Merves et al., April 2026)
6 ▪ Stand-Alone Complex or Vibercrime? Exploring the adoption and innovation of GenAI tools, coding assistants, and agents within cybercrime ecosystems (Hughes, Collier, and Thomas, April 2026)
7 ▪ MalTool: Malicious Tool Attacks on LLM Agents (Hu et al., February 2026)
8 ▪ QRS: A Rule-Synthesizing Neuro-Symbolic Triad for Autonomous Vulnerability Discovery (Tsigkourakos and Patsakis, February 2026)
9 ▪ AICrypto: Evaluating Cryptography Capabilities of Large Language Models (Wang et al., February 2026)
10 ▪ Zer0n: An AI-Assisted Vulnerability Discovery and Blockchain-Backed Integrity Framework (Parmar et al., January 2026)
11 ▪ AutoPatch: Multi-Agent Framework for Patching Real-World CVE Vulnerabilities (Seo et al., December 2025)
12 ▪ SeedAIchemy: LLM-Driven Seed Corpus Generation for Fuzzing (Wen et al., November 2025)
13 ▪ RulePilot: An LLM-Powered Agent for Security Rule Generation (Wang et al., November 2025)
14 ▪ AICrypto: A Comprehensive Benchmark for Evaluating Cryptography Capabilities of Large Language Models (Wang et al., September 2025)
15 ▪ Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search (Moss, Aug 2024)
Obtain and Develop (Software) Capabilities, Acquire Infrastructure, or Establish Accounts

>

<

‍

Jailbreak, Cost Harvesting, or Erode ML Model Integrity

Covers:

  • MITRE ATLAS Privilege Escalation, Defense Evasion, and Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Jailbreak, Cost Harvesting, or Erode ML Model Integrity

cybersecurity_tracker - Google Drive

2 ▪ EVA: Editing for Versatile Alignment against Jailbreaks (Wang et al., May 2026)
3 ▪ The Great Pretender: A Stochasticity Problem in LLM Jailbreak (Monteuuis, Chen, and Petit, May 2026)
4 ▪ Sockpuppetting: Jailbreaking LLMs by Combining Prefilling with Optimization (Dotsinski and Eustratiadis, May 2026)
5 ▪ Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack (Wang et al., May 2026)
6 ▪ MT-JailBench: A Modular Benchmark for Understanding Multi-Turn Jailbreak Attacks (Zhang et al., May 2026)
7 ▪ Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs (Yoon et al., May 2026)
8 ▪ LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments (Zhang et al., May 2026)
9 ▪ Re-Triggering Safeguards within LLMs for Jailbreak Detection (Lin et al., May 2026)
10 ▪ Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing (Lin et al., May 2026)
11 ▪ The Art of the Jailbreak: Formulating Jailbreak Attacks for LLM Security Beyond Binary Scoring (Hossain et al., May 2026)
12 ▪ Single-Configuration Attack Success Rate Is Not Enough: Jailbreak Evaluations Should Report Distributional Attack Success (Maple, Kumar, and Tapwal, May 2026)
13 ▪ Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off (Chen, Liu, and Cao, May 2026)
14 ▪ Mitigating Many-shot Jailbreak Attacks with One Single Demonstration (Chen et al., May 2026)
15 ▪ OrchJail: Jailbreaking Tool-Calling Text-to-Image Agents by Orchestration-Guided Fuzzing (Chen et al., May 2026)
16 ▪ SoK: Robustness in Large Language Models against Jailbreak Attacks (Xu et al., May 2026)
17 ▪ Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization (Fang et al., May 2026)
18 ▪ Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses (Derya and Sunar, May 2026)
19 ▪ Tracing the Dynamics of Refusal: Exploiting Latent Refusal Trajectories for Robust Jailbreak Detection (Hu et al., May 2026)
20 ▪ ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming (B\'ejar et al., May 2026)
21 ▪ SRTJ: Self-Evolving Rule-Driven Training-Free LLM Jailbreaking (Li et al., May 2026)
22 ▪ Jailbroken Frontier Models Retain Their Capabilities (Zhu et al., May 2026)
23 ▪ Dynamic Adversarial Fine-Tuning Reorganizes Refusal Geometry (Lan et al., May 2026)
24 ▪ TwinGate: Stateful Defense against Decompositional Jailbreaks in Untraceable Traffic via Asymmetric Contrastive Learning (Sun et al., May 2026)
25 ▪ One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety (Arif et al., April 2026)
26 ▪ Jailbreaking Frontier Foundation Models Through Intention Deception (Wang, Sycara, and Xie, April 2026)
27 ▪ Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards (Hammadia et al., April 2026)
28 ▪ Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem (Lin et al., April 2026)
29 ▪ Involuntary In-Context Learning: Exploiting Few-Shot Pattern Completion to Bypass Safety Alignment in GPT-5.4 (Polyakov and Kuznetsov, April 2026)
30 ▪ ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs (Liu et al., April 2026)
31 ▪ Different Paths to Harmful Compliance: Behavioral Side Effects and Mechanistic Divergence Across LLM Jailbreaks (Kabir and Tiganj, April 2026)
32 ▪ HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking (Wang et al., April 2026)
33 ▪ SafeDream: Safety World Model for Proactive Early Jailbreak Detection (Yan et al., April 2026)
34 ▪ Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization (Tang et al., April 2026)
35 ▪ TEMPLATEFUZZ: Fine-Grained Chat Template Fuzzing for Jailbreaking and Red Teaming LLMs (Shen et al., April 2026)
36 ▪ Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion (Pramanik et al., April 2026)
37 ▪ TrajGuard: Streaming Hidden-state Trajectory Detection for Decoding-time Jailbreak Defense (Liu et al., April 2026)
38 ▪ Invisible to Humans, Triggered by Agents: Stealthy Jailbreak Attacks on Mobile Vision-Language Agents (Ding et al., April 2026)
39 ▪ SelfGrader: Stable Jailbreak Detection for Large Language Models using Token-Level Logits (Zhang et al., April 2026)
40 ▪ Jailbreaking Generative AI: Multivector Phishing Threats and Transformer based Defenses (Mishra and Varshney, April 2026)
41 ▪ LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops (Fu et al., March 2026)
42 ▪ Metaphor-based Jailbreak Attacks on Text-to-Image Models (Zhang et al., March 2026)
43 ▪ Not All Tokens Are Created Equal: Query-Efficient Jailbreak Fuzzing for LLMs (Chen et al., March 2026)
44 ▪ Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models (Hong et al., March 2026)
45 ▪ Resource Consumption Threats in Large Language Models (Zhang et al., March 2026)
46 ▪ Activation Surgery: Jailbreaking White-box LLMs without Touching the Prompt (Jenny et al., March 2026)
47 ▪ Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs (Ling et al., March 2026)
48 ▪ Accelerating Suffix Jailbreak attacks with Prefix-Shared KV-cache (Wang et al., March 2026)
49 ▪ Colluding LoRA: A Composite Attack on LLM Safety Alignment (Ding, March 2026)
50 ▪ Hiding in Plain Sight: A Steganographic Approach to Stealthy LLM Jailbreaks (Geng et al., March 2026)
51 ▪ Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models (Wang, Balashankar, and Chandrasekaran, March 2026)
52 ▪ Token-Level Constraint Boundary Search for Jailbreaking Text-to-Image Models (Liu et al., March 2026)
53 ▪ Reasoning-Oriented Programming: Chaining Semantic Gadgets to Jailbreak Large Vision Language Models (Zou et al., March 2026)
54 ▪ PolyJailbreak: Cross-Modal Jailbreaking Attacks on Black-Box Multimodal LLMs (Wang et al., March 2026)
55 ▪ Two Frames Matter: A Temporal Attack for Text-to-Video Model Jailbreaking (Chen et al., March 2026)
56 ▪ SPARK: Jailbreaking T2V Models by Synergistically Prompting Auditory and Recontextualized Knowledge (Ying et al., March 2026)
57 ▪ Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Heads (Wu et al., March 2026)
58 ▪ When Memory Becomes a Vulnerability: Towards Multi-turn Jailbreak Attacks against Text-to-Image Generation Systems (Zhao et al., March 2026)
59 ▪ BitBypass: A New Direction in Jailbreaking Aligned Large Language Models with Bitstream Camouflage (Nakka and Saxena, March 2026)
60 ▪ Quantifying Frontier LLM Capabilities for Container Sandbox Escape (Marchand et al., March 2026)
61 ▪ MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs (Liu et al., March 2026)
62 ▪ Clawdrain: Exploiting Tool-Calling Chains for Stealthy Token Exhaustion in OpenClaw Agents (Dong, Feng, and Wang, March 2026)
63 ▪ Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking (Fang et al., March 2026)
64 ▪ Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search (Huang et al., February 2026)
65 ▪ Analysis of LLMs Against Prompt Injection and Jailbreak Attacks (Jaiswal et al., February 2026)
66 ▪ A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness (Luo et al., February 2026)
67 ▪ Asking Forever: Universal Activations Behind Turn Amplification in Conversational LLMs (Coalson, Fang, and Hong, February 2026)
68 ▪ Recursive language models for jailbreak detection: a procedural defense for tool-augmented agents (Shavit, February 2026)
69 ▪ Steering Dialogue Dynamics for Robustness against Multi-turn Jailbreaking Attacks (Hu, Robey, and Liu, February 2026)
70 ▪ AISA: Awakening Intrinsic Safety Awareness in Large Language Models against Jailbreak Attacks (Song, Xie, and Yin, February 2026)
71 ▪ Jailbreaking Leaves a Trace: Understanding and Detecting Jailbreak Attacks from Internal Representations of Large Language Models (Kadali and Papalexakis, February 2026)
72 ▪ from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors (Yan et al., February 2026)
73 ▪ Red-teaming the Multimodal Reasoning: Jailbreaking Vision-Language Models via Cross-modal Entanglement Attacks (Yan et al., February 2026)
74 ▪ Large Language Lobotomy: Jailbreaking Mixture-of-Experts via Expert Silencing (Lintelo, Wu, and Picek, February 2026)
75 ▪ ShallowJail: Steering Jailbreaks against Large Language Models (Liu, Pei, and Liu, February 2026)
76 ▪ TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking (Yoon et al., February 2026)
77 ▪ TrapSuffix: Proactive Defense Against Adversarial Suffixes in Jailbreaking (Du et al., February 2026)
78 ▪ A Causal Perspective for Enhancing Jailbreak Attack and Defense (Pan et al., February 2026)
79 ▪ Steering Externalities: Benign Activation Steering Unintentionally Increases Jailbreak Risk for Large Language Models (Xiong et al., February 2026)
80 ▪ When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs (Dingeto et al., February 2026)
81 ▪ How Few-shot Demonstrations Affect Prompt-based Defenses Against LLM Jailbreak Attacks (Wang et al., February 2026)
82 ▪ Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical Evidence (Fu et al., February 2026)
83 ▪ David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning (Nellessen and Kachman, February 2026)
84 ▪ Jailbreaking LLMs via Calibration (Lu, Guo, and Kong, February 2026)
85 ▪ Text is All You Need for Vision-Language Model Jailbreaking (Chen et al., February 2026)
86 ▪ ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack (Lin et al., January 2026)
87 ▪ Mask-GCG: Are All Tokens in Adversarial Suffixes Necessary for Jailbreak Attacks? (Mu et al., January 2026)
88 ▪ Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning (Yang et al., January 2026)
89 ▪ Learning to Detect Unseen Jailbreak Attacks in Large Vision-Language Models (Liang et al., January 2026)
90 ▪ LLMs Can Unlearn Refusal with Only 1,000 Benign Samples (Guo et al., January 2026)
91 ▪ LLM Jailbreak Detection for (Almost) Free! (Chen et al., January 2026)
92 ▪ GCG Attack On A Diffusion LLM (Neyroud and Corley, January 2026)
93 ▪ Eliciting Harmful Capabilities by Fine-Tuning On Safeguarded Outputs (Kaunismaa et al., January 2026)
94 ▪ TrojanPraise: Jailbreak LLMs via Benign Fine-Tuning (Xie, Song, and Luo, January 2026)
95 ▪ AJAR: Adaptive Jailbreak Architecture for Red-teaming (Dou and Yang, January 2026)
96 ▪ SpatialJB: How Text Distribution Art Becomes the "Jailbreak Key" for LLM Guardrails (Mou et al., January 2026)
97 ▪ From static to adaptive: immune memory-based jailbreak detection for large language models (Leng et al., January 2026)
98 ▪ MacPrompt: Maraconic-guided Jailbreak against Text-to-Image Models (Ye et al., January 2026)
99 ▪ PromptScreen: Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline (Rao et al., January 2026)
100 ▪ The Echo Chamber Multi-Turn LLM Jailbreak (Alobaid et al., January 2026)
101 ▪ Jailbreaking Large Language Models through Iterative Tool-Disguised Attacks via Reinforcement Learning (Wang et al., January 2026)
102 ▪ Knowledge-Driven Multi-Turn Jailbreaking on Large Language Models (Li et al., January 2026)
103 ▪ Multi-turn Jailbreaking Attack in Multi-Modal Large Language Models (Das et al., January 2026)
104 ▪ Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs (Xing et al., January 2026)
105 ▪ Jailbreaking Safeguarded Text-to-Image Models via Large Language Models (Jiang et al., January 2026)
106 ▪ $PC^2$: Politically Controversial Content Generation via Jailbreaking Attacks on GPT-based Text-to-Image Models (Choi et al., January 2026)
107 ▪ Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks (Cunningham et al., January 2026)
108 ▪ Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks (Nurlanov, Schmidt, and Bernard, January 2026)
109 ▪ Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models (Hu et al., January 2026)
110 ▪ Jailbreaking LLMs & VLMs: Mechanisms, Evaluation, and Unified Defense (Chen et al., January 2026)
111 ▪ TRYLOCK: Defense-in-Depth Against LLM Jailbreaks via Layered Preference and Representation Engineering (Thornton, January 2026)
112 ▪ How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference (Liu et al., January 2026)
113 ▪ JPU: Bridging Jailbreak Defense and Unlearning via On-Policy Path Rectification (Wang et al., January 2026)
114 ▪ Beyond Prompts: Space-Time Decoupling Control-Plane Jailbreaks in LLM Structured Output (Zhang et al., January 2026)
115 ▪ Emoji-Based Jailbreaking of Large Language Models (Gopinadh and Hussain, January 2026)
116 ▪ Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak (Gu et al., January 2026)
117 ▪ RAJ-PGA: Reasoning-Activated Jailbreak and Principle-Guided Alignment Framework for Large Reasoning Models (Chen et al., January 2026)
118 ▪ Effective and Efficient Jailbreaks of Black-Box LLMs with Cross-Behavior Attacks (Gohil, January 2026)
119 ▪ Language Model Agents Under Attack: A Cross Model-Benchmark of Profit-Seeking Behaviors in Customer Service (Zhang, January 2026)
120 ▪ Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race? (Xin et al., January 2026)
121 ▪ Involuntary Jailbreak: On Self-Prompting Attacks (Guo, Li, and Kankanhalli, December 2025)
122 ▪ EquaCode: A Multi-Strategy Jailbreak Approach for Large Language Models via Equation Solving and Code Completion (Liang, Huang, and Chen, December 2025)
123 ▪ X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability (Lu et al., December 2025)
124 ▪ Evolving Security in LLMs: A Study of Jailbreak Attacks and Defenses (Shang, Wei, and Bai, December 2025)
125 ▪ Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking (Huang et al., December 2025)
126 ▪ GateBreaker: Gate-Guided Attacks on Mixture-of-Expert LLMs (Wu et al., December 2025)
127 ▪ Odysseus: Jailbreaking Commercial Multimodal LLM-integrated Systems via Dual Steganography (Li et al., December 2025)
128 ▪ Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs (Li et al., December 2025)
129 ▪ Universal Jailbreak Suffixes Are Strong Attention Hijackers (Ben-Tov, Geva, and Sharif, December 2025)
130 ▪ Bleeding Pathways: Vanishing Discriminability in LLM Hidden States Fuels Jailbreak Attacks (Zhang et al., December 2025)
131 ▪ Efficient Jailbreak Mitigation Using Semantic Linear Classification in a Multi-Staged Pipeline (Rao et al., December 2025)
132 ▪ Breaking Minds, Breaking Systems: Jailbreaking Large Language Models via Human-like Psychological Manipulation (Liu and Lin, December 2025)
133 ▪ One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs (Tan, Yu, and Sakuma, December 2025)
134 ▪ Safe2Harm: Semantic Isomorphism Attacks for Jailbreaking Large Language Models (Yang, December 2025)
135 ▪ Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring (Hua et al., December 2025)
136 ▪ Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously (Adiletta et al., December 2025)
137 ▪ Metaphor-based Jailbreaking Attacks on Text-to-Image Models (Zhang et al., December 2025)
138 ▪ A Practical Framework for Evaluating Medical AI Security: Reproducible Assessment of Jailbreaking and Privacy Vulnerabilities Across Clinical Specialties (Wang, Zhang, and Yagemann, December 2025)
139 ▪ OmniSafeBench-MM: A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack-Defense Evaluation (Jia et al., December 2025)
140 ▪ Beyond Model Jailbreak: Systematic Dissection of the "Ten DeadlySins" in Embodied Intelligence (Huang et al., December 2025)
141 ▪ TeleAI-Safety: A comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations (Chen et al., December 2025)
142 ▪ The Trojan Knowledge: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search (Wei et al., December 2025)
143 ▪ SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism (Chen et al., December 2025)
144 ▪ Les Dissonances: Cross-Tool Harvesting and Polluting in Pool-of-Tools Empowered LLM Agents (Li et al., December 2025)
145 ▪ Immunity memory-based jailbreak detection: multi-agent adaptive guard for large language models (Leng, Zhang, and Zhang, December 2025)
146 ▪ Lockpicking LLMs: A Logit-Based Jailbreak Using Token-level Manipulation (Li et al., December 2025)
147 ▪ Involuntary Jailbreak (Guo, Li, and Kankanhalli, December 2025)
148 ▪ A Wolf in Sheep's Clothing: Bypassing Commercial LLM Guardrails via Harmless Prompt Weaving and Adaptive Tree Search (Wei et al., December 2025)
149 ▪ DefenSee: Dissecting Threat from Sight and Text - A Multi-View Defensive Pipeline for Multi-modal Jailbreaks (Wang, Fok, and Thing, December 2025)
150 ▪ Jailbreaking and Mitigation of Vulnerabilities in Large Language Models (Peng et al., November 2025)
151 ▪ Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations (Wong et al., November 2025)
152 ▪ TASO: Jailbreak LLMs via Alternative Template and Suffix Optimization (Wang et al., November 2025)
153 ▪ Beyond Jailbreak: Unveiling Risks in LLM Applications Arising from Blurred Capability Boundaries (Zhang et al., November 2025)
154 ▪ Reason2Attack: Jailbreaking Text-to-Image Models via LLM Reasoning (Zhang et al., November 2025)
155 ▪ Practical and Stealthy Touch-Guided Jailbreak Attacks on Deployed Mobile Vision-Language Agents (Ding et al., November 2025)
156 ▪ LightDefense: A Lightweight Uncertainty-Driven Defense against Jailbreaks via Shifted Token Distribution (Yang and Zhang, November 2025)
157 ▪ The Shawshank Redemption of Embodied AI: Understanding and Benchmarking Indirect Environmental Jailbreaks (Li et al., November 2025)
158 ▪ "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios (Sun et al., November 2025)
159 ▪ Scaling Patterns in Adversarial Alignment: Evidence from Multi-LLM Jailbreak Experiments (Nathanson, Williams, and Matuszek, November 2025)
160 ▪ Beyond Fixed and Dynamic Prompts: Embedded Jailbreak Templates for Advancing LLM Security (Kim, Na, and Choi, November 2025)
161 ▪ VEIL: Jailbreaking Text-to-Video Models via Visual Exploitation from Implicit Language (Ying et al., November 2025)
162 ▪ Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs (Chen et al., November 2025)
163 ▪ ForgeDAN: An Evolutionary Framework for Jailbreaking Aligned Large Language Models (Cheng et al., November 2025)
164 ▪ NegBLEURT Forest: Leveraging Inconsistencies for Detecting Jailbreak Attacks (Sleem et al., November 2025)
165 ▪ Can AI Models be Jailbroken to Phish Elderly Victims? An End-to-End Evaluation (Heiding and Lermen, November 2025)
166 ▪ GPT-5 at CTFs: Case Studies From Top-Tier Cybersecurity Events (Reworr, Petrov, and Volkov, November 2025)
167 ▪ Chain-of-Lure: A Universal Jailbreak Attack Framework using Unconstrained Synthetic Narratives (Chang et al., November 2025)
168 ▪ Graph of Attacks with Pruning: Optimizing Stealthy Jailbreak Prompt Generation for Enhanced LLM Content Moderation (Schwartz et al., November 2025)
169 ▪ UDora: A Unified Red Teaming Framework against LLM Agents by Dynamically Hijacking Their Own Reasoning (Zhang, Yang, and Li, November 2025)
170 ▪ Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs (Zhou et al., November 2025)
171 ▪ KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs (Liu et al., November 2025)
172 ▪ HumorReject: Decoupling LLM Safety from Refusal Prefix via A Little Humor (Wu et al., November 2025)
173 ▪ JailbreakZoo: Survey, Landscapes, and Horizons in Jailbreaking Large Language and Vision-Language Models (Jin et al., November 2025)
174 ▪ JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework (Zhou et al., November 2025)
175 ▪ Jailbreaking in the Haystack (Shah et al., November 2025)
176 ▪ GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs (Basani and Zhang, November 2025)
177 ▪ VERA: Variational Inference Framework for Jailbreaking Large Language Models (Lochab et al., November 2025)
178 ▪ Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models (Yang and Fu, November 2025)
179 ▪ Let the Bees Find the Weak Spots: A Path Planning Perspective on Multi-Turn Jailbreak Attacks against LLMs (Liu, Hou, and Sui, November 2025)
180 ▪ AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models (Reddy, Zagula, and Saban, November 2025)
181 ▪ An Automated Framework for Strategy Discovery, Retrieval, and Evolution in LLM Jailbreak Attacks (Liu et al., November 2025)
182 ▪ Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models (Yang et al., November 2025)
183 ▪ What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks (Kirch et al., November 2025)
184 ▪ Exploiting Latent Space Discontinuities for Building Universal LLM Jailbreaks and Data Extraction Attacks (Paim et al., November 2025)
185 ▪ Sentra-Guard: A Multilingual Human-AI Framework for Real-Time Defense Against Adversarial LLM Jailbreaks (Hasan et al., October 2025)
186 ▪ Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models (Ntais, October 2025)
187 ▪ Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses (Zhong, Fok, and Thing, October 2025)
188 ▪ The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning (Liu et al., October 2025)
189 ▪ Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency (Jiang et al., October 2025)
190 ▪ Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training (Yong and Bach, October 2025)
191 ▪ Beyond Text: Multimodal Jailbreaking of Vision-Language and Audio Models through Perceptually Simple Transformations (Kumar et al., October 2025)
192 ▪ Learning to Detect Unknown Jailbreak Attacks in Large Vision-Language Models (Liang et al., October 2025)
193 ▪ HarmNet: A Framework for Adaptive Multi-Turn Jailbreak Attacks on Large Language Models (Narula et al., October 2025)
194 ▪ BreakFun: Jailbreaking LLMs via Schema Exploitation (Oskooei and Aktas, October 2025)
195 ▪ VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models (Liao, Lochab, and Zhang, October 2025)
196 ▪ Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks (Wang et al., October 2025)
197 ▪ PRISON: Unmasking the Criminal Potential of Large Language Models (Wu et al., October 2025)
198 ▪ Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling (Zhang et al., October 2025)
199 ▪ Active Honeypot Guardrail System: Probing and Confirming Multi-Turn LLM Jailbreaks (Wu, Wang, and Liao, October 2025)
200 ▪ SoK: Evaluating Jailbreak Guardrails for Large Language Models (Wang et al., October 2025)
201 ▪ Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts (Zhang et al., October 2025)
202 ▪ Bag of Tricks for Subverting Reasoning-based Safety Guardrails (Chen et al., October 2025)
203 ▪ ArtPerception: ASCII Art-based Jailbreak on LLMs with Recognition Pre-test (Yang et al., October 2025)
204 ▪ MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation (Zhu et al., October 2025)
205 ▪ The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections (Nasr et al., October 2025)
206 ▪ Pattern Enhanced Multi-Turn Jailbreaking: Exploiting Structural Vulnerabilities in Large Language Models (Nihal et al., October 2025)
207 ▪ PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization (Liu et al., October 2025)
208 ▪ MetaDefense: Defending Finetuning-based Jailbreak Attack Before and During Generation (Jiang and Pan, October 2025)
209 ▪ Effective and Stealthy One-Shot Jailbreaks on Deployed Mobile Vision-Language Agents (Ding et al., October 2025)
210 ▪ Jailbreak Attack Initializations as Extractors of Compliance Directions (Levi et al., October 2025)
211 ▪ AutoDAN-Reasoning: Enhancing Strategies Exploration based Jailbreak Attacks with Test-Time Scaling (Liu and Xiao, October 2025)
212 ▪ Auditing Pay-Per-Token in Large Language Models (Velasco, Tsirtsis, and Gomez-Rodriguez, October 2025)
213 ▪ DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization (Huang et al., October 2025)
214 ▪ Imperceptible Jailbreaking against Large Language Models (Gao et al., October 2025)
215 ▪ NEXUS: Network Exploration for eXploiting Unsafe Sequences in Multi-Turn LLM Jailbreaks (Asl et al., October 2025)
216 ▪ JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models (Peng et al., October 2025)
217 ▪ XBreaking: Explainable Artificial Intelligence for Jailbreaking LLMs (Arazzi et al., October 2025)
218 ▪ PrisonBreak: Jailbreaking Large Language Models with at Most Twenty-Five Targeted Bit-flips (Coalson et al., October 2025)
219 ▪ Untargeted Jailbreak Attack (Huang et al., October 2025)
220 ▪ Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs (Xie, Song, and Luo, October 2025)
221 ▪ Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks (Saha et al., October 2025)
222 ▪ Fine-Tuning Jailbreaks under Highly Constrained Black-Box Settings: A Three-Pronged Approach (Li, Wang, and Li, October 2025)
223 ▪ Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge (Dou et al., October 2025)
224 ▪ AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning Models (Liang et al., October 2025)
225 ▪ STAC: When Innocent Tools Form Dangerous Chains to Jailbreak LLM Agents (Li et al., October 2025)
226 ▪ Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents (Das, Sandler, and Fioretto, September 2025)
227 ▪ Formalization Driven LLM Prompt Jailbreaking via Reinforcement Learning (Wang et al., September 2025)
228 ▪ Takedown: How It's Done in Modern Coding Agent Exploits (Lee et al., September 2025)
229 ▪ Bidirectional Intention Inference Enhances LLMs' Defense Against Multi-Turn Jailbreak Attacks (Tong et al., September 2025)
230 ▪ Activation-Guided Local Editing for Jailbreaking Attacks (Wang et al., August 2025)
231 ▪ Probabilistic Modeling of Jailbreak on Multimodal LLMs: From Quantification to Application (Xu et al., August 2025)
232 ▪ Enhancing Jailbreak Attacks on LLMs via Persona Prompts (Zhang et al., July 2025)
233 ▪ PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking (Zou et al., July 2025)
234 ▪ NCCR: to Evaluate the Robustness of Neural Networks and Adversarial Examples (Shi, July 2025)
235 ▪ How Much Do Large Language Model Cheat on Evaluation? Benchmarking Overestimation under the One-Time-Pad-Based Framework (Liang et al., July 2025)
236 ▪ From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem (Mao et al., July 2025)
237 ▪ Exploiting Jailbreaking Vulnerabilities in Generative AI to Bypass Ethical Safeguards for Facilitating Phishing Attacks (Mishra and Varshney, July 2025)
238 ▪ Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility (Murphy et al., July 2025)
239 ▪ Representation Bending for Large Language Model Safety (Yousefpour et al., July 2025)
240 ▪ GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing (Zhang et al., July 2025)
241 ▪ Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning (Liu et al., July 2025)
242 ▪ GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods (Huang et al., July 2025)
243 ▪ On Jailbreaking Quantized Language Models Through Fault Injection Attacks (Zahran et al., July 2025)
244 ▪ Feint and Attack: Attention-Based Strategies for Jailbreaking and Protecting LLMs (Pu et al., July 2025)
245 ▪ CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations (Li et al., July 2025)
246 ▪ Three Minds, One Legend: Jailbreak Large Reasoning Model with Adaptive Stacked Ciphers<br><br> (Nguyen et al., May 2025)
247 ▪ Prompt Inference Attack on Distributed Large Language Model Inference Frameworks (Luo, Yu, and Xiao, May 2025)
248 ▪ JULI: Jailbreak Large Language Models by Self-Introspection (Wang, Hu, and Wagnetr, May 2025)
249 ▪ AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models (Chen et al, May 2025)
250 ▪ Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search (Arel and Zhou, May 2025)
251 ▪ Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate (Qi et al, Apr 2025)
252 ▪ Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models (Ke et al, Apr 2025)
253 ▪ h4rm3l: A language for Composable Jailbreak Attack Synthesis (Doumbouya et al, Mar 2025)
254 ▪ Siege: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search (Zhou, Mar 2025)
255 ▪ Adversarial Training for Multimodal Large Language Models against Jailbreak Attacks (Lu et al, Mar 2025)
256 ▪ Dagger Behind Smile: Fool LLMs with a Happy Ending Story (Song et al, Feb 2025)
257 ▪ Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks (Fang et al, Feb 2025)
258 ▪ An Engorgio Prompt Makes Large Language Model Babble on (Dong et al, Feb 2025)
259 ▪ Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models (Ying et al, Mar 2025)
260 ▪ LLMs can be Dangerous Reasoners: Analyzing-based Jailbreak Attack on Large Language Models (Lin et al, Feb 2025)
261 ▪ Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation (Armstrong et al, Jan 2025)
262 ▪ Data-Free Model-Related Attacks: Unleashing the Potential of Generative AI (Ye et al, Jan 2025)
263 ▪ GreedyPixel: Fine-Grained Black-Box Adversarial Attack Via Greedy Algorithm (Wang et al, Jan 2025)
264 ▪ MOS-Attack: A Scalable Multi-objective Adversarial Attack Framework (Guo et al, Jan 2025)
265 ▪ Siren: A Learning-Based Multi-Turn Attack Framework for Simulating Real-World Human Jailbreak Behaviors (Zhao and Zhang, Jan 2025)
266 ▪ Gandalf the Red: Adaptive Security for LLMs (Pfister et al, Jan 2025)
267 ▪ DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models (Yin, Liu, and Xu, Jan 2025)
268 ▪ Infecting Generative AI With Viruses (Noever and McKee, Jan 2025)
269 ▪ BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger (Chen et al, Jan 2025)
270 ▪ Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment (Ghosal et al, Dec 2024)
271 ▪ Automated Progressive Red Teaming (Jiang et al, Dec 2024)
272 ▪ Crabs: Consuming Resrouce via Auto-generation for LLM-DoS Attack under Black-box Settings (Zhang et al, Dec 2024)
273 ▪ LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses (Lin et al, Dec 2024)
274 ▪ When LLM Meets DRL: Advancing Jailbreaking Efficiency via DRL-guided Search (Chen et al, Dec 2024)
275 ▪ Data to Defense: The Role of Curation in Customizing LLMs Against Jailbreaking Attacks (Liu et al, Dec 2024)
276 ▪ DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints (Zhao et al, Dec 2024)
277 ▪ SATA: A Paradigm for LLM Jailbreak via Simple Assistive Task Linkage (Dong et al, Dec 2024)
278 ▪ JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs (Li et al, Dec 2024)
279 ▪ Watertox: The Art of Simplicity in Universal Attacks A Cross-Model Framework for Robust Adversarial Generation (Gao et al, Dec 2024)
280 ▪ CAMH: Advancing Model Hijacking Attack in Machine Learning (He et al, Dec 2024)
281 ▪ Best-of-N Jailbreaking (Hughes et al, Dec 2024)
282 ▪ AICAttack: Adversarial Image Captioning Attack with Attention-Based Optimization (Li et al, Dec 2024)
283 ▪ AdvPrefix: An Objective for Nuanced LLM Jailbreaks (Zhu et al, Dec 2024)
284 ▪ Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach (Wang et al, Dec 2024)
285 ▪ Stealthy Multi-Task Adversarial Attacks (Guo et al, Nov 2024)
286 ▪ AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs (Liu et al, Nov 2024)
287 ▪ BlackDAN: A Black-Box Multi-Objective Approach for Effective and Contextual Jailbreaking of Large Language Models (Wang et al, Nov 2024)
288 ▪ JailBreakV: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks (Luo et al, Nov 2024)
289 ▪ SQL Injection Jailbreak: a structural disaster of large language models (Zhao et al, Nov 2024)
290 ▪ LLMStinger: Jailbreaking LLMs using RL fine-tuned LLMs (Jha, Arora, and Ganesh, Nov 2024)
291 ▪ AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks (Zeng et al, Nov 2024)
292 ▪ DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers (Li et al, Nov 2024)
293 ▪ SequentialBreak: Large Language Models Can be Fooled by Embedding Jailbreak Prompts into Sequential Prompt Chains (Saiem et al, Nov 2024)
294 ▪ Transferable Ensemble Black-box Jailbreak Attacks on Large Language Models (Yang and Fu, Oct 2024)
295 ▪ Pruning for Protection: Increasing Jailbreak Resistance in Aligned LLMs Without Fine-Tuning (Hasan, Rugina, and Wang, Oct 2024)
296 ▪ Tree of Attacks: Jailbreaking Black-Box LLMs Automatically (Mehrotra et al, Oct 2024)
297 ▪ Fight Back Against Jailbreaking via Prompt Adversarial Tuning (Mo et al, Oct 2024)
298 ▪ HijackRAG: Hijacking Attacks against Retrieval-Augmented Large Language Models (Zhang et al, Oct 2024)
299 ▪ Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector (Huang et al, Oct 2024)
300 ▪ Improved Few-Shot Jailbreaking Can Circumvent Aligned Language Models and Their Defenses (Zheng et al, Oct 2024)
301 ▪ Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents (Yang et al, Oct 2024)
302 ▪ Transferable Adversarial Attacks on SAM and Its Downstream Models (Xia et al, Oct 2024)
303 ▪ Remote Timing Attacks on Efficient Language Model Inference (Carlini and Nasr, Oct 2024)
304 ▪ MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defenses for Multimodal Large Language Models (Weng et al, Oct 2024)
305 ▪ RePD: Defending Jailbreak Attack through a Retrieval-based Prompt Decomposition Process (Wang, Liu, and Xiao, Oct 2024)
306 ▪ Refusal-Trained LLMs Are Easily Jailbroken As Browser Agents (Kumar et al, Oct 2024)
307 ▪ Universal Black-Box Reward Poisoning Attack against Offline Reinforcement Learning (Xu, Gumaste, and Singh, Oct 2024)
308 ▪ Adversarial Attacks on Large Language Models Using Regularized Relaxation (Chacko et al, Oct 2024)
309 ▪ Faster-GCG: Efficient Discrete Optimization Jailbreak Attacks against Aligned Large Language Models (Li et al, Oct 2024)
310 ▪ Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables (Chen et al, Oct 2024)
311 ▪ Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks (Peng et al, Oct 2024)
312 ▪ Effective and Evasive Fuzz Testing-Driven Jailbreaking Attacks against LLMs (Gong et al, Oct 2024)
313 ▪ Functional Homotopy: Smoothing Discrete Optimization via Continuous Parameters for LLM Jailbreak Attacks (Wang et al, Oct 2024)
314 ▪ Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models (Shen et al, Oct 2024)
315 ▪ PathSeeker: Exploring LLM Security Vulnerabilities with a Reinforcement Learning-Based Jailbreak Approach (Lin et al, Oct 2024)
316 ▪ Rethinking and Defending Protective Perturbation in Personalized Diffusion Models (Liu et al, Oct 2024)
317 ▪ Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models (Yu et al, Sep 2024)
318 ▪ Read Over the Lines: Attacking LLMs and Toxicity Detection Systems with ASCII Art to Mask Profanity (Berezin, Farahbakhsh, Crespi, Oct 2024)
319 ▪ Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI (Rawat et al, Sep 2024)
320 ▪ Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn Interaction (Zhang et al, Sep 2024)
321 ▪ Adversarial Attacks on Parts of Speech: An Empirical Study in Text-to-Image Generation (Shahariar et al, Sep 2024)
322 ▪ Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack (Russinovich, Salem, and Eldan, Sep 2024)
323 ▪ VulZoo: A Comprehensive Vulnerability Intelligence Dataset (Ruan et al, Sep 2024)
324 ▪ Adversarial Attacks on Machine Learning-Aided Visualizations (Fujiwara et al, Sep 2024)
325 ▪ Jailbreaking Large Language Models with Symbolic Mathematics (Bethany et al, Sep 2024)
326 ▪ Image Hijacks: Adversarial Images can Control Generative Models at Runtime (Bailey et al, Sep 2024)
327 ▪ LLM Whisperer: An Inconspicuous Attack to Bias LLM Responses (Lin et al, Sep 2024)
328 ▪ Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents (Shafran, Schuster, and Shmatikov, Sep 2024)
329 ▪ Security Attacks on LLM-based Code Completion Tools (Cheng et al, Sep 2024)
330 ▪ Adversarial Attacks to Multi-Modal Models (Dou et al, Sep 2024)
331 ▪ HSF: Defending against Jailbreak Attacks with Hidden State Filtering (Qian et al, Sep 2024)
332 ▪ Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language Models (Kalavasis et al, Sep 2024)
333 ▪ Recent Advances in Attack and Defense Approaches of Large Language Models (Cui et al, Sep 2024)
334 ▪ SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner (Wang et al, Sep 2024)
335 ▪ LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet (Li et al, Sep 2024)
336 ▪ Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models (Ma et al, Sep 2024)
337 ▪ Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models (An et al, Sep 2024)
338 ▪ The Dark Side of Function Calling: Pathways to Jailbreaking Large Language Models (Wu et al, Aug 2024)
339 ▪ Detecting AI Flaws: Target-Driven Attacks on Internal Faults in Language Models (Du et al, Aug 2024)
340 ▪ LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet (Li et al, Aug 2024)
341 ▪ Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything (Zou et al, Aug 2024)
342 ▪ A StrongREJECT for Empty Jailbreaks (Souly et al, Aug 2024)
343 ▪ RT-Attack: Jailbreaking Text-to-Image Models via Random Token (Gao et al, Aug 2024)
344 ▪ CAMH: Advancing Model Hijacking Attack in Machine Learning (He et al, Aug 2024)
345 ▪ RT-Attack: Jailbreaking Text-to-Image Models via Random Token (Gao et al, Aug 2024)
346 ▪ BaThe: Defense against the Jailbreak Attack in Multimodal Large Language Models by Treating Harmful Instruction as Backdoor Trigger (Chen et al, Aug 2024)
347 ▪ Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks (Zhao et al, Aug 2024)
348 ▪ A Survey of Trojan Attacks and Defenses to Deep Neural Networks (Jin et al, Aug 2024)
349 ▪ MMJ-Bench: A Comprehensive Study on Jailbreak Attacks and Defensesfor Vision Language Models (Weng et al, Aug 2024)
350 ▪ Trojan Activation Attack: Red-Teaming Large Language Models using Activation Steering for Safety-Alignment (Wang and Shu, Aug 2023)
351 ▪ Resilience in Online Federated Learning: Mitigating Model-Poisoning Attacks via Partial Sharing (Lari et al, Aug 2024)
352 ▪ EnJa: Ensemble Jailbreak on Large Language Models (Zhang et al, Aug 2024)
353 ▪ Can Reinforcement Learning Unlock the Hidden Dangers in Aligned Large Language Models? (Bahrami, Vishwamitra, and Najafirad, Aug 2024)
354 ▪ Jailbreaking Text-to-Image Models with LLM-Based Agents (Dong et al, Aug 2024)
355 ▪ Can LLMs be Fooled? Investigating Vulnerabilities in LLMs (Abdali et al, Jul 2024)
356 ▪ Figure it Out: Analyzing-based Jailbreak Attack on Large Language Models (Lu et al, Jul 2024)
357 ▪ Vera Verto: Multimodal Hijacking Attack (Zhang et al, Jul 2024)
Jailbreak, Cost Harvesting, or Erode ML Model Integrity

>

<

‍

Proxy AI ML Model (Simulations)

Covers:

  • MITRE ATLAS ML Attack Staging

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Proxy AI ML Model (Simulations)

cybersecurity_tracker - Google Drive

2 ▪ PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior (Flemings and Annavaram, May 2026)
3 ▪ Searching for Privacy Risks in LLM Agents via Simulation (Zhang and Yang, May 2026)
4 ▪ HBEE: Human Behavioral Entropy Engine -- Pre-Registered Multi-Agent LLM Simulation of Peer-Suspicion-Based Detection Inversion (Ferrel, May 2026)
5 ▪ Threat-Oriented Digital Twinning for Security Evaluation of Autonomous Platforms (Neubert, Kandel, and Pek\"oz, April 2026)
6 ▪ Text-Based Personas for Simulating User Privacy Decisions (Fawaz et al., March 2026)
7 ▪ Evaluating Generalization Mechanisms in Autonomous Cyber Attack Agents (Luk\'a\v{s} et al., March 2026)
8 ▪ How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors? (Li et al., February 2026)
9 ▪ Towards Production-Worthy Simulation for Autonomous Cyber Operations (Tholl et al., February 2026)
10 ▪ CyberExplorer: Benchmarking LLM Offensive Security Capabilities in a Real-World Attacking Simulation Environment (Rani et al., February 2026)
11 ▪ Identifying Adversary Tactics and Techniques in Malware Binaries with an LLM Agent (Xuan et al., February 2026)
12 ▪ Bypassing AI Control Protocols via Agent-as-a-Proxy Attacks (Isbarov and Kantarcioglu, February 2026)
13 ▪ VirtualCrime: Evaluating Criminal Potential of Large Language Models via Sandbox Simulation (Tang et al., January 2026)
14 ▪ MCP Bridge: A Lightweight, LLM-Agnostic RESTful Proxy for Model Context Protocol Servers (Ahmadi, Sharif, and Banad, January 2026)
15 ▪ The Imitation Game: Using Large Language Models as Chatbots to Combat Chat-Based Cybercrimes (Yao et al., December 2025)
16 ▪ Chimera: Harnessing Multi-Agent LLMs for Automatic Insider Threat Simulation (Yu et al., December 2025)
17 ▪ ScamAgents: How AI Agents Can Simulate Human-Level Scam Calls (Badhe, December 2025)
18 ▪ HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities (Ren et al., October 2025)
19 ▪ Neutral Agent-based Adversarial Policy Learning against Deep Reinforcement Learning in Multi-party Open Systems (Peng et al., October 2025)
20 ▪ SCART: Simulation of Cyber Attacks for Real-Time (Rahimi et al., October 2025)
21 ▪ LegalSim: Multi-Agent Simulation of Legal Systems for Discovering Procedural Exploits (Badhe, October 2025)
22 ▪ Secret Collusion among AI Agents: Multi-Agent Deception via Steganography (Motwani et al., July 2025)
Proxy AI ML Model (Simulations)

>

<

‍

Verify Attack (Efficacy)

Covers:

  • MITRE ATLAS ML Attack Staging

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Verify Attack (Efficacy)

cybersecurity_tracker - Google Drive

2 ▪ Quantifying LLM Safety Degradation Under Repeated Attacks Using Survival Analysis (Topol, May 2026)
3 ▪ CTFusion: A CTF-based Benchmark for LLM Agent Evaluation (Lee, Bae, and Yun, May 2026)
4 ▪ Behavioral Integrity Verification for AI Agent Skills (Wu, Li, and Liu, May 2026)
5 ▪ SOCpilot: Verifying Policy Compliance for LLM-Assisted Incident Response (Barbieri et al., May 2026)
6 ▪ Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture the Flag Challenges (Al-Kaswan et al., May 2026)
7 ▪ Latent Adversarial Detection: Adaptive Probing of LLM Activations for Multi-Turn Attack Detection (Kulkarni, May 2026)
8 ▪ Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode (Ji et al., April 2026)
9 ▪ The Surprising Universality of LLM Outputs: A Real-Time Verification Primitive (Bogdan and Valois-Franklin, April 2026)
10 ▪ MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors (Li et al., April 2026)
11 ▪ Do Agents Dream of Root Shells? Partial-Credit Evaluation of LLM Agents in Capture The Flag Challenges (Al-Kaswan et al., April 2026)
12 ▪ Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery (Agarwal, April 2026)
13 ▪ RealVuln: Benchmarking Rule-Based, General-Purpose LLM, and Security-Specialized Scanners on Real-World Code (Pellew and Raza, April 2026)
14 ▪ Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis (Gajjar, April 2026)
15 ▪ PoC-Adapt: Semantic-Aware Automated Vulnerability Reproduction with LLM Multi-Agents and Reinforcement Learning-Driven Adaptive Policy (Duy et al., April 2026)
16 ▪ Guiding Symbolic Execution with Static Analysis and LLMs for Vulnerability Discovery (Shafiuzzaman et al., April 2026)
17 ▪ Dynamic Free-Rider Detection in Federated Learning via Simulated Attack Patterns (Nakamura, April 2026)
18 ▪ OrgForge-IT: A Verifiable Synthetic Benchmark for LLM-Based Insider Threat Detection (Flynt, March 2026)
19 ▪ SkillProbe: Security Auditing for Emerging Agent Skill Marketplaces via Multi-Agent Collaboration (Guo et al., March 2026)
20 ▪ Towards Verifiable AI with Lightweight Cryptographic Proofs of Inference (Anchuri et al., March 2026)
21 ▪ When Scanners Lie: Evaluator Instability in LLM Red-Teaming (Erez et al., March 2026)
22 ▪ Security and Quality in LLM-Generated Code: A Multi-Language, Multi-Model Analysis (Kharma et al., March 2026)
23 ▪ Before You Hand Over the Wheel: Evaluating LLMs for Security Incident Analysis (Jajodia et al., March 2026)
24 ▪ Beyond Input Guardrails: Reconstructing Cross-Agent Semantic Flows for Execution-Aware Attack Detection (Wei et al., March 2026)
25 ▪ AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding (Ma et al., March 2026)
26 ▪ ZeroDayBench: Evaluating LLM Agents on Unseen Zero-Day Vulnerabilities for Cyberdefense (Lau et al., March 2026)
27 ▪ Can LLMs Hack Enterprise Networks? -- Replicated Computational Results (RCR) Report (Happe and Cito, March 2026)
28 ▪ DualSentinel: A Lightweight Framework for Detecting Targeted Attacks in Black-box LLM via Dual Entropy Lull Pattern (Pang et al., March 2026)
29 ▪ vEcho: A Paradigm Shift from Vulnerability Verification to Proactive Discovery with Large Language Models (Jiang et al., March 2026)
30 ▪ IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computation (Guo et al., February 2026)
31 ▪ Evaluating the Reliability of Digital Forensic Evidence Discovered by Large Language Model: A Case Study (Khatiwala, Addai, and Xu, February 2026)
32 ▪ Red-Teaming Claude Opus and ChatGPT-based Security Advisors for Trusted Execution Environments (Mukherjee, February 2026)
33 ▪ Mind the Gap: Evaluating LLMs for High-Level Malicious Package Detection vs. Fine-Grained Indicator Identification (Ryan et al., February 2026)
34 ▪ Peak + Accumulation: A Proxy-Level Scoring Formula for Multi-Turn LLM Attack Detection (Corll, February 2026)
35 ▪ Learning-Based Automated Adversarial Red-Teaming for Robustness Evaluation of Large Language Models (Wei et al., February 2026)
36 ▪ VideoSTF: Stress-Testing Output Repetition in Video Large Language Models (Cao et al., February 2026)
37 ▪ Benchmarking Knowledge-Extraction Attack and Defense on Retrieval-Augmented Generation (Qi et al., February 2026)
38 ▪ Capability-Based Scaling Trends for LLM-Based Red-Teaming (Panfilov et al., February 2026)
39 ▪ Towards Real-World Industrial-Scale Verification: LLM-Driven Theorem Proving on seL4 (Zhang et al., February 2026)
40 ▪ Next-generation cyberattack detection with large language models: anomaly analysis across heterogeneous logs (Chagna and Goldschmidt, February 2026)
41 ▪ On the Difficulty of Selecting Few-Shot Examples for Effective LLM-based Vulnerability Detection (Hannan et al., February 2026)
42 ▪ SVIP: Towards Verifiable Inference of Open-source Large Language Models (Sun et al., February 2026)
43 ▪ GradingAttack: Attacking Large Language Models Towards Short Answer Grading Ability (Li et al., February 2026)
44 ▪ Evaluating Large Language Models for Security Bug Report Prediction (Soltaniani, Razzaq, and Ghafari, February 2026)
45 ▪ Lightweight LLMs for Network Attack Detection in IoT Networks (Sudasinghe, Liyanage, and Pussewalage, January 2026)
46 ▪ Holmes: An Evidence-Grounded LLM Agent for Auditable DDoS Investigation in Cloud Networks (Chen et al., January 2026)
47 ▪ Proactively Detecting Threats: A Novel Approach Using LLMs (Chawla and Prasad, January 2026)
48 ▪ LLMs as verification oracles for Solidity (Bartoletti, Lipparini, and Pompianu, January 2026)
49 ▪ Beyond Single Bugs: Benchmarking Large Language Models for Multi-Vulnerability Detection (Pushkar et al., December 2025)
50 ▪ Evaluating Large Language Models for Line-Level Vulnerability Localization (Zhang et al., December 2025)
51 ▪ VET Your Agent: Towards Host-Independent Autonomy via Verifiable Execution Traces (Grigor et al., December 2025)
52 ▪ Exact Verification of Graph Neural Networks with Incremental Constraint Solving (Liu, Lu, and Kwiatkowska, December 2025)
53 ▪ Clip-and-Verify: Linear Constraint-Driven Domain Clipping for Accelerating Neural Network Verification (Zhou et al., December 2025)
54 ▪ Verifying LLM Inference to Detect Model Weight Exfiltration (Rinberg et al., December 2025)
55 ▪ From Description to Score: Can LLMs Quantify Vulnerabilities? (Jafarikhah et al., December 2025)
56 ▪ Large Language Models Cannot Reliably Detect Vulnerabilities in JavaScript: The First Systematic Benchmark and Evaluation (Fei et al., December 2025)
57 ▪ Distillability of LLM Security Logic: Predicting Attack Success Rate of Outline Filling Attack via Ranking Regression (Zhang et al., December 2025)
58 ▪ Accuracy and Efficiency Trade-Offs in LLM-Based Malware Detection and Explanation: A Comparative Study of Parameter Tuning vs. Full Fine-Tuning (Gravereaux and Islam, November 2025)
59 ▪ AttackPilot: Autonomous Inference Attacks Against ML Services With LLM-Based Agents (Wu et al., November 2025)
60 ▪ LLM-CSEC: Empirical Evaluation of Security in C/C++ Code Generated by Large Language Models (Shahid, Ahmed, and Ranjan, November 2025)
61 ▪ Guided Reasoning in LLM-Driven Penetration Testing Using Structured Attack Trees (Nakano et al., November 2025)
62 ▪ CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning (Deason et al., November 2025)
63 ▪ Verifying LLM Inference to Prevent Model Weight Exfiltration (Rinberg et al., November 2025)
64 ▪ AthenaBench: A Dynamic Benchmark for Evaluating LLMs in Cyber Threat Intelligence (Alam et al., November 2025)
65 ▪ Scalable GPU-Based Integrity Verification for Large Machine Learning Models (Spoczynski and Melara, October 2025)
66 ▪ Floating-Point Neural Network Verification at the Software Level (Manino et al., October 2025)
67 ▪ A Systematic Approach to Predict the Impact of Cybersecurity Vulnerabilities Using LLMs (H{\o}st, Lison, and Moonen, October 2025)
68 ▪ Safeguarding Efficacy in Large Language Models: Evaluating Resistance to Human-Written and Algorithmic Adversarial Prompts (Downey-Webb, Jogunola, and Ajao, October 2025)
69 ▪ LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet? (Liu et al., October 2025)
70 ▪ CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale (Wang et al., October 2025)
71 ▪ VeriLLM: A Lightweight Framework for Publicly Verifiable Decentralized Inference (Wang et al., September 2025)
72 ▪ Automated Vulnerability Validation and Verification: A Large Language Model Approach (Lotfi, Katsis, and Bertino, September 2025)
73 ▪ DREAM: Scalable Red Teaming for Text-to-Image Generative Systems via Distribution Modeling (Li et al., July 2025)
74 ▪ LRCTI: A Large Language Model-Based Framework for Multi-Step Evidence Retrieval and Reasoning in Cyber Threat Intelligence Credibility Verification (Tang et al., July 2025)
Verify Attack (Efficacy)

>

<

‍

Insecure Output Handling

Covers:

  • OWASP LLM 02: Insecure Output Handling

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Insecure Output Handling

cybersecurity_tracker - Google Drive

2 ▪ "Tab, Tab, Bug": Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs (Lyu et al., May 2026)
3 ▪ Towards Secure Logging: Characterizing and Benchmarking Logging Code Security Issues with LLMs (Yuan et al., April 2026)
4 ▪ Surgical Repair of Insecure Code Generation in LLMs (Sandoval, Dolan-Gavitt, and Garg, April 2026)
5 ▪ From IOCs to Regex: Automating CTI Operationalization for SOC with LLMs (University et al., April 2026)
6 ▪ How Secure is Code Generated by ChatGPT? (Khoury et al., April 2026)
7 ▪ Causality Laundering: Denial-Feedback Leakage in Tool-Calling LLM Agents (Chinaei, April 2026)
8 ▪ Privacy Guard & Token Parsimony by Prompt and Context Handling and LLM Routing (Langiu, April 2026)
9 ▪ How Vulnerable Are Edge LLMs? (Ding et al., March 2026)
10 ▪ FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation (Wu et al., March 2026)
11 ▪ Vulnerabilities in Partial TEE-Shielded LLM Inference with Precomputed Noise (Saini, Jiang, and Liu, February 2026)
12 ▪ CryptoGen: Secure Transformer Generation with Encrypted KV-Cache Reuse (Zhang et al., February 2026)
13 ▪ "Tab, Tab, Bug'': Security Pitfalls of Next Edit Suggestions in AI-Integrated IDEs (Lyu et al., February 2026)
14 ▪ Supporting Students in Navigating LLM-Generated Insecure Code (Park et al., November 2025)
15 ▪ Think Fast: Real-Time IoT Intrusion Reasoning Using IDS and LLMs at the Edge Gateway (Jamshidi et al., November 2025)
16 ▪ GenSIaC: Toward Security-Aware Infrastructure-as-Code Generation with Large Language Models (Li et al., November 2025)
17 ▪ DrunkAgent: Stealthy Memory Corruption in LLM-Powered Recommender Agents (Yang et al., October 2025)
18 ▪ TypePilot: Leveraging the Scala Type System for Secure LLM-generated Code (Sternfeld, Kucharavy, and Dolamic, October 2025)
19 ▪ Semantic Encryption: Secure and Effective Interaction with Cloud-based Large Language Models via Semantic Transformation (Chen et al., August 2025)
20 ▪ eX-NIDS: A Framework for Explainable Network Intrusion Detection Leveraging Large Language Models (Houssel et al., July 2025)
21 ▪ Accelerating Automatic Program Repair with Dual Retrieval-Augmented Fine-Tuning and Patch Generation on Large Language Models (Guo et al., July 2025)
22 ▪ ETrace:Event-Driven Vulnerability Detection in Smart Contracts via LLM-Based Trace Analysis (Peng et al., July 2025)
23 ▪ DESIGN: Encrypted GNN Inference via Server-Side Input Graph Pruning (Zhao et al., July 2025)
24 ▪ Exploiting Defenses against GAN-Based Feature Inference Attacks in Federated Learning (Suo et al, Apr 2025)
25 ▪ Bayes-Nash Generative Privacy Against Membership Inference Attacks (Zhang et al, Feb 2025)
26 ▪ A hierarchical approach for assessing the vulnerability of tree-based classification models to membership inference attack (Preen and Smith, Feb 2025)
27 ▪ The Early Bird Catches the Leak: Unveiling Timing Side Channels in LLM Serving Systems (Song et al, Feb 2025)
28 ▪ LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states (Ibanez-Lissen et al, Jan 2025)
29 ▪ Time Will Tell: Timing Side Channels via Output Token Count in Large Language Models (Zhang, Saileshwar, and Lie, Dec 2024)
30 ▪ Privacy-Preserving Low-Rank Adaptation against Membership Inference Attacks for Latent Diffusion Models (Luo et al, Dec 2024)
31 ▪ Practical Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration (Fu et al, Nov 2024)
32 ▪ InputSnatch: Stealing Input in LLM Services via Timing Side-Channel Attacks (Zheng et al, Nov 2024)
33 ▪ Towards Black-Box Membership Inference Attack for Diffusion Models (Li et al, Nov 2024)
34 ▪ Protection against Source Inference Attacks in Federated Learning using Unary Encoding and Shuffling (Athanasiou, Jung, and Palmidessi, Nov 2024)
35 ▪ OSLO: One-Shot Label-Only Membership Inference Attacks (Peng et al, Oct 2024)
36 ▪ Detecting Training Data of Large Language Models via Expectation Maximization (Kim et al, Oct 2024)
37 ▪ Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models (Suriyakumar et al, Oct 2024)
38 ▪ Black-box Membership Inference Attacks against Fine-tuned Diffusion Models (Pang and Wang, Sep 2024)
39 ▪ Is Difficulty Calibration All We Need? Towards More Practical Membership Inference Attacks (He et al, Sep 2024)
40 ▪ Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based Policies (Wang et al, Aug 2024)
41 ▪ Membership Inference Attack Against Masked Image Modeling (Li et al, Aug 2024)
42 ▪ Pathway to Secure and Trustworthy 6G for LLMs: Attacks, Defense, and Opportunities (Khowaja et al, Aug 2024)
43 ▪ Synthetic Image Learning: Preserving Performance and Preventing Membership Inference Attacks (Lomurno and Matteucci, Jul 2024)
44 ▪ Thermometer: Towards Universal Calibration for Large Language Models (Shen et al, Jun 2024)
Insecure Output Handling

>

<

‍

Sensitive Information Disclosure

Covers:

  • OWASP LLM 06: Sensitive Information Disclosure

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Sensitive Information Disclosure

cybersecurity_tracker - Google Drive

2 ▪ Do Skill Descriptions Tell the Truth? Detecting Undisclosed Security Behaviors in Code-Backed LLM Skills (He et al., May 2026)
3 ▪ Differentially Private Synthetic Text Generation for Retrieval-Augmented Generation (RAG) (Mori et al., May 2026)
4 ▪ Reconstruction of Personally Identifiable Information from Supervised Finetuned Models (Furukawa and Oprea, May 2026)
5 ▪ Can You Keep a Secret? Involuntary Information Leakage in Language Model Writing (Holtzman and West, May 2026)
6 ▪ Deep Learning under Fractional-Order Differential Privacy (Partohaghighi and Marcia, May 2026)
7 ▪ Profiling for Pennies: Unveiling the Privacy Iceberg of LLM Agents (Chen et al., May 2026)
8 ▪ How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study (Wang et al., May 2026)
9 ▪ From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists (Cecconello et al., May 2026)
10 ▪ Graph Reconstruction from Differentially Private GNN Explanations (Sahoo, Shivottam, and Mishra, May 2026)
11 ▪ Evaluating Retrieval-Augmented Generation for Explainable Malware Analysis (Ng and Fard, May 2026)
12 ▪ PIIGuard: Mitigating PII Harvesting under Adversarial Sanitization (Liu, Zha, and Chen, May 2026)
13 ▪ Metric-Normalized Posterior Leakage (mPL): Attacker-Aligned Privacy for Joint Consumption (Chen et al., May 2026)
14 ▪ Revisiting Privacy Leakage in Machine Unlearning: Membership Inference Beyond the Forgotten Set (Fu et al., May 2026)
15 ▪ E-MIA: Exam-Style Black-Box Membership Inference Attacks against RAG Systems (Guan et al., May 2026)
16 ▪ When RAG Chatbots Expose Their Backend: An Anonymized Case Study of Privacy and Security Risks in Patient-Facing Medical AI (Madrid-Garc\'ia and Rujas, May 2026)
17 ▪ Detecting Malicious Concepts without Image Generation in AI-Generated Content (AIGC) (Xu et al., May 2026)
18 ▪ Tracking Conversations: Measuring Content and Identity Exposure on AI Chatbots (Jazlan et al., May 2026)
19 ▪ Spore: Efficient and Training-Free Privacy Extraction Attack on LLMs via Inference-Time Hybrid Probing (Cui et al., April 2026)
20 ▪ Privacy Leakage via Output Label Space and Differentially Private Continual Learning (Tobaben et al., April 2026)
21 ▪ Behavioral Canaries: Auditing Private Retrieved Context Usage in RL Fine-Tuning (Chen, Yuan, and Kairouz, April 2026)
22 ▪ Hidden Secrets in the arXiv: Discovering, Analyzing, and Preventing Unintentional Information Disclosure in Source Files of Scientific Preprints (Pennekamp et al., April 2026)
23 ▪ DP-FlogTinyLLM: Differentially private federated log anomaly detection using Tiny LLMs (Thompson, Sen, and Bhattacharya, April 2026)
24 ▪ Evaluating Answer Leakage Robustness of LLM Tutors against Adversarial Student Attacks (Zhao, Kne\v{z}evi\'c, and K\"aser, April 2026)
25 ▪ Breaking Euston: Recovering Private Inputs from Secure Inference by Exploiting Subspace Leakage (Zhao and Wang, April 2026)
26 ▪ PolicyGapper: Automated Detection of Inconsistencies Between Google Play Data Safety Sections and Privacy Policies Using LLMs (Ferrari et al., April 2026)
27 ▪ Too Private to Tell: Practical Token Theft Attacks on Apple Intelligence (Zhou et al., April 2026)
28 ▪ Automated Profile Inference with Language Model Agents (Du et al., April 2026)
29 ▪ A2-DIDM: Privacy-preserving Accumulator-enabled Auditing for Distributed Identity of DNN Model (Xie et al., April 2026)
30 ▪ CoLA: A Choice Leakage Attack Framework to Expose Privacy Risks in Subset Training (Li et al., April 2026)
31 ▪ Fully Homomorphic Encryption on Llama 3 model for privacy preserving LLM inference (Abdennebi, Kara, and Lahlou, April 2026)
32 ▪ LLM-Redactor: An Empirical Evaluation of Eight Techniques for Privacy-Preserving LLM Requests (Agyemang et al., April 2026)
33 ▪ Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game (Xie et al., April 2026)
34 ▪ ADAM: A Systematic Data Extraction Attack on Agent Memory via Adaptive Querying (Lyu et al., April 2026)
35 ▪ Conversations Risk Detection LLMs in Financial Agents via Multi-Stage Generative Rollout (Jiang and Wu, April 2026)
36 ▪ Security Concerns in Generative AI Coding Assistants: Insights from Online Discussions on GitHub Copilot (Ferreyra et al., April 2026)
37 ▪ ConfusionPrompt: Practical Private Inference for Online Large Language Models (Mai et al., April 2026)
38 ▪ Towards Privacy-Preserving Large Language Model: Text-free Inference Through Alignment and Adaptation (Yoon et al., April 2026)
39 ▪ Evaluating LLM-based Personal Information Extraction and Countermeasures (Liu et al., April 2026)
40 ▪ Exposing the Ghost in the Transformer: Abnormal Detection for Large Language Models via Hidden State Forensics (Zhou et al., April 2026)
41 ▪ No Attacker Needed: Unintentional Cross-User Contamination in Shared-State LLM Agents (Yang et al., April 2026)
42 ▪ Observable Channels, Not Just Storage: Evaluating Privacy Leakage in LLM Agent Pipelines (Huang et al., March 2026)
43 ▪ Towards Privacy-Preserving LLM Inference via Covariant Obfuscation (Technical Report) (Lin et al., March 2026)
44 ▪ "Elementary, My Dear Watson." Detecting Malicious Skills via Neuro-Symbolic Reasoning across Heterogeneous Artifacts (Wang et al., March 2026)
45 ▪ Not All Entities are Created Equal: A Dynamic Anonymization Framework for Privacy-Preserving Retrieval-Augmented Generation (Zhu et al., March 2026)
46 ▪ Computing Maximal Per-Record Leakage and Leakage-Distortion Functions for Privacy Mechanisms under Entropy-Constrained Adversaries (Wu et al., March 2026)
47 ▪ Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information (Zhan et al., March 2026)
48 ▪ A Critical Review on the Effectiveness and Privacy Threats of Membership Inference Attacks (Jebreel, S\'anchez, and Domingo-Ferrer, March 2026)
49 ▪ Beyond Theoretical Bounds: Empirical Privacy Loss Calibration for Text Rewriting Under Local Differential Privacy (Li et al., March 2026)
50 ▪ CIPL: A Target-Independent Framework for Channel-Inversion Privacy Leakage in Agents (Huang, Hou, and Meng, March 2026)
51 ▪ RedacBench: Can AI Erase Your Secrets? (Jeon, Kim, and Shin, March 2026)
52 ▪ PlanTwin: Privacy-Preserving Planning Abstractions for Cloud-Assisted LLM Agents (Yu et al., March 2026)
53 ▪ Revisiting Label Inference Attacks in Vertical Federated Learning: Why They Are Vulnerable and How to Defend (Liu et al., March 2026)
54 ▪ WebPII: Benchmarking Visual PII Detection for Computer-Use Agents (Zhao, March 2026)
55 ▪ SEAL-Tag: Self-Tag Evidence Aggregation with Probabilistic Circuits for PII-Safe Retrieval-Augmented Generation (Xie, Li, and Cheng, March 2026)
56 ▪ VisualLeakBench: Auditing the Fragility of Large Vision-Language Models against PII Leakage and Social Engineering (Wang et al., March 2026)
57 ▪ AEX: Non-Intrusive Multi-Hop Attestation and Provenance for LLM APIs (Guan, March 2026)
58 ▪ Membership Inference for Contrastive Pre-training Models with Text-only PII Queries (Cheng et al., March 2026)
59 ▪ A Decision-Theoretic Formalisation of Steganography With Applications to LLM Monitoring (Anwar et al., March 2026)
60 ▪ Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights (Fang and Kim, March 2026)
61 ▪ STAMP: Selective Task-Aware Mechanism for Text Privacy (Tian et al., March 2026)
62 ▪ WebWeaver: Breaking Topology Confidentiality in LLM Multi-Agent Systems with Stealthy Context-Based Inference (Xiong et al., March 2026)
63 ▪ CLIOPATRA: Extracting Private Information from LLM Insights (Annamalai, Cristofaro, and Kairouz, March 2026)
64 ▪ Automated TEE Adaptation with LLMs: Identifying, Transforming, and Porting Sensitive Functions in Programs (Han et al., March 2026)
65 ▪ Exposing Citation Vulnerabilities in Generative Engines (Mochizuki et al., March 2026)
66 ▪ BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprints (Gill, Isak, and Dressman, March 2026)
67 ▪ Real Money, Fake Models: Deceptive Model Claims in Shadow APIs (Zhang et al., March 2026)
68 ▪ Towards Privacy-Preserving LLM Inference via Collaborative Obfuscation (Technical Report) (Lin et al., March 2026)
69 ▪ Your Inference Request Will Become a Black Box: Confidential Inference for Cloud-based Large Language Models (Huang et al., March 2026)
70 ▪ Assessing Deanonymization Risks with Stylometry-Assisted LLM Agent (Zhang and Zhang, February 2026)
71 ▪ Personal Information Parroting in Language Models (Subramani, Ghate, and Diab, February 2026)
72 ▪ A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage (Xin et al., February 2026)
73 ▪ FeatureBleed: Inferring Private Enriched Attributes From Sparsity-Optimized AI Accelerators (Asher et al., February 2026)
74 ▪ Discovering Universal Activation Directions for PII Leakage in Language Models (Marchyok et al., February 2026)
75 ▪ Large-scale online deanonymization with LLMs (Lermen et al., February 2026)
76 ▪ NLP Privacy Risk Identification in Social Media (NLP-PRISM): A Survey (Goswami, Kumar, and Das, February 2026)
77 ▪ PII-Bench: Evaluating Query-Aware Privacy Protection Systems (Shen et al., February 2026)
78 ▪ SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs (Shaaban and Elmahallawy, February 2026)
79 ▪ Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs (Yan et al., February 2026)
80 ▪ CAPID: Context-Aware PII Detection for Question-Answering Systems (Ponomarenko et al., February 2026)
81 ▪ Stateless Yet Not Forgetful: Implicit Memory as a Hidden Channel in LLMs (Salem, Paverd, and Abdelnabi, February 2026)
82 ▪ Retrieval Pivot Attacks in Hybrid RAG: Measuring and Mitigating Amplified Leakage from Vector Seeds to Graph Expansion (Thornton, February 2026)
83 ▪ Taipan: A Query-free Transfer-based Multiple Sensitive Attribute Inference Attack Solely from Publicly Released Graphs (Song and Palanisamy, February 2026)
84 ▪ Do Vision-Language Models Respect Contextual Integrity in Location Disclosure? (Yang et al., February 2026)
85 ▪ PriMod4AI: Lifecycle-Aware Privacy Threat Modeling for AI Systems using LLM (Savaliya et al., February 2026)
86 ▪ Evaluating the Vulnerability Landscape of LLM-Generated Smart Contracts (Do, Sohrabi, and Hassan, February 2026)
87 ▪ No More Hidden Pitfalls? Exposing Smart Contract Bad Practices with LLM-Powered Hybrid Analysis (Li et al., February 2026)
88 ▪ Security Analysis of ChatGPT: Threats and Privacy Risks (Xiang, Li, and Li, February 2026)
89 ▪ Enhancing Smart Contract Vulnerability Detection in DApps Leveraging Fine-Tuned LLM (Bu et al., February 2026)
90 ▪ Decoupling Generalizability and Membership Privacy Risks in Neural Networks (Fang and Kim, February 2026)
91 ▪ Semantic Leakage from Image Embeddings (Chen et al., February 2026)
92 ▪ Protecting Private Code in IDE Autocomplete using Differential Privacy (Grigorenko et al., February 2026)
93 ▪ Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models (Westphal, Navaie, and Rosas, February 2026)
94 ▪ Okara: Detection and Attribution of TLS Man-in-the-Middle Vulnerabilities in Android Apps with Foundation Models (Yang et al., February 2026)
95 ▪ AlienLM: Alienization of Language for API-Boundary Privacy in Black-Box LLMs (Kim and Kang, February 2026)
96 ▪ OD-Stega: LLM-Based Relatively Secure Steganography via Optimized Distributions (Huang et al., January 2026)
97 ▪ Enhancing Membership Inference Attacks on Diffusion Models from a Frequency-Domain Perspective (Lian et al., January 2026)
98 ▪ Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models (Luo et al., January 2026)
99 ▪ Mitigating Sensitive Information Leakage in LLMs4Code through Machine Unlearning (Gu et al., January 2026)
100 ▪ CanaryBench: Stress Testing Privacy Leakage in Cluster-Level Conversation Summaries (Mehta, January 2026)
101 ▪ Beyond Data Privacy: New Privacy Risks for Large Language Models (Du et al., January 2026)
102 ▪ Connect the Dots: Knowledge Graph-Guided Crawler Attack on Retrieval-Augmented Generation Systems (Yao et al., January 2026)
103 ▪ NeuroFilter: Privacy Guardrails for Conversational LLM Agents (Das and Fioretto, January 2026)
104 ▪ VidLeaks: Membership Inference Attacks Against Text-to-Video Models (Wang et al., January 2026)
105 ▪ On Membership Inference Attacks in Knowledge Distillation (Cui, Zhang, and Pei, January 2026)
106 ▪ Automated Generation of Accurate Privacy Captions From Android Source Code Using Large Language Models (Jain et al., January 2026)
107 ▪ Leveraging Membership Inference Attacks for Privacy Measurement in Federated Learning for Remote Sensing Images (Duong et al., January 2026)
108 ▪ PII-VisBench: Evaluating Personally Identifiable Information Safety in Vision Language Models Along a Continuum of Visibility (Shahariar et al., January 2026)
109 ▪ Agentic LLMs as Powerful Deanonymizers: Re-identification of Participants in the Anthropic Interviewer Dataset (Li, January 2026)
110 ▪ DP-MGTD: Privacy-Preserving Machine-Generated Text Detection via Adaptive Differentially Private Entity Sanitization (Wang et al., January 2026)
111 ▪ SoK: Privacy Risks and Mitigations in Retrieval-Augmented Generation Systems (Bodea et al., January 2026)
112 ▪ MemHunter: Automated and Verifiable Memorization Detection at Dataset-scale in LLMs (Wu et al., January 2026)
113 ▪ I Large Language Models possono nascondere un testo in un altro testo della stessa lunghezza (Norelli and Bronstein, January 2026)
114 ▪ InfoDecom: Decomposing Information for Defending Against Privacy Leakage in Split Inference (Deng, Lu, and Duan, January 2026)
115 ▪ Privacy Bias in Language Models: A Contextual Integrity-based Auditing Metric (Shvartzshnaider and Duddu, December 2025)
116 ▪ Detecting Malicious Entra OAuth Apps with LLM-Based Permission Risk Scoring (Mahara, December 2025)
117 ▪ PerProb: Indirectly Evaluating Memorization in Large Language Models (Liao et al., December 2025)
118 ▪ CTIGuardian: A Few-Shot Framework for Mitigating Privacy Leakage in Fine-Tuned LLMs (Arachchige et al., December 2025)
119 ▪ Taint-Based Code Slicing for LLMs-based Malicious NPM Package Detection (Nguyen et al., December 2025)
120 ▪ Gradient-Free Privacy Leakage in Federated Language Models through Selective Weight Tampering (Rashid et al., December 2025)
121 ▪ Argus: A Multi-Agent Sensitive Information Leakage Detection Framework Based on Hierarchical Reference Relationships (Wang et al., December 2025)
122 ▪ Exposing and Defending Membership Leakage in Vulnerability Prediction Models (Liao et al., December 2025)
123 ▪ Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach (Yang et al., December 2025)
124 ▪ CKG-LLM: LLM-Assisted Detection of Smart Contract Access Control Vulnerabilities Based on Knowledge Graphs (Li et al., December 2025)
125 ▪ Sell Data to AI Algorithms Without Revealing It: Secure Data Valuation and Sharing via Homomorphic Encryption (Yang et al., December 2025)
126 ▪ When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs (Chen et al., December 2025)
127 ▪ WildCode: An Empirical Analysis of Code Generated by ChatGPT (Khanmohammadi et al., December 2025)
128 ▪ Towards Contextual Sensitive Data Detection (Telkamp and Hulsebos, December 2025)
129 ▪ Privacy Risks and Preservation Methods in Explainable Artificial Intelligence: A Scoping Review (Allana, Kankanhalli, and Dara, December 2025)
130 ▪ Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference (Luo et al., December 2025)
131 ▪ GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents (Zhang et al., December 2025)
132 ▪ Vision Token Masking Alone Cannot Prevent PHI Leakage in Medical Document OCR: A Systematic Evaluation (Young, November 2025)
133 ▪ Privacy Auditing of Multi-domain Graph Pre-trained Model under Membership Inference Attacks (Luo et al., November 2025)
134 ▪ Membership Inference Attacks Beyond Overfitting (Khalil et al., November 2025)
135 ▪ Password Strength Analysis Through Social Network Data Exposure: A Combined Approach Relying on Data Reconstruction and Generative Models (Atzori et al., November 2025)
136 ▪ Privacy Preserving In-Context-Learning Framework for Large Language Models (Bhusal et al., November 2025)
137 ▪ Confidential Prompting: Privacy-preserving LLM Inference on Cloud (Li, Gim, and Zhong, November 2025)
138 ▪ Quantifying Privacy Leakage in Split Inference via Fisher-Approximated Shannon Information Analysis (Deng et al., November 2025)
139 ▪ Observational Auditing of Label Privacy (Kalemaj et al., November 2025)
140 ▪ Explainable Transformer-Based Email Phishing Classification with Adversarial Robustness (P, November 2025)
141 ▪ BudgetLeak: Membership Inference Attacks on RAG Systems via the Generation Budget Side Channel (Li et al., November 2025)
142 ▪ Privacy Challenges and Solutions in Retrieval-Augmented Generation-Enhanced LLMs for Healthcare Chatbots: A Review of Applications, Risks, and Future Directions (Guan et al., November 2025)
143 ▪ How Worrying Are Privacy Attacks Against Machine Learning? (Domingo-Ferrer, November 2025)
144 ▪ Private-RAG: Answering Multiple Queries with LLMs while Keeping Your Data Private (Wu et al., November 2025)
145 ▪ SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought (Batra et al., November 2025)
146 ▪ $\mathbf{S^2LM}$: Towards Semantic Steganography via Large Language Models (Wu et al., November 2025)
147 ▪ Whisper Leak: a side-channel attack on Large Language Models (McDonald and Or, November 2025)
148 ▪ Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework (Li et al., November 2025)
149 ▪ PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases (Bae et al., November 2025)
150 ▪ Exploring the limits of strong membership inference attacks on large language models (Hayes et al., November 2025)
151 ▪ FTSmartAudit: A Knowledge Distillation-Enhanced Framework for Automated Smart Contract Auditing Using Fine-Tuned LLMs (Wei et al., November 2025)
152 ▪ Prevalence of Security and Privacy Risk-Inducing Usage of AI-based Conversational Agents (Grosse and Ebert, November 2025)
153 ▪ NetEcho: From Real-World Streaming Side-Channels to Full LLM Conversation Recovery (Zhang et al., October 2025)
154 ▪ Learning to Attack: Uncovering Privacy Risks in Sequential Data Releases (Cui, Zhang, and Pei, October 2025)
155 ▪ Differential Privacy: Gradient Leakage Attacks in Federated Learning Environments (Fernandez-de-Retana et al., October 2025)
156 ▪ Membership Inference Attacks on Recommender System: A Survey (He et al., October 2025)
157 ▪ LLMs can hide text in other text of the same length (Norelli and Bronstein, October 2025)
158 ▪ LLMs can hide text in other text of the same length.ipynb (Norelli and Bronstein, October 2025)
159 ▪ The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models (Dodd et al., October 2025)
160 ▪ CircuitGuard: Mitigating LLM Memorization in RTL Code Generation Against IP Leakage (Mashnoor et al., October 2025)
161 ▪ Exploring Membership Inference Vulnerabilities in Clinical Large Language Models (Nemecek et al., October 2025)
162 ▪ Evaluating Large Language Models in detecting Secrets in Android Apps (Alecci et al., October 2025)
163 ▪ The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers (Makroo et al., October 2025)
164 ▪ Membership Inference over Diffusion-models-based Synthetic Tabular Data (Cheng and Bahmani, October 2025)
165 ▪ DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space Smoothing (Qiao et al., October 2025)
166 ▪ AndroByte: LLM-Driven Privacy Analysis through Bytecode Summarization and Dynamic Dataflow Call Graph Generation (Khatun et al., October 2025)
167 ▪ Lost in the Averages: A New Specific Setup to Evaluate Membership Inference Attacks Against Machine Learning Models (Kr\v{c}o et al., October 2025)
168 ▪ Early Signs of Steganographic Capabilities in Frontier LLMs (Zolkowski et al., October 2025)
169 ▪ Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving Technique (Li et al., October 2025)
170 ▪ The Model's Language Matters: A Comparative Privacy Analysis of LLMs (Mishra, Boutet, and Magnana, October 2025)
171 ▪ PATCH: Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit PatcHing (Hughes et al., October 2025)
172 ▪ Membership Inference Attacks on LLM-based Recommender Systems (He et al., October 2025)
173 ▪ Empirical Comparison of Membership Inference Attacks in Deep Transfer Learning (Bai et al., October 2025)
174 ▪ Membership Inference Attacks on Tokenizers of Large Language Models (Tong et al., October 2025)
175 ▪ Can We Infer Confidential Properties of Training Data from LLMs? (Huang et al., October 2025)
176 ▪ Rethinking Exact Unlearning under Exposure: Extracting Forgotten Data under Exact Unlearning in Large Language Model (Wu et al., October 2025)
177 ▪ Privacy Leakage Overshadowed by Views of AI: A Study on Human Oversight of Privacy in Language Model Agent (Zhang, Guo, and Li, October 2025)
178 ▪ MIA-EPT: Membership Inference Attack via Error Prediction for Tabular Data (German et al., October 2025)
179 ▪ Activation Functions Considered Harmful: Recovering Neural Network Weights through Controlled Channels (Spielman et al., October 2025)
180 ▪ Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers (Green et al., October 2025)
181 ▪ An Ethically Grounded LLM-Based Approach to Insider Threat Synthesis and Detection (Gelman, Hastings, and Kenley, October 2025)
182 ▪ Inducing Uncertainty on Open-Weight Models for Test-Time Privacy in Image Recognition (Ashiq et al., October 2025)
183 ▪ Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models (Zhang et al., October 2025)
184 ▪ When Speculation Spills Secrets: Side Channels via Speculative Decoding In LLMs (Wei et al., September 2025)
185 ▪ Sanitize Your Responses: Mitigating Privacy Leakage in Large Language Models (Fu et al., September 2025)
186 ▪ Active Authentication via Korean Keystrokes Under Varying LLM Assistance and Cognitive Contexts (Roh and Kumar, September 2025)
187 ▪ MaskSQL: Safeguarding Privacy for LLM-Based Text-to-SQL via Abstraction (Waterloo et al., September 2025)
188 ▪ RecPS: Privacy Risk Scoring for Recommender Systems (He, Gu, and Chen, July 2025)
189 ▪ Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents (Fang et al., July 2025)
190 ▪ Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents (Ngong et al., July 2025)
191 ▪ Let's Measure the Elephant in the Room: Facilitating Personalized Automated Analysis of Privacy Policies at Scale (Zhao et al., July 2025)
192 ▪ Auditing Prompt Caching in Language Model APIs (Gu et al., July 2025)
193 ▪ Measuring the Accuracy and Effectiveness of PII Removal Services (He et al., July 2025)
194 ▪ Clio-X: AWeb3 Solution for Privacy-Preserving AI Access to Digital Archives (Lemieux et al., July 2025)
195 ▪ PBa-LLM: Privacy- and Bias-aware NLP using Named-Entity Recognition (NER) (Mancera et al., July 2025)
196 ▪ The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation (Xiong et al., July 2025)
197 ▪ Real-Time Privacy Risk Measurement with Privacy Tokens for Gradient Leakage (Meng et al, Feb 2025)
198 ▪ Exploring Privacy and Fairness Risks in Sharing Diffusion Models: An Adversarial Perspective (Luo et al, Sep 2024)
199 ▪ LLM-PBE: Assessing Data Privacy in Large Language Models (Li et al, Sep 2024)
200 ▪ PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action (Shao et al, Sep 2024)
201 ▪ Privacy-preserving Universal Adversarial Defense for Black-box Models (Li et al, Aug 2024)
202 ▪ DePrompt: Desensitization and Evaluation of Personal Identifiable Information in Large Language Model Prompts (Sun et al, Aug 2024)
203 ▪ Casper: Prompt Sanitization for Protecting User Privacy in Web-Based Large Language Models (Chong et al, Aug 2024)
Sensitive Information Disclosure

>

<

‍

Insecure Plugin Design and Plugin Compromise

Covers:

  • OWASP LLM 07: Insecure Plugin Design
  • MITRE ATLAS Execution & Privilege Escalation

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Insecure Plugin Design and Plugin Compromise

cybersecurity_tracker - Google Drive

2 ▪ AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills (Zhuang et al., May 2026)
3 ▪ Trust Me, Import This: Dependency Steering Attacks via Malicious Agent Skills (Liu et al., May 2026)
4 ▪ Heimdallr: Characterizing and Detecting LLM-Induced Security Risks in GitHub CI Workflows (Ruan et al., May 2026)
5 ▪ SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents (Ouyang et al., May 2026)
6 ▪ ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers (Liu et al., March 2026)
7 ▪ Paladin: A Policy Framework for Securing Cloud APIs by Combining Application Context with Generative AI (Priya, Stephen, and Natarajan, March 2026)
8 ▪ OpenPort Protocol: A Security Governance Specification for AI Agent Tool Access (Zhu et al., February 2026)
9 ▪ Securing LLM-as-a-Service for Small Businesses: An Industry Case Study of a Distributed Chatbot Deployment Platform (Xie et al., January 2026)
10 ▪ Keep the Lights On, Keep the Lengths in Check: Plug-In Adversarial Detection for Time-Series LLMs in Energy Forecasting (Ma et al., December 2025)
11 ▪ When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot Plugins (Kaya et al., November 2025)
12 ▪ Security study based on the Chatgptplugin system: ldentifying Security Vulnerabilities (Ren, July 2025)
13 ▪ When Developer Aid Becomes Security Debt: A Systematic Analysis of Insecure Behaviors in LLM Coding Agents (Kozak, Moghaddam, and Sivaraman, July 2025)
14 ▪ PromptChain: A Decentralized Web3 Architecture for Managing AI Prompts as Digital Assets (Bara, July 2025)
15 ▪ We Urgently Need Privilege Management in MCP: A Measurement of API Usage in MCP Ecosystems (Li et al., July 2025)
16 ▪ The Philosopher's Stone: Trojaning Plugins of Large Language Models (Dong et al, Sep 2024)
17 ▪ LLM Platform Security: Applying a Systematic Evaluation Framework to OpenAI's ChatGPT Plugins (Iqbal, Kohno, and Roesner, Jul 2024)
Insecure Plugin Design and Plugin Compromise

>

<

‍

Hallucination Squatting and Phishing

Covers:

  • MITRE ATLAS Initial Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Hallucination Squatting and Phishing

cybersecurity_tracker - Google Drive

2 ▪ SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing (Chen et al., May 2026)
3 ▪ REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations (Liang et al., May 2026)
4 ▪ Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data (Vafa, Roy, and Nilizadeh, May 2026)
5 ▪ LLM Ghostbusters: Surgical Hallucination Suppression via Adaptive Unlearning (Spracklen et al., May 2026)
6 ▪ CyberCane: Neuro-Symbolic RAG for Privacy-Preserving Phishing Detection with Formal Ontology Reasoning (Hakim et al., April 2026)
7 ▪ GuardPhish: Securing Open-Source LLMs from Phishing Abuse (Mishra, Varshney, and Sahithi, April 2026)
8 ▪ Segment-Level Coherence for Robust Harmful Intent Probing in LLMs (He et al., April 2026)
9 ▪ Reducing Hallucination in Enterprise AI Workflows via Hybrid Utility Minimum Bayes Risk (HUMBR) (Fang et al., April 2026)
10 ▪ Hijacking Text Heritage: Hiding the Human Signature through Homoglyphic Substitution (Dilworth, April 2026)
11 ▪ The Phish, The Spam, and The Valid: Generating Feature-Rich Emails for Benchmarking LLMs (Toth, Bisztray, and Gruschka, March 2026)
12 ▪ Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models (Guesmi and Shafique, March 2026)
13 ▪ Tool Receipts, Not Zero-Knowledge Proofs: Practical Hallucination Detection for AI Agents (Basu, March 2026)
14 ▪ PhishDebate: An LLM-Based Multi-Agent Framework for Phishing Website Detection (Li et al., March 2026)
15 ▪ Breaking Semantic-Aware Watermarks via LLM-Guided Coherence-Preserving Semantic Injection (Gao et al., February 2026)
16 ▪ PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring (Liu et al., February 2026)
17 ▪ GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models (Xu et al., February 2026)
18 ▪ Hallucination-Resistant Security Planning with a Large Language Model (Hammar, Alpcan, and Lupu, February 2026)
19 ▪ Benchmarking Large Language Models for Zero-shot and Few-shot Phishing URL Detection (Hasan and BusiReddyGari, February 2026)
20 ▪ When Agents "Misremember" Collectively: Exploring the Mandela Effect in LLM-based Multi-Agent Systems (Xu et al., February 2026)
21 ▪ Eliciting Least-to-Most Reasoning for Phishing URL Detection (Trikilis et al., January 2026)
22 ▪ Predictive Coding and Information Bottleneck for Hallucination Detection in Large Language Models (Bhatt, January 2026)
23 ▪ CORVUS: Red-Teaming Hallucination Detectors via Internal Signal Camouflage in Large Language Models (Min et al., January 2026)
24 ▪ Can Large Language Models Automate Phishing Warning Explanations? A Controlled Experiment on Effectiveness and User Perception (Cau et al., December 2025)
25 ▪ SoK: Exposing the Generation and Detection Gaps in LLM-Generated Phishing Through Examination of Generation Methods, Content Characteristics, and Countermeasures (Chen et al., November 2025)
26 ▪ Trustworthiness Calibration Framework for Phishing Email Detection Using Large Language Models (Ganiuly and Smaiyl, November 2025)
27 ▪ Cross-Lingual Summarization as a Black-Box Watermark Removal Attack (Ganesan, October 2025)
28 ▪ MTRE: Multi-Token Reliability Estimation for Hallucination Detection in VLMs (Zollicoffer, Vu, and Bhattarai, October 2025)
29 ▪ BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? (Jiang et al., October 2025)
30 ▪ HauntAttack: When Attack Follows Reasoning as a Shadow (Ma et al., October 2025)
31 ▪ Elevating Cyber Threat Intelligence against Disinformation Campaigns with LLM-based Concept Extraction and the FakeCTI Dataset (Cotroneo, Natella, and Orbinato, October 2025)
32 ▪ Robust ML-based Detection of Conventional, LLM-Generated, and Adversarial Phishing Emails Using Advanced Text Preprocessing (Kulal et al., October 2025)
33 ▪ DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation (Ahn, Park, and Han, October 2025)
34 ▪ The Security Threat of Compressed Projectors in Large Vision-Language Models (Zhang et al., October 2025)
35 ▪ SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinations (Liang et al., October 2025)
36 ▪ Talking Like a Phisher: LLM-Based Attacks on Voice Phishing Classifiers (Li et al., July 2025)
37 ▪ DomainLynx: Leveraging Large Language Models for Enhanced Domain Squatting Detection (Chiba et al, Oct 2024)
38 ▪ We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (Spracklen et al, Sep 2024)
Hallucination Squatting and Phishing

>

<

‍

Persistence

Covers:

  • MITRE ATLAS Persistence

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Persistence

cybersecurity_tracker - Google Drive

2 ▪ Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents (Leong, May 2026)
3 ▪ Autonomous LLM Agent Worms: Cross-Platform Propagation, Automated Discovery and Temporal Re-Entry Defense (Zha and Wang, May 2026)
4 ▪ Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents (Tallam, April 2026)
5 ▪ ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore (Zheng et al., March 2026)
6 ▪ Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections (Yang et al., February 2026)
7 ▪ MemoryGraft: Persistent Compromise of LLM Agents via Poisoned Experience Retrieval (Srivastava and He, December 2025)
Persistence

>

<

‍

Backdoor ML Model and Craft Adversarial Data

Covers:

  • MITRE ATLAS ML Attack Staging

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Backdoor ML Model and Craft Adversarial Data

cybersecurity_tracker - Google Drive

2 ▪ Angel or Demon: Investigating the Plasticity Interventions' Impact on Backdoor Threats in Deep Reinforcement Learning (Ma et al., May 2026)
3 ▪ MetaBackdoor: Exploiting Positional Encoding as a Backdoor Attack Surface in LLMs (Wen et al., May 2026)
4 ▪ To See is Not to Learn: Protecting Multimodal Data from Unauthorized Fine-Tuning of Large Vision-Language Model (Zhao et al., May 2026)
5 ▪ Backdoor Threats in Variational Quantum Circuits: Taxonomy, Attacks, and Defenses (Jiang and Chen, May 2026)
6 ▪ Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks (Eggen et al., May 2026)
7 ▪ LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters (Ahn et al., May 2026)
8 ▪ DiffusionHijack: Supply-Chain PRNG Backdoor Attack on Diffusion Models and Quantum Random Number Defense (You et al., May 2026)
9 ▪ BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models (Rachapudi et al., May 2026)
10 ▪ FedSurrogate: Backdoor Defense in Federated Learning via Layer Criticality and Surrogate Replacement (Abacha et al., May 2026)
11 ▪ Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors (Yang et al., May 2026)
12 ▪ BadDLM: Backdooring Diffusion Language Models with Diverse Targets (Zhai et al., May 2026)
13 ▪ Enhancing Adversarial Robustness in Network Intrusion Detection: A Layer-wise Adaptive Regularization Approach (Nasir et al., May 2026)
14 ▪ Hammer and Anvil: Toward a Theory of Backdoors in Federated Learning (Fenaux et al., May 2026)
15 ▪ Activation Differences Reveal Backdoors: A Comparison of SAE Architectures (Kumar, May 2026)
16 ▪ Cross-Modal Backdoors in Multimodal Large Language Models (Wang et al., May 2026)
17 ▪ DeTrigger: A Gradient-Centric Approach to Backdoor Attack Mitigation in Federated Learning (Lee et al., May 2026)
18 ▪ Backdoor Mitigation in Object Detection via Adversarial Fine-Tuning (Dunnett et al., May 2026)
19 ▪ Stateful Agent Backdoor (Dai et al., May 2026)
20 ▪ Coward: Collision-based OOD Watermarking for Practical Proactive Federated Backdoor Detection (Li et al., May 2026)
21 ▪ The Adversarial Discount - AI, Signal Correlation, and the Cybersecurity Arms Race (Bono, May 2026)
22 ▪ Laundering AI Authority with Adversarial Examples (Zhang et al., May 2026)
23 ▪ Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions (Choudhary et al., May 2026)
24 ▪ Detecting Adversarial Data via Provable Adversarial Noise Amplification (Mumcu and Yilmaz, May 2026)
25 ▪ Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration (Das et al., May 2026)
26 ▪ VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models (Liu and Lao, May 2026)
27 ▪ Checkerboard: A Simple, Effective, Efficient and Learning-free Clean Label Backdoor Attack with Low Poisoning Budget (Yang et al., May 2026)
28 ▪ Attention Is Where You Attack (Srivastava and Panda, May 2026)
29 ▪ One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness (Deguchi, Chousa, and Sakai, May 2026)
30 ▪ Low Rank Adaptation for Adversarial Perturbation (Liu et al., May 2026)
31 ▪ Understanding Adversarial Transferability in Vision-Language Models for Autonomous Driving: A Cross-Architecture Analysis (Fernandez et al., May 2026)
32 ▪ Variational Autoencoder-Based Black-Box Adversarial Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, April 2026)
33 ▪ Invisible Hands: Gray-Box Bit Flip Attack for Steering LLMs Without Knowledge of Gradients, Data, and Weights (Almalky et al., April 2026)
34 ▪ CacheTrap: Unveiling a Stealthier Gray-Box Trojan against LLMs (Nahian et al., April 2026)
35 ▪ MEASER: Malware embedding attacks on open-source LLMs (Tan et al., April 2026)
36 ▪ Prototype-Guided Robust Learning against Backdoor Attacks (Guo et al., April 2026)
37 ▪ Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training (Zhao et al., April 2026)
38 ▪ DETOUR: A Practical Backdoor Attack against Object Detection (Liu et al., April 2026)
39 ▪ Defusing the Trigger: Plug-and-Play Defense for Backdoored LLMs via Tail-Risk Intrinsic Geometric Smoothing (Fan et al., April 2026)
40 ▪ Toward Polymorphic Backdoor against Semantic Communication via Intensity-Based Poisoning (Yang et al., April 2026)
41 ▪ Adversarial Malware Generation in Linux ELF Binaries via Semantic-Preserving Transformations (Hrdonka and Jure\v{c}ek, April 2026)
42 ▪ Adversarial Co-Evolution of Malware and Detection Models: A Bilevel Optimization Perspective (Jure\v{c}kov\'a et al., April 2026)
43 ▪ Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers (Wei et al., April 2026)
44 ▪ PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers (Liu et al., April 2026)
45 ▪ Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients (Yuan et al., April 2026)
46 ▪ Penny Wise, Pixel Foolish: Bypassing Price Constraints in Multimodal Agents via Visual Adversarial Perturbations (Qian and Kang, April 2026)
47 ▪ Beyond Text Prompts: Precise Concept Erasure through Text-Image Collaboration (Li et al., April 2026)
48 ▪ TopFeaRe: Locating Critical State of Adversarial Resilience for Graphs Regarding Topology-Feature Entanglement (Fan et al., April 2026)
49 ▪ Improving Clean Accuracy via a Tangent-Space Perspective on Adversarial Training (Yi, Lai, and Li, April 2026)
50 ▪ Robustness of Vision Foundation Models to Common Perturbations (Liu et al., April 2026)
51 ▪ NeuroTrace: Inference Provenance-Based Detection of Adversarial Examples (Hmida et al., April 2026)
52 ▪ INTARG: Informed Real-Time Adversarial Attack Generation for Time-Series Regression (Tokgoz et al., April 2026)
53 ▪ Scaling Exposes the Trigger: Input-Level Backdoor Detection in Text-to-Image Diffusion Models via Cross-Attention Scaling (Li et al., April 2026)
54 ▪ Compiling Activation Steering into Weights via Null-Space Constraints for Stealthy Backdoors (Yin et al., April 2026)
55 ▪ Spatiotemporal-Aware Bit-Flip Injection on DNN-based Advanced Driver Assistance Systems (extended version) (Zhao et al., April 2026)
56 ▪ Look Twice before You Leap: A Rational Framework for Localized Adversarial Anonymization (Duan et al., April 2026)
57 ▪ Property-Preserving Hashing for $\ell_1$-Distance Predicates: Applications to Countering Adversarial Input Attacks (Asghar, Zhang, and Kaafar, April 2026)
58 ▪ Defending against Backdoor Attacks via Module Switching (Li et al., April 2026)
59 ▪ Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models (Truong and Le, April 2026)
60 ▪ Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward (Guo et al., April 2026)
61 ▪ BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning (Tie et al., April 2026)
62 ▪ Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor (Chang et al., April 2026)
63 ▪ CLIP-Inspector: Model-Level Backdoor Detection for Prompt-Tuned CLIP via OOD Trigger Inversion (Jindal et al., April 2026)
64 ▪ Follow My Eyes: Backdoor Attacks on VLM-based Scanpath Prediction (Romero et al., April 2026)
65 ▪ BoBa: Boosting Backdoor Detection through Data Distribution Inference in Federated Learning (Jiang et al., April 2026)
66 ▪ BadImplant: Injection-based Multi-Targeted Graph Backdoor Attack (Khan, Miah, and Bi, April 2026)
67 ▪ CAAP: Capture-Aware Adversarial Patch Attacks on Palmprint Recognition Models (Liu et al., April 2026)
68 ▪ MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning (Zeng et al., April 2026)
69 ▪ SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems (Feng et al., April 2026)
70 ▪ Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models (Fortier et al., April 2026)
71 ▪ Stealthy and Adjustable Text-Guided Backdoor Attacks on Multimodal Pretrained Models (Zhang et al., April 2026)
72 ▪ Your LLM Agent Can Leak Your Data: Data Exfiltration via Backdoored Tool Use (Zhang and Pei, April 2026)
73 ▪ FABLE: A Localized, Targeted Adversarial Attack on Weather Forecasting Models (Deng et al., April 2026)
74 ▪ Explainability-Guided Adversarial Attacks on Transformer-Based Malware Detectors Using Control Flow Graphs (Wheeler, Aryal, and Gupta, April 2026)
75 ▪ ProtoGuard-SL: Prototype Consistency Based Backdoor Defense for Vertical Split Learning (Shui et al., April 2026)
76 ▪ S$^4$ST: A Strong, Self-transferable, faSt, and Simple Scale Transformation for Transferable Targeted Attack (Liu et al., April 2026)
77 ▪ Street-Legal Physical-World Adversarial Rim for License Plates (Kalidasu and Ganapathy, April 2026)
78 ▪ Backdoor Attacks on Decentralised Post-Training (Ersoy et al., April 2026)
79 ▪ Towards Physically Realizable Adversarial Attenuation Patch against SAR Object Detection (Zhang, Qin, and Wang, April 2026)
80 ▪ Spike-PTSD: A Bio-Plausible Adversarial Example Attack on Spiking Neural Networks via PTSD-Inspired Spike Scaling (Jin et al., April 2026)
81 ▪ Diffusion-Guided Adversarial Perturbation Injection for Generalizable Defense Against Facial Manipulations (Li et al., April 2026)
82 ▪ Adversarial Attenuation Patch Attack for SAR Object Detection (Zhang, Qin, and Wang, April 2026)
83 ▪ Dummy-Aware Weighted Attack (DAWA): Breaking the Safe Sink in Dummy Class Defenses (Yu et al., April 2026)
84 ▪ Beyond Corner Patches: Semantics-Aware Backdoor Attack in Federated Learning (Herath, Zhao, and Bagchi, April 2026)
85 ▪ Trojan-Speak: Bypassing Constitutional Classifiers with No Jailbreak Tax via Adversarial Finetuning (Sel et al., April 2026)
86 ▪ SNEAKDOOR: Stealthy Backdoor Attacks against Distribution Matching-based Dataset Condensation (Yang et al., April 2026)
87 ▪ Graph-Aware Stealthy Poison-Text Backdoors for Text-Attributed Graphs (Luo et al., March 2026)
88 ▪ FlowPure: Continuous Normalizing Flows for Adversarial Purification (Collaert et al., March 2026)
89 ▪ Seeking Flat Minima over Diverse Surrogates for Improved Adversarial Transferability: A Theoretical Framework and Algorithmic Instantiation (Zheng et al., March 2026)
90 ▪ FL-PBM: Pre-Training Backdoor Mitigation for Federated Learning (Wehbi et al., March 2026)
91 ▪ Mitigating Backdoor Attacks in Federated Learning Using PPA and MiniMax Game Theory (Wehbi et al., March 2026)
92 ▪ Detection of Adversarial Attacks in Robotic Perception (Sharawy, Nakshbandiand, and Grigorescu, March 2026)
93 ▪ Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models (Yao et al., March 2026)
94 ▪ Attacking AI Accelerators by Leveraging Arithmetic Properties of Addition (Heidary and Joardar, March 2026)
95 ▪ A Channel-Triggered Backdoor Attack on Wireless Semantic Image Reconstruction (Wan et al., March 2026)
96 ▪ On the Vulnerability of Deep Automatic Modulation Classifiers to Explainable Backdoor Threats (Salmi and Bogucka, March 2026)
97 ▪ Physical Backdoor Attack Against Deep Learning-Based Modulation Classification (Salmi and Bogucka, March 2026)
98 ▪ IrisFP: Adversarial-Example-based Model Fingerprinting with Enhanced Uniqueness and Robustness (Geng et al., March 2026)
99 ▪ Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach (Liao et al., March 2026)
100 ▪ Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs (Panfilov et al., March 2026)
101 ▪ Precision-Varying Prediction (PVP): Robustifying ASR systems against adversarial attacks (Pizarro, Narasimhan, and Fischer, March 2026)
102 ▪ Adversarial Vulnerabilities in Neural Operator Digital Twins: Gradient-Free Attacks on Nuclear Thermal-Hydraulic Surrogates (Roy et al., March 2026)
103 ▪ TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches (Huang et al., March 2026)
104 ▪ AgentRAE: Remote Action Execution through Notification-based Visual Backdoors against Screenshots-based Mobile GUI Agents (Luo et al., March 2026)
105 ▪ Architectural Backdoors for Within-Batch Data Stealing and Model Inference Manipulation (K\"uchler et al., March 2026)
106 ▪ Adversarial Attacks on Locally Private Graph Neural Networks (Kharagpur et al., March 2026)
107 ▪ Graph-Aware Text-Only Backdoor Poisoning for Text-Attributed Graphs (Luo et al., March 2026)
108 ▪ Trojan horse hunt in deep forecasting models: Insights from the European Space Agency competition (Kotowski et al., March 2026)
109 ▪ Trojan's Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance (Liu et al., March 2026)
110 ▪ Attack by Unlearning: Unlearning-Induced Adversarial Attacks on Graph Neural Networks (Zhang, Wang, and Wang, March 2026)
111 ▪ MAED: Mathematical Activation Error Detection for Mitigating Physical Fault Attacks in DNN Inference (Ahmadi et al., March 2026)
112 ▪ STEP: Detecting Audio Backdoor Attacks via Stability-based Trigger Exposure Profiling (Wang et al., March 2026)
113 ▪ Adversarial attacks against Modern Vision-Language Models (Torre, March 2026)
114 ▪ Coded Robust Aggregation for Distributed Learning under Byzantine Attacks (Li, Xiao, and Skoglund, March 2026)
115 ▪ Poisoning the Pixels: Revisiting Backdoor Attacks on Semantic Segmentation (Zhang et al., March 2026)
116 ▪ BadLLM-TG: A Backdoor Defender powered by LLM Trigger Generator (Zhang et al., March 2026)
117 ▪ Inevitable Encounters: Backdoor Attacks Involving Lossy Compression (Li, Chen, and Chen, March 2026)
118 ▪ Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference (Li and Kim, March 2026)
119 ▪ Test-Time Attention Purification for Backdoored Large Vision Language Models (Zhang et al., March 2026)
120 ▪ RTD-Guard: A Black-Box Textual Adversarial Detection Framework via Replacement Token Detection (Zhu et al., March 2026)
121 ▪ Purify Once, Edit Freely: Breaking Image Protections under Model Mismatch (Zhao et al., March 2026)
122 ▪ Cascade: Composing Software-Hardware Attack Gadgets for Adversarial Threat Amplification in Compound AI Systems (Banerjee et al., March 2026)
123 ▪ Delayed Backdoor Attacks: Exploring the Temporal Dimension as a New Attack Surface in Pre-Trained Models (Ding et al., March 2026)
124 ▪ Forging the Unforgeable: On the Feasibility of Counterfeit Watermarks in Backdoor-Based Dataset Ownership Verification (Li et al., March 2026)
125 ▪ Backdoor Directions in Vision Transformers (Karayalcin et al., March 2026)
126 ▪ Repurposing Backdoors for Good: Ephemeral Intrinsic Proofs for Verifiable Aggregation in Cross-silo Federated Learning (Qin, Yang, and Tang, March 2026)
127 ▪ Detecting and Eliminating Neural Network Backdoors Through Active Paths with Application to Intrusion Detection (H{\o}yheim et al., March 2026)
128 ▪ Amnesia: Adversarial Semantic Layer Specific Activation Steering in Large Language Models (Raza et al., March 2026)
129 ▪ TASER: Task-Aware Spectral Energy Refine for Backdoor Suppression in UAV Swarms Decentralized Federated Learning (Huang and Yang, March 2026)
130 ▪ Targeted Bit-Flip Attacks on LLM-Based Agents (Wang et al., March 2026)
131 ▪ VisPoison: An Effective Backdoor Attack Framework for Tabular Data Visualization Models (Li et al., March 2026)
132 ▪ Removing the Trigger, Not the Backdoor: Alternative Triggers and Latent Backdoors (Abad et al., March 2026)
133 ▪ NetDiffuser: Deceiving DNN-Based Network Attack Detection Systems with Diffusion-Generated Adversarial Traffic (Kumar et al., March 2026)
134 ▪ DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models (Xu et al., March 2026)
135 ▪ SFIBA: Spatial-based Full-target Invisible Backdoor Attacks (Yin et al., March 2026)
136 ▪ SlowBA: An efficiency backdoor attack towards VLM-based GUI agents (Li et al., March 2026)
137 ▪ Backdoor4Good: Benchmarking Beneficial Uses of Backdoors in LLMs (Li et al., March 2026)
138 ▪ Structure-Aware Distributed Backdoor Attacks in Federated Learning (Jian et al., March 2026)
139 ▪ Sleeper Cell: Injecting Latent Malice Temporal Backdoors into Tool-Using LLMs (Pallakonda et al., March 2026)
140 ▪ DSBA: Dynamic Stealthy Backdoor Attack with Collaborative Optimization in Self-Supervised Learning (Wang et al., March 2026)
141 ▪ TraceGuard: Process-Guided Firewall against Reasoning Backdoors in Large Language Models (Guo et al., March 2026)
142 ▪ Physical Evaluation of Naturalistic Adversarial Patches for Camera-Based Traffic-Sign Detection (D'Urso et al., March 2026)
143 ▪ BadRSSD: Backdoor Attacks on Regularized Self-Supervised Diffusion Models (Wang et al., March 2026)
144 ▪ IU: Imperceptible Universal Backdoor Attack (Lin et al., March 2026)
145 ▪ ProtegoFed: Backdoor-Free Federated Instruction Tuning with Interspersed Poisoned Data (Zhao et al., March 2026)
146 ▪ DropVLA: An Action-Level Backdoor Attack on Vision--Language--Action Models (Xu et al., February 2026)
147 ▪ Self-Purification Mitigates Backdoors in Multimodal Diffusion Language Models (Wan et al., February 2026)
148 ▪ Silent Until Sparse: Backdoor Attacks on Semi-Structured Sparsity (Guo et al., February 2026)
149 ▪ Is the Trigger Essential? A Feature-Based Triggerless Backdoor Attack in Vertical Federated Learning (Liu et al., February 2026)
150 ▪ When Backdoors Go Beyond Triggers: Semantic Drift in Diffusion Models Under Encoder Attacks (Chen and Zhu, February 2026)
151 ▪ Can a Teenager Fool an AI? Evaluating Low-Cost Cosmetic Attacks on Age Estimation Systems (Shen et al., February 2026)
152 ▪ SafePickle: Robust and Generic ML Detection of Malicious Pickle-based ML Models (Ohayon, Gilkarov, and Dubin, February 2026)
153 ▪ DCInject: Persistent Backdoor Attacks via Frequency Manipulation in Personal Federated Learning (Birhan et al., February 2026)
154 ▪ TrapFlow: Controllable Website Fingerprinting Defense via Dynamic Backdoor Learning (Liang et al., February 2026)
155 ▪ TFL: Targeted Bit-Flip Attack on Large Language Model (Guo, Chakrabarti, and Fan, February 2026)
156 ▪ Cert-SSBD: Certified Backdoor Defense with Sample-Specific Smoothing Noises (Qiao et al., February 2026)
157 ▪ Revisiting Backdoor Threat in Federated Instruction Tuning from a Signal Aggregation Perspective (Zhao, Hu, and Liu, February 2026)
158 ▪ Weight space Detection of Backdoors in LoRA Adapters (Merenciano et al., February 2026)
159 ▪ Exploiting Layer-Specific Vulnerabilities to Backdoor Attack in Federated Learning (Foroughi et al., February 2026)
160 ▪ Backdooring Bias in Large Language Models (Das et al., February 2026)
161 ▪ Backdoor Attacks on Contrastive Continual Learning for IoT Systems (Tim and D, February 2026)
162 ▪ PBP: Post-training Backdoor Purification for Malware Classifiers (Nguyen et al., February 2026)
163 ▪ Transferable Backdoor Attacks for Code Models via Sharpness-Aware Adversarial Perturbation (Chang et al., February 2026)
164 ▪ Kill it with FIRE: On Leveraging Latent Space Directions for Runtime Backdoor Mitigation in Deep Neural Networks (Ahlers et al., February 2026)
165 ▪ Understanding and Enhancing Encoder-based Adversarial Transferability against Large Vision-Language Models (Zhang et al., February 2026)
166 ▪ BadSNN: Backdoor Attacks on Spiking Neural Networks via Adversarial Spiking Neuron (Miah, Vu, and Bi, February 2026)
167 ▪ Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations (Miah and Bi, February 2026)
168 ▪ Trojans in Artificial Intelligence (TrojAI) Final Report (Reese et al., February 2026)
169 ▪ AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models (Li et al., February 2026)
170 ▪ Plato's Form: Toward Backdoor Defense-as-a-Service for LLMs with Prototype Representations (Chen et al., February 2026)
171 ▪ Dependable Artificial Intelligence with Reliability and Security (DAIReS): A Unified Syndrome Decoding Approach for Hallucination and Backdoor Trigger Detection (Studies et al., February 2026)
172 ▪ BadTemplate: A Training-Free Backdoor Attack via Chat Template Against Large Language Models (Wang et al., February 2026)
173 ▪ Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning (Rathbun et al., February 2026)
174 ▪ Semantic-level Backdoor Attack against Text-to-Image Diffusion Models (Chen et al., February 2026)
175 ▪ Semantic Consensus Decoding: Backdoor Defense for Verilog Code Generation (Yang et al., February 2026)
176 ▪ Inference-Time Backdoors via Hidden Instructions in LLM Chat Templates (Fogel et al., February 2026)
177 ▪ Time Is All It Takes: Spike-Retiming Attacks on Event-Driven Spiking Neural Networks (Yu et al., February 2026)
178 ▪ The Trigger in the Haystack: Extracting and Reconstructing LLM Backdoor Triggers (Bullwinkel et al., February 2026)
179 ▪ DF-LoGiT: Data-Free Logic-Gated Backdoor Attacks in Vision Transformers (Shen et al., February 2026)
180 ▪ Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection (Mukherjee et al., February 2026)
181 ▪ HPE: Hallucinated Positive Entanglement for Backdoor Attacks in Federated Self-Supervised Learning (Wang et al., February 2026)
182 ▪ Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency (Wang et al., February 2026)
183 ▪ RPP: A Certified Poisoned-Sample Detection Framework for Backdoor Attacks under Dataset Imbalance (Lin et al., February 2026)
184 ▪ Impact of Phonetics on Speaker Identity in Adversarial Voice Attack (Dar et al., February 2026)
185 ▪ Beauty and the Beast: Imperceptible Perturbations Against Diffusion-Based Face Swapping via Directional Attribute Editing (Huang and Li, February 2026)
186 ▪ An Effective Energy Mask-based Adversarial Evasion Attacks against Misclassification in Speaker Recognition Systems (Park and Kim, February 2026)
187 ▪ Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation (Li et al., January 2026)
188 ▪ Hardware-Triggered Backdoors (M\"oller et al., January 2026)
189 ▪ BadDet+: Robust Backdoor Attacks for Object Detection (Dunnett et al., January 2026)
190 ▪ SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning (Khan, Chandio, and Anwar, January 2026)
191 ▪ From Internal Diagnosis to External Auditing: A VLM-Driven Paradigm for Online Test-Time Backdoor Defense (Xu et al., January 2026)
192 ▪ TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion (Sreekumar et al., January 2026)
193 ▪ Multi-Targeted Graph Backdoor Attack (Khan, Miah, and Bi, January 2026)
194 ▪ SpooFL: Spoofing Federated Learning (Baglin, Zhu, and Hadfield, January 2026)
195 ▪ Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs (Lu et al., January 2026)
196 ▪ SilentDrift: Exploiting Action Chunking for Stealthy Backdoor Attacks on Vision-Language-Action Models (Xu et al., January 2026)
197 ▪ DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing (Hu et al., January 2026)
198 ▪ SoK: On the Survivability of Backdoor Attacks on Unconstrained Face Recognition Systems (Roux et al., January 2026)
199 ▪ SecureSplit: Mitigating Backdoor Attacks in Split Learning (Dou et al., January 2026)
200 ▪ MirageNet:A Secure, Efficient, and Scalable On-Device Model Protection in Heterogeneous TEE and GPU System (Zheng, Cheng, and Ding, January 2026)
201 ▪ DUAP: Dual-task Universal Adversarial Perturbations Against Voice Control Systems (Sun et al., January 2026)
202 ▪ Malware Classification using Diluted Convolutional Neural Network with Fast Gradient Sign Method (Anand et al., January 2026)
203 ▪ BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation (Wu et al., January 2026)
204 ▪ STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio (Park et al., January 2026)
205 ▪ Double Strike: Breaking Approximation-Based Side-Channel Countermeasures for DNNs (Casalino et al., January 2026)
206 ▪ ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models (Xu et al., January 2026)
207 ▪ BadPatches: Routing-aware Backdoor Attacks on Vision Mixture of Experts (Chan, Lintelo, and Picek, January 2026)
208 ▪ HogVul: Black-box Adversarial Code Generation Framework Against LM-based Vulnerability Detectors (Yang et al., January 2026)
209 ▪ Unified Framework for Qualifying Security Boundary of PUFs Against Machine Learning Attacks (Fei et al., January 2026)
210 ▪ State Backdoor: Towards Stealthy Real-world Poisoning Attack on Vision-Language-Action Model in State Space (Guo et al., January 2026)
211 ▪ Inhibitory Attacks on Backdoor-based Fingerprinting for Large Language Models (Fu et al., January 2026)
212 ▪ Beyond Immediate Activation: Temporally Decoupled Backdoor Attacks on Time Series Forecasting (Liu et al., January 2026)
213 ▪ Adversarial Contrastive Learning for LLM Quantization Attacks (Song et al., January 2026)
214 ▪ Non-omniscient backdoor injection with one poison sample: Proving the one-poison hypothesis for linear regression, linear classification, and 2-layer ReLU neural networks (Peinemann et al., January 2026)
215 ▪ SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models (Xue, Zhang, and Xie, January 2026)
216 ▪ SLIP: Soft Label Mechanism and Key-Extraction-Guided CoT-based Defense Against Instruction Backdoor in APIs (Wu et al., January 2026)
217 ▪ Coward: Collision-based Watermark for Proactive Federated Backdoor Detection (Li et al., January 2026)
218 ▪ From Chat Control to Robot Control: The Backdoors Left Open for the Sake of Safety (Akalin and Giaretta, January 2026)
219 ▪ IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection (Zhu et al., January 2026)
220 ▪ Crafting Adversarial Inputs for Large Vision-Language Models Using Black-Box Optimization (Guan, Jin, and Wang, January 2026)
221 ▪ NADD: Amplifying Noise for Effective Diffusion-based Adversarial Purification (Nguyen et al., January 2026)
222 ▪ The Trojan in the Vocabulary: Stealthy Sabotage of LLM Composition (Liu et al., January 2026)
223 ▪ Rectifying Adversarial Examples Using Their Vulnerabilities (Morimoto, Morita, and Ono, January 2026)
224 ▪ BadBlocks: Lightweight and Stealthy Backdoor Threat in Text-to-Image Diffusion Models (Pan et al., January 2026)
225 ▪ Breaking Audio Large Language Models by Attacking Only the Encoder: A Universal Targeted Latent-Space Audio Attack (Ziv, Lapid, and Sipper, January 2026)
226 ▪ Training-Free Color-Aware Adversarial Diffusion Sanitization for Diffusion Stegomalware Defense at Security Gateways (Frants and Agaian, January 2026)
227 ▪ RobustMask: Certified Robustness against Adversarial Neural Ranking Attack via Randomized Masking (Liu et al., December 2025)
228 ▪ Backdoor Attacks on Prompt-Driven Video Segmentation Foundation Models (Zhang et al., December 2025)
229 ▪ Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation (Li et al., December 2025)
230 ▪ LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors (Lan and Na\"it-Abdesselam, December 2025)
231 ▪ WGLE:Backdoor-free and Multi-bit Black-box Watermarking for Graph Neural Networks (Li et al., December 2025)
232 ▪ CoTDeceptor:Adversarial Code Obfuscation Against CoT-Enhanced LLM Code Agents (Li et al., December 2025)
233 ▪ Collision-based Watermark for Detecting Backdoor Manipulation in Federated Learning (Li et al., December 2025)
234 ▪ ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures (Yang et al., December 2025)
235 ▪ One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models (Fang et al., December 2025)
236 ▪ Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA Models (Chen et al., December 2025)
237 ▪ Adversarial Robustness of Vision in Open Foundation Models (Fox, Buchanan, and Papadopoulos, December 2025)
238 ▪ Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier? (Lu et al., December 2025)
239 ▪ PHANTOM: Progressive High-fidelity Adversarial Network for Threat Object Modeling (Al-Karaki, Khan, and Athamneh, December 2025)
240 ▪ Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs (Cui et al., December 2025)
241 ▪ CIS-BA: Continuous Interaction Space Based Backdoor Attack for Object Detection in the Real-World (Zhao et al., December 2025)
242 ▪ Backdoors in DRL: Four Environments Focusing on In-distribution Triggers (Ashcraft et al., December 2025)
243 ▪ Trojan Cleansing with Neural Collapse (Gu et al., December 2025)
244 ▪ We Can Always Catch You: Detecting Adversarial Patched Objects WITH or WITHOUT Signature (Li et al., December 2025)
245 ▪ Bi-Erasing: A Bidirectional Framework for Concept Removal in Diffusion Models (Chen, Wang, and Li, December 2025)
246 ▪ Less Is More: Sparse and Cooperative Perturbation for Point Cloud Attacks (Tang et al., December 2025)
247 ▪ Adversarial Attacks Against Deep Learning-Based Radio Frequency Fingerprint Identification (Ma et al., December 2025)
248 ▪ Authority Backdoor: A Certifiable Backdoor Mechanism for Authoring DNNs (Yang et al., December 2025)
249 ▪ Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs (Betley et al., December 2025)
250 ▪ ByteShield: Adversarially Robust End-to-End Malware Detection through Byte Masking (Gibert and Many\`a, December 2025)
251 ▪ FBA$^2$D: Frequency-based Black-box Attack for AI-generated Image Detection (Chen et al., December 2025)
252 ▪ AdLift: Lifting Adversarial Perturbations to Safeguard 3D Gaussian Splatting Assets Against Instruction-Driven Editing (Hong et al., December 2025)
253 ▪ Quantization Blindspots: How Model Compression Breaks Backdoor Defenses (Pandey and Ye, December 2025)
254 ▪ Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models (Zhao et al., December 2025)
255 ▪ Edge-Only Universal Adversarial Attacks in Distributed Learning (Rossolini et al., December 2025)
256 ▪ UltraClean: A Simple Framework to Train Robust Neural Networks against Backdoor Attacks (Zhao and Lao, December 2025)
257 ▪ SafeGenes: Evaluating the Adversarial Robustness of Genomic Foundation Models (Zhan, Barbour, and Moore, December 2025)
258 ▪ Bones of Contention: Exploring Query-Efficient Attacks against Skeleton Recognition Systems (Cao et al., December 2025)
259 ▪ Superpixel Attack: Enhancing Black-box Adversarial Attack with Image-driven Division Areas (Oe et al., December 2025)
260 ▪ Concept-Guided Backdoor Attack on Vision Language Models (Shen et al., December 2025)
261 ▪ TrojanLoC: LLM-based Framework for RTL Trojan Localization (Xiao et al., December 2025)
262 ▪ CacheTrap: Injecting Trojans in LLMs without Leaving any Traces in Inputs or Weights (Nahian et al., December 2025)
263 ▪ Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning (Zhang et al., December 2025)
264 ▪ On the Effectiveness of Adversarial Training on Malware Classifiers (Bostani et al., November 2025)
265 ▪ Special-Character Adversarial Attacks on Open-Source Language Model (Sarabamoun, November 2025)
266 ▪ Illuminating the Black Box: Real-Time Monitoring of Backdoor Unlearning in CNNs via Explainable AI (Hoang, November 2025)
267 ▪ CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion (Xia et al., November 2025)
268 ▪ BackFed: An Efficient & Standardized Benchmark Suite for Backdoor Attacks in Federated Learning (Dao et al., November 2025)
269 ▪ Cross-LLM Generalization of Behavioral Backdoor Detection in AI Agent Supply Chains (Sanna, November 2025)
270 ▪ IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding (Li et al., November 2025)
271 ▪ DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs (Guo et al., November 2025)
272 ▪ A Novel and Practical Universal Adversarial Perturbations against Deep Reinforcement Learning based Intrusion Detection Systems (Zhang et al., November 2025)
273 ▪ Towards Effective, Stealthy, and Persistent Backdoor Attacks Targeting Graph Foundation Models (Luo et al., November 2025)
274 ▪ Evaluating Adversarial Vulnerabilities in Modern Large Language Models (Perel, November 2025)
275 ▪ Defending the Edge: Representative-Attention Defense against Backdoor Attacks in Federated Learning (Obioma, Sun, and Mustafa, November 2025)
276 ▪ Steering in the Shadows: Causal Amplification for Activation Space Attacks in Large Language Models (Xu et al., November 2025)
277 ▪ AutoBackdoor: Automating Backdoor Attacks via LLM Agents (Li et al., November 2025)
278 ▪ Transferable Dual-Domain Feature Importance Attack against AI-Generated Image Detector (Zhu et al., November 2025)
279 ▪ Attacking Autonomous Driving Agents with Adversarial Machine Learning: A Holistic Evaluation with the CARLA Leaderboard (Wong et al., November 2025)
280 ▪ High Dimensional Distributed Gradient Descent with Arbitrary Number of Byzantine Attackers (Liu et al., November 2025)
281 ▪ BadTime: An Effective Backdoor Attack on Multivariate Long-Term Time Series Forecasting (Xiang et al., November 2025)
282 ▪ TooBadRL: Trigger Optimization to Boost Effectiveness of Backdoor Attacks on Deep Reinforcement Learning (Zhang et al., November 2025)
283 ▪ Watch Out for the Lifespan: Evaluating Backdoor Attacks Against Federated Model Adaptation (Vuillod, Moellic, and Dutertre, November 2025)
284 ▪ Steganographic Backdoor Attacks in NLP: Ultra-Low Poisoning and Defense Evasion (Xue et al., November 2025)
285 ▪ Dynamic Black-box Backdoor Attacks on IoT Sensory Data (Chathoth and Lee, November 2025)
286 ▪ Uncovering and Aligning Anomalous Attention Heads to Defend Against NLP Backdoor Attacks (Jin et al., November 2025)
287 ▪ DiffProtect: Generate Adversarial Examples with Diffusion Models for Facial Privacy Protection (Liu et al., November 2025)
288 ▪ NeuroStrike: Neuron-Level Attacks on Aligned LLMs (Wu et al., November 2025)
289 ▪ Isolate Trigger: Detecting and Eliminating Adaptive Backdoor Attacks (Sun et al., November 2025)
290 ▪ Backdooring CLIP through Concept Confusion (Hu et al., November 2025)
291 ▪ The 'Sure' Trap: Multi-Scale Poisoning Analysis of Stealthy Compliance-Only Backdoors in Fine-Tuned Large Language Models (Tan, Huang, and Li, November 2025)
292 ▪ Calibrated Adversarial Sampling: Multi-Armed Bandit-Guided Generalization Against Unforeseen Attacks (Wang et al., November 2025)
293 ▪ MedFedPure: A Medical Federated Framework with MAE-based Detection and Diffusion Purification for Inference-Time Attacks (Karami et al., November 2025)
294 ▪ Enhancing All-to-X Backdoor Attacks with Optimized Target Class Mapping (Wang et al., November 2025)
295 ▪ SoK: The Last Line of Defense: On Backdoor Defense Evaluation (Abad et al., November 2025)
296 ▪ Efficient Adversarial Malware Defense via Trust-Based Raw Override and Confidence-Adaptive Bit-Depth Reduction (Chaudhary and Doppalpudi, November 2025)
297 ▪ AttackVLA: Benchmarking Adversarial and Backdoor Attacks on Vision-Language-Action Models (Li et al., November 2025)
298 ▪ BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning (Wang and Zhao, November 2025)
299 ▪ Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model Merging (Li et al., November 2025)
300 ▪ Transferable Hypergraph Attack via Injecting Nodes into Pivotal Hyperedges (He et al., November 2025)
301 ▪ Are Neural Networks Collision Resistant? (Benedetti et al., November 2025)
302 ▪ Trapped by Their Own Light: Deployable and Stealth Retroreflective Patch Attacks on Traffic Sign Recognition Systems (Tsuruoka et al., November 2025)
303 ▪ Improving Adversarial Transferability with Neighbourhood Gradient Information (Guo et al., November 2025)
304 ▪ Robust Backdoor Removal by Reconstructing Trigger-Activated Changes in Latent Representation (Iwahana et al., November 2025)
305 ▪ Unveiling Hidden Threats: Using Fractal Triggers to Boost Stealthiness of Distributed Backdoor Attacks in Federated Learning (Wang, Shen, and Lam, November 2025)
306 ▪ Breaking the Stealth-Potency Trade-off in Clean-Image Backdoors with Generative Trigger Optimization (Xu et al., November 2025)
307 ▪ E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis (Zhang et al., November 2025)
308 ▪ From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge (Lu et al., November 2025)
309 ▪ CatBack: Universal Backdoor Attacks on Tabular Data via Categorical Encoding (Tajalli, Koffas, and Picek, November 2025)
310 ▪ Feature compression is the root cause of adversarial fragility in neural network classifiers (Gao et al., November 2025)
311 ▪ Lorica: A Synergistic Fine-Tuning Framework for Advancing Personalized Adversarial Robustness (Qi et al., November 2025)
312 ▪ ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training (Yao et al., November 2025)
313 ▪ ShadowLogic: Backdoors in Any Whitebox LLM (Schulz, Kawasaki, and Ring, November 2025)
314 ▪ Rethinking Robust Adversarial Concept Erasure in Diffusion Models (Yin, Tian, and Zhang, November 2025)
315 ▪ GSE: Group-wise Sparse and Explainable Adversarial Attacks (Sadiku, Wagner, and Pokutta, October 2025)
316 ▪ SSCL-BW: Sample-Specific Clean-Label Backdoor Watermarking for Dataset Ownership Verification (Wang et al., October 2025)
317 ▪ Hammering the Diagnosis: Rowhammer-Induced Stealthy Trojan Attacks on ViT-Based Medical Imaging (Latibari et al., October 2025)
318 ▪ Your Compiler is Backdooring Your Model: Understanding and Exploiting Compilation Inconsistency Vulnerabilities in Deep Learning Compilers (Chen et al., October 2025)
319 ▪ ME: Trigger Element Combination Backdoor Attack on Copyright Infringement (Yang et al., October 2025)
320 ▪ Breaking Agent Backbones: Evaluating the Security of Backbone LLMs in AI Agents (Bazinska et al., October 2025)
321 ▪ Cross-Paradigm Graph Backdoor Attacks with Promptable Subgraph Triggers (Liu et al., October 2025)
322 ▪ Boosting Adversarial Transferability with Spatial Adversarial Alignment (Chen et al., October 2025)
323 ▪ FPT-Noise: Dynamic Scene-Aware Counterattack for Test-Time Adversarial Defense in Vision-Language Models (Deng et al., October 2025)
324 ▪ NeuPerm: Disrupting Malware Hidden in Neural Network Parameters by Leveraging Permutation Symmetry (Gilkarov and Dubin, October 2025)
325 ▪ Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection (Yi et al., October 2025)
326 ▪ Exploring the Effect of DNN Depth on Adversarial Attacks in Network Intrusion Detection Systems (ElShehaby and Matrawy, October 2025)
327 ▪ Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation (Muri et al., October 2025)
328 ▪ Enhancing Adversarial Transferability with Adversarial Weight Tuning (Chen et al., October 2025)
329 ▪ Colliding with Adversaries at ECML-PKDD 2025 Adversarial Attack Competition 1st Prize Solution (Stefanopoulos and Voskou, October 2025)
330 ▪ Can Transformer Memory Be Corrupted? Investigating Cache-Side Vulnerabilities in Large Language Models (Hossain et al., October 2025)
331 ▪ UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks (Phute et al., October 2025)
332 ▪ A Versatile Framework for Designing Group-Sparse Adversarial Attacks (Heshmati et al., October 2025)
333 ▪ Patronus: Safeguarding Text-to-Image Models against White-Box Adversaries (Li et al., October 2025)
334 ▪ PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models (Seddik et al., October 2025)
335 ▪ Stealthy Dual-Trigger Backdoors: Attacking Prompt Tuning in LM-Empowered Graph Foundation Models (Xue et al., October 2025)
336 ▪ An Information Asymmetry Game for Trigger-based DNN Model Watermarking (Huang et al., October 2025)
337 ▪ Improving Transferability of Adversarial Examples via Bayesian Attacks (Li et al., October 2025)
338 ▪ Who Speaks for the Trigger? Dynamic Expert Routing in Backdoored Mixture-of-Experts Transformers (Zhao et al., October 2025)
339 ▪ Injection, Attack and Erasure: Revocable Backdoor Attacks via Machine Unlearning (Song et al., October 2025)
340 ▪ How Vulnerable Is My Learned Policy? Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies (Kalra et al., October 2025)
341 ▪ Temporal Logic-Based Multi-Vehicle Backdoor Attacks against Offline RL Agents in End-to-end Autonomous Driving (Chen et al., October 2025)
342 ▪ DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent (Zhu et al., October 2025)
343 ▪ Adversarial Attacks on Downstream Weather Forecasting Models: Application to Tropical Cyclone Trajectory Prediction (Deng et al., October 2025)
344 ▪ Collaborative Shadows: Distributed Backdoor Attacks in LLM-Based Multi-Agent Systems (Zhu et al., October 2025)
345 ▪ TabVLA: Targeted Backdoor Attacks on Vision-Language-Action Models (Xu et al., October 2025)
346 ▪ SASER: Stego attacks on open-source LLMs (Tan et al., October 2025)
347 ▪ Rounding-Guided Backdoor Injection in Deep Learning Model Quantization (Chen et al., October 2025)
348 ▪ Goal-oriented Backdoor Attack against Vision-Language-Action Models via Physical Objects (Zhou et al., October 2025)
349 ▪ GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis (Dutta et al., October 2025)
350 ▪ Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning (Gloaguen et al., October 2025)
351 ▪ Backdoor Vectors: a Task Arithmetic View on Backdoor Attacks and Defenses (Pawlak et al., October 2025)
352 ▪ Fewer Weights, More Problems: A Practical Attack on LLM Pruning (Egashira et al., October 2025)
353 ▪ Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs (Hu et al., October 2025)
354 ▪ Attacking the Spike: On the Transferability and Security of Spiking Neural Networks to Adversarial Examples (Xu et al., October 2025)
355 ▪ Unsupervised Backdoor Detection and Mitigation for Spiking Neural Networks (Li et al., October 2025)
356 ▪ A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and Triggers (Wu et al., October 2025)
357 ▪ From Poisoned to Aware: Fostering Backdoor Self-Awareness in LLMs (Shen et al., October 2025)
358 ▪ Universal Adversarial Perturbation Attacks On Modern Behavior Cloning Policies (Kalra et al., October 2025)
359 ▪ Cooperative Decentralized Backdoor Attacks on Vertical Federated Learning (Lee et al., October 2025)
360 ▪ Backdoors in Code Summarizers: How Bad Is It? (Wang et al., October 2025)
361 ▪ Machine Unlearning Meets Adversarial Robustness via Constrained Interventions on LLMs (Rezkellah and Dakhmouche, October 2025)
362 ▪ Adversarial training with restricted data manipulation (Benfield et al., October 2025)
363 ▪ NatGVD: Natural Adversarial Example Attack towards Graph-based Vulnerability Detection (Rath et al., October 2025)
364 ▪ P2P: A Poison-to-Poison Remedy for Reliable Backdoor Defense in LLMs (Zhao et al., October 2025)
365 ▪ Backdoor-Powered Prompt Injection Attacks Nullify Defense Methods (Chen et al., October 2025)
366 ▪ Attack logics, not outputs: Towards efficient robustification of deep neural networks by falsifying concept-based properties (Dankworth and Schwalbe, October 2025)
367 ▪ Mutual Information Guided Backdoor Mitigation for Pre-trained Encoders (Han et al., October 2025)
368 ▪ A Statistical Method for Attack-Agnostic Adversarial Attack Detection with Compressive Sensing Comparison (Wimalasuriya and Tragoudas, October 2025)
369 ▪ Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks (Zhang et al., October 2025)
370 ▪ Mirage Fools the Ear, Mute Hides the Truth: Precise Targeted Adversarial Attacks on Polyphonic Sound Event Detection Systems (Su et al., October 2025)
371 ▪ Towards Imperceptible Adversarial Defense: A Gradient-Driven Shield against Facial Manipulations (Li et al., October 2025)
372 ▪ Evaluating the Robustness of a Production Malware Detection System to Transferable Adversarial Attacks (Nasr et al., October 2025)
373 ▪ Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation (Chaudhari et al., October 2025)
374 ▪ A Backdoor-based Explainable AI Benchmark for High Fidelity Evaluation of Attributions (Yang et al., October 2025)
375 ▪ Backdoor Attacks Against Speech Language Models (Fortier et al., October 2025)
376 ▪ Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness (Takahashi et al., October 2025)
377 ▪ Has the Two-Decade-Old Prophecy Come True? Artificial Bad Intelligence Triggered by Merely a Single-Bit Flip in Large Language Models (Yan et al., October 2025)
378 ▪ Cut the Deadwood Out: Backdoor Purification via Guided Module Substitution (Tong et al., October 2025)
379 ▪ Backdoor Attribution: Elucidating and Controlling Backdoor in Language Models (Yu et al., October 2025)
380 ▪ Stealthy Yet Effective: Distribution-Preserving Backdoor Attacks on Graph Classification (Wang et al., October 2025)
381 ▪ TrojanStego: Your Language Model Can Secretly Be A Steganographic Privacy Leaking Agent (Meier et al., September 2025)
382 ▪ Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack (Chen et al., September 2025)
383 ▪ GPM: The Gaussian Pancake Mechanism for Planting Undetectable Backdoors in Differential Privacy (Sun and He, September 2025)
384 ▪ DUP: Detection-guided Unlearning for Backdoor Purification in Language Models (Hu et al., August 2025)
385 ▪ Practical, Generalizable and Robust Backdoor Attacks on Text-to-Image Diffusion Models (Dai et al., August 2025)
386 ▪ BeDKD: Backdoor Defense based on Dynamic Knowledge Distillation and Directional Mapping Modulator (Wu et al., August 2025)
387 ▪ Backdoor Attacks on Deep Learning Face Detection (Roux et al., August 2025)
388 ▪ SDBA: A Stealthy and Long-Lasting Durable Backdoor Attack in Federated Learning (Choe et al., July 2025)
389 ▪ FedBAP: Backdoor Defense via Benign Adversarial Perturbation in Federated Learning (Yan et al., July 2025)
390 ▪ Persistent Backdoor Attacks in Continual Learning (Guo, Kumar, and Tourani, July 2025)
391 ▪ ConSeg: Contextual Backdoor Attack Against Semantic Segmentation (Abbasi et al., July 2025)
392 ▪ Gungnir: Exploiting Stylistic Features in Images for Backdoor Attacks on Diffusion Models (Pan et al., July 2025)
393 ▪ BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion Model (Lin et al., July 2025)
394 ▪ VTarbel: Targeted Label Attack with Minimal Knowledge on Detector-enhanced Vertical Federated Learning (Tan et al., July 2025)
395 ▪ BLAST: A Stealthy Backdoor Leverage Attack against Cooperative Multi-Agent Deep Reinforcement Learning based Systems (Fang et al., July 2025)
396 ▪ Invisible Textual Backdoor Attacks based on Dual-Trigger (Hou et al., July 2025)
397 ▪ An Adversarial-Driven Experimental Study on Deep Learning for RF Fingerprinting (Cao et al., July 2025)
398 ▪ Architectural Backdoors in Deep Learning: A Survey of Vulnerabilities, Detection, and Defense (Childress, Collyer, and Knapp, July 2025)
399 ▪ Non-Adaptive Adversarial Face Generation (Kim et al., July 2025)
400 ▪ Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks (Gosch et al., July 2025)
401 ▪ Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs (Sivapiromrat et al., July 2025)
402 ▪ 3S-Attack: Spatial, Spectral and Semantic Invisible Backdoor Attack Against DNN Models (Yin et al., July 2025)
403 ▪ No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms (Kazdan et al., July 2025)
404 ▪ BURN: Backdoor Unlearning via Adversarial Boundary Analysis (Su et al., July 2025)
405 ▪ Breaking PEFT Limitations: Leveraging Weak-to-Strong Knowledge Transfer for Backdoor Attacks in LLMs (Zhao et al., July 2025)
406 ▪ One Surrogate to Fool Them All: Universal, Transferable, and Targeted Adversarial Attacks with CLIP (Xu et al., July 2025)
407 ▪ Can We Trust Embodied Agents? Exploring Backdoor Attacks against Embodied LLM-based Decision-Making Systems (Jiao et al, Apr 2025)
408 ▪ How to Backdoor Consistency Models (Wang and Kantarcioglu, Feb 2025)
409 ▪ Towards Backdoor Stealthiness in Model Parameter Space (Xu et al, Jan 2025)
410 ▪ Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data (Zhang et al, Jan 2025)
411 ▪ A Backdoor Attack Scheme with Invisible Triggers Based on Model Architecture Modification (Ma et al, Jan 2025)
412 ▪ Double Landmines: Invisible Textual Backdoor Attacks based on Dual-Trigger (Hou et al, Dec 2024)
413 ▪ Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models (Li et al, Dec 2024)
414 ▪ UOR: Universal Backdoor Attacks on Pre-trained Language Models (Du et al, Dec 2024)
415 ▪ Client-Side Patching against Backdoor Attacks in Federated Learning (Molina-Coronado, Dec 2024)
416 ▪ Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models (Li et al, Dec 2024)
417 ▪ Backdoor Learning Curves: Explaining Backdoor Poisoning Beyond Influence Functions (Cinà et al, Dec 2024)
418 ▪ Proactive Adversarial Defense: Harnessing Prompt Tuning in Vision-Language Models to Detect Unseen Backdoored Images (Stein et al, Dec 2024)
419 ▪ Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP (Singh, Croce, and Hein, Dec 2024)
420 ▪ RTL-Breaker: Assessing the Security of LLMs against Backdoor Attacks on HDL Code Generation (Mankali et al, Dec 2024)
421 ▪ Gracefully Filtering Backdoor Samples for Generative Large Language Models without Retraining (Wu et al, Dec 2024)
422 ▪ LoBAM: LoRA-Based Backdoor Attack on Model Merging (Yin et al, Dec 2024)
423 ▪ Towards Clean-Label Backdoor Attacks in the Physical World (Dao et al, Nov 2024)
424 ▪ Unlearn to Relearn Backdoors: Deferred Backdoor Functionality Attacks on Deep Learning Models (Shin and Park, Nov 2024)
425 ▪ Memory Backdoor Attacks on Neural Networks (Luzon et al, Nov 2024)
426 ▪ AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection (Lu et al, Nov 2024)
427 ▪ When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations (Ge et al, Nov 2024)
428 ▪ Backdoor defense, learnability and obfuscation (Christiano et al, Nov 2024)
429 ▪ Combinational Backdoor Attack against Customized Text-to-Image Models (Jiang et al, Nov 2024)
430 ▪ Backdoor Attack Against Vision Transformers via Attention Gradient-Based Image Erosion (Guo et al, Nov 2024)
431 ▪ Backdoor Attacks against Image-to-Image Networks (Jiang et al, Nov 2024)
432 ▪ Infighting in the Dark: Multi-Labels Backdoor Attack in Federated Learning (Li et al, Nov 2024)
433 ▪ Planting Undetectable Backdoors in Machine Learning Models (Goldwasser et al, Nov 2024)
434 ▪ On the Credibility of Backdoor Attacks Against Object Detectors in the Physical World (Gia Doan et al, Oct 2024)
435 ▪ Model X-ray:Detecting Backdoored Models via Decision Boundary (Su et al, Oct 2024)
436 ▪ Dullahan: Stealthy Backdoor Attack against Without-Label-Sharing Split Learning (Pu et al, Oct 2024)
437 ▪ Mind Your Questions Towards Backdoor Attacks on Text-to-Visualization Models (Li et al, Oct 2024)
438 ▪ BadCM: Invisible Backdoor Attack Against Cross-Modal Learning (Zheng et al, Oct 2024)
439 ▪ BACKTIME: Backdoor Attacks on Multivariate Time Series Forecasting (Lin et al, Oct 2024)
440 ▪ Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context Learning (Zhao et al, Oct 2024)
441 ▪ Hidden in Plain Sound: Environmental Backdoor Poisoning Attacks on Whisper, and Mitigations (Bartolini, Stoyanov, and Giaretta, Sep 2024)
442 ▪ EmoBack: Backdoor Attacks Against Speaker Identification Using Emotional Prosody (Schoof et al, Sep 2024)
443 ▪ Context is the Key: Backdoor Attacks for In-Context Learning with Vision Transformers (Abad et al, Sep 2024)
444 ▪ Concealing Backdoor Model Updates in Federated Learning by Trigger-Optimized Data Poisoning (Zhang, Gong, and Reiter, Sep 2024)
445 ▪ TERD: A Unified Framework for Safeguarding Diffusion Models Against Backdoors (Mo et al , Sep 2024)
446 ▪ INK: Inheritable Natural Backdoor Attack Against Model Distillation (Liu et al , Sep 2024)
447 ▪ Exploiting the Vulnerability of Large Language Models via Defense-Aware Architectural Backdoor (Miah and Bi, Sep 2024)
448 ▪ Rethinking Backdoor Detection Evaluation for Language Models (Yan et al, Sep 2024)
449 ▪ Shortcuts Everywhere and Nowhere: Exploring Multi-Trigger Backdoor Attacks (Li et al, Aug 2024)
450 ▪ Transferring Backdoors between Large Language Models by Knowledge Distillation (Cheng et al, Aug 2024)
451 ▪ Revocable Backdoor for Deep Model Trading (Zu et al, Aug 2024)
452 ▪ BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning (Wu et al, Jul 2024)
453 ▪ Diff-Cleanse: Identifying and Mitigating Backdoor Attacks in Diffusion Models (Hao et al, Jul 2024)
Backdoor ML Model and Craft Adversarial Data

>

<

‍

Supply Chain Vulnerabilities and Compromise

Covers:

  • OWASP LLM 05/OWASP ML 06/MITRE ATLAS Initial Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Supply Chain Vulnerabilities and Compromise

cybersecurity_tracker - Google Drive

2 ▪ Exploiting LLM Agent Supply Chains via Payload-less Skills (Liu et al., May 2026)
3 ▪ Under the Hood of SKILL.md: Semantic Supply-chain Attacks on AI Agent Skill Registry (Saha, Faghih, and Feizi, May 2026)
4 ▪ Cryptographic Registry Provenance: Structural Defense Against Dependency Confusion in AI Package Ecosystems (McCann, May 2026)
5 ▪ Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors (Li et al., May 2026)
6 ▪ SOK: A Taxonomy of Attack Vectors and Defense Strategies for Agentic Supply Chain Runtime (Jiang et al., April 2026)
7 ▪ Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP Ecosystem (Zhao et al., April 2026)
8 ▪ Your Agent Is Mine: Measuring Malicious Intermediary Attacks on the LLM Supply Chain (Liu et al., April 2026)
9 ▪ Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems (Qu et al., April 2026)
10 ▪ How Can ChatGPT Support Human Security Testers to Help Mitigate Supply Chain Attacks? (Zhang et al., March 2026)
11 ▪ SynthChain: A Synthetic Benchmark and Forensic Analysis of Advanced and Stealthy Software Supply Chain Attacks (Tan et al., March 2026)
12 ▪ On the (In)Security of Loading Machine Learning Models (Digregorio et al., March 2026)
13 ▪ Formal Analysis and Supply Chain Security for Agentic AI Skills (Bhardwaj, March 2026)
14 ▪ LLM Scalability Risk for Agentic-AI and Model Supply Chain Security (Ahi, Agrawal, and Valizadeh, February 2026)
15 ▪ Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines (Jamshidi et al., February 2026)
16 ▪ AIBoMGen: Generating an AI Bill of Materials for Secure, Transparent, and Compliant Model Training (Vandendriessche et al., January 2026)
17 ▪ Understanding Security Risks of AI Agents' Dependency Updates (Singla et al., January 2026)
18 ▪ Securing the AI Supply Chain: What Can We Learn From Developer-Reported Security Issues and Solutions of AI Projects? (Nguyen, Le, and Babar, December 2025)
19 ▪ CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution (Zhu et al., December 2025)
20 ▪ A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution (Zhu et al., December 2025)
21 ▪ Identifying the Supply Chain of AI for Trustworthiness and Risk Management in Critical Applications (Sheh and Geappen, November 2025)
22 ▪ Supply Chain Exploitation of Secure ROS 2 Systems: A Proof-of-Concept on Autonomous Platform Compromise via Keystore Exfiltration (Sakib et al., November 2025)
23 ▪ Lexo: Eliminating Stealthy Supply-Chain Attacks via LLM-Assisted Program Regeneration (Lamprou et al., October 2025)
24 ▪ Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain (Boisvert et al., October 2025)
25 ▪ ExclaveFL: Providing Transparency to Federated Learning using Exclaves (Guo et al., August 2025)
26 ▪ VDGraph: A Graph-Theoretic Approach to Unlock Insights from SBOM and SCA Data (Xia et al., July 2025)
27 ▪ Understanding the Supply Chain and Risks of Large Language Model Applications (Ma et al., July 2025)
28 ▪ PyPitfall: Dependency Chaos and Software Supply Chain Vulnerabilities in Python (Mahon, Hou, and Yao, July 2025)
29 ▪ Security Enclave Architecture for Heterogeneous Security Primitives for Supply-Chain Attacks (Raj et al., July 2025)
30 ▪ Exploiting Leaderboards for Large-Scale Distribution of Malicious Models (Suri et al., July 2025)
31 ▪ Rugsafe: A multichain protocol for recovering from and defending against Rug Pulls (Pharr and Hussain, July 2025)
Supply Chain Vulnerabilities and Compromise

>

<

‍

Excessive Agency, Agentic Manipulation, Agentic Systems

We added 'agentic' manipulation to this subcategory.

Covers:

  • OWASP LLM 08: Excessive Agency

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Excessive Agency, Agentic Manipulation, Agentic Systems

cybersecurity_tracker - Google Drive

2 ▪ Progent: Securing AI Agents with Privilege Control (Shi et al., May 2026)
3 ▪ Do Coding Agents Understand Least-Privilege Authorization? (Yan et al., May 2026)
4 ▪ Quantitative Certification of Agentic Tool Selection (Yeon, Chaudhary, and Singh, May 2026)
5 ▪ Language-Based Agent Control (Zhou, D'Antoni, and Polikarpova, May 2026)
6 ▪ Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety (Bilal et al., May 2026)
7 ▪ No Attack Required: Semantic Fuzzing for Specification Violations in Agent Skills (Li et al., May 2026)
8 ▪ A Security Analysis of the OpenClaw AI Agent Framework (Suwansathit, Zhang, and Gu, May 2026)
9 ▪ Attacks and Mitigations for Distributed Governance of Agentic AI under Byzantine Adversaries (Laws, Oprea, and Nita-Rotaru, May 2026)
10 ▪ SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces (Jin et al., May 2026)
11 ▪ Proteus: A Self-Evolving Red Team for Agent Skill Ecosystems (Zhou, May 2026)
12 ▪ Five Attacks on x402 Agentic Payment Protocol (Li, Wang, and Wang, May 2026)
13 ▪ Safety Context Injection: Inference-Time Safety Alignment via Static Filtering and Agentic Analysis (Xu et al., May 2026)
14 ▪ FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems (Li et al., May 2026)
15 ▪ Digital Identity for Agentic Systems: Toward a Portable Authorization Standard for Autonomous Agents (Madhira, May 2026)
16 ▪ Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution (Fendley et al., May 2026)
17 ▪ Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw (Yao et al., May 2026)
18 ▪ The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck (Fan et al., May 2026)
19 ▪ Portable Agent Memory: A Protocol for Cryptographically-Verified Memory Transfer Across Heterogeneous AI Agents (Ravindran, May 2026)
20 ▪ AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents (Rassul and Rashid, May 2026)
21 ▪ The Authorization-Execution Gap Is a Major Safety and Security Problem in Open-World Agents (Wu et al., May 2026)
22 ▪ Formal Policy Enforcement for Real-World Agentic Systems (Palumbo et al., May 2026)
23 ▪ MATRA: Modeling the Attack Surface of Agentic AI Systems -- OpenClaw Case Study (hamme et al., May 2026)
24 ▪ Evaluating Tool Cloning in Agentic-AI Ecosystems (Kim et al., May 2026)
25 ▪ Agentic Fuzzing: Opportunities and Challenges (Park and Yun, May 2026)
26 ▪ Nautilus Compass: Black-box Persona Drift Detection for Production LLM Agents (Wang, May 2026)
27 ▪ Security Risks in Tool-Enabled AI Agents: A Systematic Analysis of Privileged Execution Environments (Goel, May 2026)
28 ▪ When Child Inherits: Modeling and Exploiting Subagent Spawn in Multi-Agent Networks (Cai, Zhang, and Hei, May 2026)
29 ▪ MAGIQ: A Post-Quantum Multi-Agentic AI Governance System with Provable Security (Avizeh et al., May 2026)
30 ▪ Demystifying and Detecting Agentic Workflow Injection Vulnerabilities in GitHub Actions (Wang et al., May 2026)
31 ▪ Language Models Can Autonomously Hack and Self-Replicate (Air et al., May 2026)
32 ▪ Agentic AI and the Industrialization of Cyber Offense: Forecast, Consequences, and Defensive Priorities for Enterprises and the Mittelstand (Koch, May 2026)
33 ▪ AgentDyn: Are Your Agent Security Defenses Deployable in Real-World Dynamic Environments? (Li et al., May 2026)
34 ▪ Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches (David and Gervais, May 2026)
35 ▪ Constraining Host-Level Abuse in Self-Hosted Computer-Use Agents via TEE-Backed Isolation (Lu et al., May 2026)
36 ▪ ClawGuard: Out-of-Band Detection of LLM Agent Workflow Hijacking via EM Side Channel (Gan et al., May 2026)
37 ▪ SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills (Wu et al., May 2026)
38 ▪ WAAA! Web Adversaries Against Agentic Browsers (Datta et al., May 2026)
39 ▪ Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use (Arceo and Narsing, May 2026)
40 ▪ AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use (Yang, May 2026)
41 ▪ Agentic Vulnerability Reasoning on Windows COM Binaries (Lee, Kim, and Zhang, May 2026)
42 ▪ Redefining AI Red Teaming in the Agentic Era: From Weeks to Hours (Dheekonda, Pearce, and Landers, May 2026)
43 ▪ Stable Agentic Control: Tool-Mediated LLM Architecture for Autonomous Cyber Defense (Prinos et al., May 2026)
44 ▪ When Agents Handle Secrets: A Survey of Confidential Computing for Agentic AI (Forough, Kogias, and Haddadi, May 2026)
45 ▪ Towards Agentic Runtime Healing (Sun et al., May 2026)
46 ▪ Beyond Crash: Hijacking Your Autonomous Vehicle for Fun and Profit (Sun et al., May 2026)
47 ▪ When Alignment Isn't Enough: Response-Path Attacks on LLM Agents (Luo et al., May 2026)
48 ▪ Architectural Obsolescence of Unhardened Agentic-AI Runtimes (Metere, May 2026)
49 ▪ AgenticVM: Agentic AI for Adaptive Software Vulnerability Management (Arifin et al., May 2026)
50 ▪ Self-Adaptive Multi-Agent LLM-Based Security Pattern Selection for IoT Systems (Jamshidi et al., May 2026)
51 ▪ Alignment Contracts for Agentic Security Systems (David, Guarnieri, and Gervais, May 2026)
52 ▪ Ambient Persuasion in a Deployed AI Agent: Unauthorized Escalation Following Routine Non-Adversarial Content Exposure (Cuadros and Maiga, May 2026)
53 ▪ Chronology of Multi-Agent Interactions for Provenance of Evolving Information (Chang and Echizen, May 2026)
54 ▪ From surveillance to signalling: escalation channels as environmental controls for agentic AI (Gomez, May 2026)
55 ▪ From Prompt to Physical Actuation: Holistic Threat Modeling of LLM-Enabled Robotic Systems (Nagaraja, Bahsi, and Cunha, May 2026)
56 ▪ Agent Name Service (ANS): A Proof-of-Concept Trust Layer for Secure AI Agent Discovery, Identity, and Governance in Kubernetes (Mittal and Cruz, May 2026)
57 ▪ A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron? (Chen et al., April 2026)
58 ▪ Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents (Witt et al., April 2026)
59 ▪ SecMate: Multi-Agent Adaptive Cybersecurity Troubleshooting with Tri-Context Personalization (Meidan et al., April 2026)
60 ▪ Towards Agentic Investigation of Security Alerts (Eilertsen, Mavroeidis, and Grov, April 2026)
61 ▪ From CRUD to Autonomous Agents: Formal Validation and Zero-Trust Security for Semantic Gateways in AI-Native Enterprise Systems (Peyrano, April 2026)
62 ▪ AgentDID: Trustless Identity Authentication for AI Agents (Xu et al., April 2026)
63 ▪ Structured Security Auditing and Robustness Enhancement for Untrusted Agent Skills (Lv et al., April 2026)
64 ▪ SUDP: Secret-Use Delegation Protocol for Agentic Systems (Yu, Geng, and Knottenbelt, April 2026)
65 ▪ Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture (Xiang, April 2026)
66 ▪ AI Identity: Standards, Gaps, and Research Directions for AI Agents (Otsuka, Toyoda, and Leung, April 2026)
67 ▪ AgentWard: A Lifecycle Security Architecture for Autonomous AI Agents (Zhang et al., April 2026)
68 ▪ When the Agent Is the Adversary: Architectural Requirements for Agentic AI Containment After the April 2026 Frontier Model Escape (Mitchell, April 2026)
69 ▪ Ghost in the Agent: Redefining Information Flow Tracking for LLM Agents (Cai et al., April 2026)
70 ▪ From Stateless Queries to Autonomous Actions: A Layered Security Framework for Agentic AI Systems (Chu, April 2026)
71 ▪ AgentBound: Securing Execution Boundaries of AI Agents (B\"uhler et al., April 2026)
72 ▪ Automation-Exploit: A Multi-Agent LLM Framework for Adaptive Offensive Security with Digital Twin-Based Risk-Mitigated Exploitation (Andreucci and Castiglione, April 2026)
73 ▪ Sovereign Agentic Loops: Decoupling AI Reasoning from Execution in Real-World Systems (He and Yu, April 2026)
74 ▪ Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms (Azarafrooz, April 2026)
75 ▪ Breaking MCP with Function Hijacking Attacks: Novel Threats for Function Calling and Agentic Models (Belkhiter et al., April 2026)
76 ▪ AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation (Roy and Singh, April 2026)
77 ▪ Whispers in the Machine: Confidentiality in Agentic Systems (Evertz et al., April 2026)
78 ▪ ClawCoin: An Agentic AI-Native Cryptocurrency for Decentralized Agent Economies (Li et al., April 2026)
79 ▪ HadAgent: Harness-Aware Decentralized Agentic AI Serving with Proof-of-Inference Blockchain Consensus (Jimenez et al., April 2026)
80 ▪ An AI Agent Execution Environment to Safeguard User Data (Stanley et al., April 2026)
81 ▪ Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps (Chona, Kozlov, and Kumar, April 2026)
82 ▪ Towards Optimal Agentic Architectures for Offensive Security Tasks (David and Gervais, April 2026)
83 ▪ Owner-Harm: A Missing Threat Model for AI Agent Safety (Zhang and Jiang, April 2026)
84 ▪ From Craft to Kernel: A Governance-First Execution Architecture and Semantic ISA for Agentic Computers (Wen et al., April 2026)
85 ▪ Evaluating Privilege Usage of Agents with Real-World Tools (Zhang et al., April 2026)
86 ▪ From Admission to Invariants: Measuring Deviation in Delegated Agent Systems (Fernandez, April 2026)
87 ▪ Atomic Decision Boundaries: A Structural Requirement for Guaranteeing Execution-Time Admissibility in Autonomous Systems (Fernandez, April 2026)
88 ▪ AgenTEE: Confidential LLM Agent Execution on Edge Devices (Abdollahi et al., April 2026)
89 ▪ CapSeal: Capability-Sealed Secret Mediation for Secure Agent Execution (Jin, Guo, and Cheung, April 2026)
90 ▪ A Survey on the Security of Long-Term Memory in LLM Agents: Toward Mnemonic Sovereignty (Lin, Li, and Chen, April 2026)
91 ▪ CAMP: Cumulative Agentic Masking and Pruning for Privacy Protection in Multi-Turn LLM Conversations (Panjwani, April 2026)
92 ▪ HarmfulSkillBench: How Do Harmful Skills Weaponize Your Agents? (Jiang et al., April 2026)
93 ▪ An Agentic Workflow for Detecting Personally Identifiable Information in Crash Narratives (Ma et al., April 2026)
94 ▪ SoK: Security of Autonomous LLM Agents in Agentic Commerce (Mao et al., April 2026)
95 ▪ CBCL: Safe Self-Extending Agent Communication (O'Connor, April 2026)
96 ▪ Challenges and Future Directions in Agentic Reverse Engineering Systems (Radey, West, and Fawaz, April 2026)
97 ▪ Can Agents Secure Hardware? Evaluating Agentic LLM-Driven Obfuscation for IP Protection (Ghimire et al., April 2026)
98 ▪ To trust or not to trust: Attention-based Trust Management for LLM Multi-Agent Systems (He et al., April 2026)
99 ▪ Policy-Invisible Violations in LLM-Based Agents (Wu and Gong, April 2026)
100 ▪ Parallax: Why AI Agents That Think Must Never Act (Fokou, April 2026)
101 ▪ Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents (Sidik and Rokach, April 2026)
102 ▪ Beyond RAG for Cyber Threat Intelligence: A Systematic Evaluation of Graph-Based and Agentic Retrieval (Hamzic et al., April 2026)
103 ▪ Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering (Stantchev, April 2026)
104 ▪ The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems (Zhang et al., April 2026)
105 ▪ The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use Agents (Ding et al., April 2026)
106 ▪ Self-Sovereign Agent (Qu et al., April 2026)
107 ▪ VCAO: Verifier-Centered Agentic Orchestration for Strategic OS Vulnerability Discovery (Mishra, April 2026)
108 ▪ ACF: A Collaborative Framework for Agent Covert Communication under Cognitive Asymmetry (Wu et al., April 2026)
109 ▪ ARuleCon: Agentic Security Rule Conversion (Xu et al., April 2026)
110 ▪ SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills (Hou and Yang, April 2026)
111 ▪ ClawLess: A Security Model of AI Agents (Lu et al., April 2026)
112 ▪ ZitPit: Consumer-Side Admission Control for Agentic Software Intake (Taylor et al., April 2026)
113 ▪ Foundations for Agentic AI Investigations from the Forensic Analysis of OpenClaw (Gruber and Hilgert, April 2026)
114 ▪ LanG -- A Governance-Aware Agentic AI Platform for Unified Security Operations (Abdennebi et al., April 2026)
115 ▪ Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents (Zhang et al., April 2026)
116 ▪ Your Agent, Their Asset: A Real-World Safety Analysis of OpenClaw (Wang et al., April 2026)
117 ▪ Undetectable Conversations Between AI Agents via Pseudorandom Noise-Resilient Key Exchange (Vaikuntanathan and Zamir, April 2026)
118 ▪ HDP: A Lightweight Cryptographic Protocol for Human Delegation Provenance in Agentic AI Systems (Dalugoda, April 2026)
119 ▪ Governance-Constrained Agentic AI: Blockchain-Enforced Human Oversight for Safety-Critical Wildfire Monitoring (Akarma et al., April 2026)
120 ▪ AutoVerifier: An Agentic Automated Verification Framework Using Large Language Models (Du et al., April 2026)
121 ▪ Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis (Li et al., April 2026)
122 ▪ SentinelAgent: Intent-Verified Delegation Chains for Securing Federal Multi-Agent AI Systems (Patil, April 2026)
123 ▪ Type-Checked Compliance: Deterministic Guardrails for Agentic Financial Systems Using Lean 4 Theorem Proving (Rashie and Rashi, April 2026)
124 ▪ APEX: Agent Payment Execution with Policy for Autonomous Agent API Access (Uddin et al., April 2026)
125 ▪ Multi-Agent LLM Governance for Safe Two-Timescale Reinforcement Learning in SDN-IoT Defense (Jamshidi et al., April 2026)
126 ▪ AutoMIA: Improved Baselines for Membership Inference Attack via Agentic Self-Exploration (Liu et al., April 2026)
127 ▪ Do Phone-Use Agents Respect Your Privacy? (Tang et al., April 2026)
128 ▪ SafeClaw-R: Towards Safe and Secure Multi-Agent Personal Assistants (Wang et al., April 2026)
129 ▪ Multi-Agent Actor-Critics in Autonomous Cyber Defense (Wang and Dechene, March 2026)
130 ▪ "What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents (Peng, March 2026)
131 ▪ Evaluating Privilege Usage of Agents on Real-World Tools (Zhang et al., March 2026)
132 ▪ Red-MIRROR: Agentic LLM-based Autonomous Penetration Testing with Reflective Verification and Knowledge-augmented Interaction (Khang et al., March 2026)
133 ▪ Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization (Kong et al., March 2026)
134 ▪ Clawed and Dangerous: Can We Trust Open Agentic Systems? (Chen et al., March 2026)
135 ▪ AIP: Agent Identity Protocol for Verifiable Delegation Across MCP and A2A (Prakash, March 2026)
136 ▪ Infrastructure for Valuable, Tradable, and Verifiable Agent Memory (Li et al., March 2026)
137 ▪ AgentRFC: Security Design Principles and Conformance Testing for Agent Protocols (Zheng and Zhang, March 2026)
138 ▪ SoK: The Attack Surface of Agentic AI -- Tools, and Autonomy (Dehghantanha and Homayoun, March 2026)
139 ▪ Agent-Sentry: Bounding LLM Agents via Execution Provenance (Sequeira et al., March 2026)
140 ▪ STRIATUM-CTF: A Protocol-Driven Agentic Framework for General-Purpose CTF Solving (Hugglestone et al., March 2026)
141 ▪ When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents (Lu et al., March 2026)
142 ▪ Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions (Xiang et al., March 2026)
143 ▪ Before the Tool Call: Deterministic Pre-Action Authorization for Autonomous AI Agents (Uchibeke, March 2026)
144 ▪ AC4A: Access Control for Agents (Sharma and Grossman, March 2026)
145 ▪ MANA: Towards Efficient Mobile Ad Detection via Multimodal Agentic UI Navigation (Zhao et al., March 2026)
146 ▪ Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning (Zhang et al., March 2026)
147 ▪ An Agentic Multi-Agent Architecture for Cybersecurity Risk Management (Gupta et al., March 2026)
148 ▪ The Verifier Tax: Horizon Dependent Safety Success Tradeoffs in Tool Using LLM Agents (Sah et al., March 2026)
149 ▪ A402: Binding Cryptocurrency Payments to Service Execution for Agentic Commerce (Li et al., March 2026)
150 ▪ Access Controlled Website Interaction for Agentic AI with Delegated Critical Tasks (Kim and Kim, March 2026)
151 ▪ Security, privacy, and agentic AI in a regulatory view: From definitions and distinctions to provisions and reflections (Zhang and Maharjan, March 2026)
152 ▪ Agent Control Protocol: Admission Control for Agent Actions (Fernandez, March 2026)
153 ▪ Capability-Priced Micro-Markets: A Micro-Economic Framework for the Agentic Web over HTTP 402 (Huang et al., March 2026)
154 ▪ Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare (Maiti, March 2026)
155 ▪ LAAF: Logic-layer Automated Attack Framework A Systematic Red-Teaming Methodology for LPCI Vulnerabilities in Agentic Large Language Model Systems (Atta et al., March 2026)
156 ▪ PAuth - Precise Task-Scoped Authorization For Agents (Sharma et al., March 2026)
157 ▪ Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback (Jiralerspong, Kondrup, and Bengio, March 2026)
158 ▪ ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems (Zhang et al., March 2026)
159 ▪ ILION: Deterministic Pre-Execution Safety Gates for Agentic AI Systems (Chitan, March 2026)
160 ▪ From Storage to Steering: Memory Control Flow Attacks on LLM Agents (Xu et al., March 2026)
161 ▪ Towards Agentic Honeynet Configuration (Mirra et al., March 2026)
162 ▪ Sovereign-OS: A Charter-Governed Operating System for Autonomous AI Agents with Verifiable Fiscal Discipline (Yuan et al., March 2026)
163 ▪ Defensible Design for OpenClaw: Securing Autonomous Tool-Invoking Agents (Li, Li, and Li, March 2026)
164 ▪ Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats (Deng et al., March 2026)
165 ▪ The Attack and Defense Landscape of Agentic AI: A Comprehensive Survey (Kim et al., March 2026)
166 ▪ Re-Evaluating EVMBench: Are AI Agents Ready for Smart Contract Security? (Peng, Wu, and Zhou, March 2026)
167 ▪ Execution Is the New Attack Surface: Survivability-Aware Agentic Crypto Trading with OpenClaw-Style Local Executors (Borjigin et al., March 2026)
168 ▪ SBOMs into Agentic AIBOMs: Schema Extensions, Agentic Orchestration, and Reproducibility Evaluation (Radanliev et al., March 2026)
169 ▪ ProvAgent: Threat Detection Based on Identity-Behavior Binding and Multi-Agent Collaborative Attack Investigation (Yan et al., March 2026)
170 ▪ AgenticCyOps: Securing Multi-Agentic AI Integration in Enterprise Cyber Operations (Mitra et al., March 2026)
171 ▪ Security Considerations for Multi-agent Systems (Nguyen, Ndebugre, and Arremsetty, March 2026)
172 ▪ SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directions (Mishra et al., March 2026)
173 ▪ From Thinker to Society: Security in Hierarchical Autonomy Evolution of AI Agents (Zhang et al., March 2026)
174 ▪ Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practice (Ge, March 2026)
175 ▪ Information-Theoretic Privacy Control for Sequential Multi-Agent LLM Systems (Asif and Amiri, March 2026)
176 ▪ Evolving Deception: When Agents Evolve, Deception Wins (Ying et al., March 2026)
177 ▪ CyberSleuth: Autonomous Blue-Team LLM Agent for Web Attack Forensics (Fumero et al., March 2026)
178 ▪ AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows (Ngong et al., March 2026)
179 ▪ Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization (Mumcu and Yilmaz, March 2026)
180 ▪ On the Suitability of LLM-Driven Agents for Dark Pattern Audits (Sun, Vekaria, and Nithyanand, March 2026)
181 ▪ From Secure Agentic AI to Secure Agentic Web: Challenges, Threats, and Future Directions (Deng, Gui, and Zhang, March 2026)
182 ▪ Atomicity for Agents: Exposing, Exploiting, and Mitigating TOCTOU Vulnerabilities in Browser-Use Agents (Jiang et al., March 2026)
183 ▪ LiaisonAgent: An Multi-Agent Framework for Autonomous Risk Investigation and Governance (Tang, Qing, and Chen, March 2026)
184 ▪ Adversarial Intent is a Latent Variable: Stateful Trust Inference for Securing Multimodal Agentic RAG (Singh et al., February 2026)
185 ▪ "Are You Sure?": An Empirical Study of Human Perception Vulnerability in LLM-Driven Agentic Systems (Li et al., February 2026)
186 ▪ SoK: Agentic Skills -- Beyond Tool Use in LLM Agents (Jiang et al., February 2026)
187 ▪ The LLMbda Calculus: AI Agents, Conversations, and Information Flow (Garby, Gordon, and Sands, February 2026)
188 ▪ Agentic AI as a Cybersecurity Attack Surface: Threats, Exploits, and Defenses in Runtime Supply Chains (Jiang et al., February 2026)
189 ▪ Security Risks of AI Agents Hiring Humans: An Empirical Marketplace Study (Mehta, February 2026)
190 ▪ OpenSage: Self-programming Agent Generation Engine (Li et al., February 2026)
191 ▪ Policy Compiler for Secure Agentic Systems (Palumbo et al., February 2026)
192 ▪ Intellicise Wireless Networks Meet Agentic AI: A Security and Privacy Perspective (Meng et al., February 2026)
193 ▪ Overthinking Loops in Agents: A Structural Risk via MCP Tools (Lee et al., February 2026)
194 ▪ AXE: An Agentic eXploit Engine for Confirming Zero-Day Vulnerability Reports (Sajadi et al., February 2026)
195 ▪ In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach (Gao, Hammar, and Li, February 2026)
196 ▪ Agentic AI for Cybersecurity: A Meta-Cognitive Architecture for Governable Autonomy (Kojukhov and Bovshover, February 2026)
197 ▪ Optimizing Agent Planning for Security and Autonomy (Kolluri et al., February 2026)
198 ▪ Agentic Knowledge Distillation: Autonomous Training of Small Language Models for SMS Threat Detection (ElZemity et al., February 2026)
199 ▪ Authenticated Workflows: A Systems Approach to Protecting Agentic AI (Rajagopalan and Rao, February 2026)
200 ▪ Trustworthy Agentic AI Requires Deterministic Architectural Boundaries (Bhattarai and Vu, February 2026)
201 ▪ Focus Session: LLM4PQC -- An Agentic Framework for Accurate and Efficient Synthesis of PQC Cores (Perera et al., February 2026)
202 ▪ Autonomous Action Runtime Management(AARM):A System Specification for Securing AI-Driven Actions at Runtime (Errico, February 2026)
203 ▪ Agent-Fence: Mapping Security Vulnerabilities Across Deep Research Agents (Puppala et al., February 2026)
204 ▪ AgentSys: Secure and Dynamic LLM Agents Through Explicit Hierarchical Memory Management (Wen et al., February 2026)
205 ▪ Aegis: Towards Governance, Integrity, and Security of AI Voice Agents (Li, Chen, and Wei, February 2026)
206 ▪ Patch-to-PoC: A Systematic Study of Agentic LLM Systems for Linux Kernel N-Day Reproduction (Pu et al., February 2026)
207 ▪ Malicious Agent Skills in the Wild: A Large-Scale Security Empirical Study (Liu et al., February 2026)
208 ▪ Zero-Trust Runtime Verification for Agentic Payment Protocols: Mitigating Replay and Context-Binding Failures in AP2 (Lan et al., February 2026)
209 ▪ CVE-Factory: Scaling Expert-Level Agentic Tasks for Code Security Vulnerability (Luo et al., February 2026)
210 ▪ Human Society-Inspired Approaches to Agentic AI Security: The 4C Framework (Abuadbba et al., February 2026)
211 ▪ Protocol Agent: What If Agents Could Use Cryptography In Everyday Life? (Rossi, February 2026)
212 ▪ Semantic-Aware Advanced Persistent Threat Detection Using Autoencoders on LLM-Encoded System Logs (Mohammed et al., February 2026)
213 ▪ StepShield: When, Not Whether to Intervene on Rogue Agents (Felicia et al., January 2026)
214 ▪ AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection (Charoenwet et al., January 2026)
215 ▪ Multi-Agent Collaborative Intrusion Detection for Low-Altitude Economy IoT: An LLM-Enhanced Agentic AI Framework (Li et al., January 2026)
216 ▪ Secure Intellicise Wireless Network: Agentic AI for Coverless Semantic Steganography Communication (Meng et al., January 2026)
217 ▪ An LLM Agent-based Framework for Whaling Countermeasures (Miyamoto, Iimura, and Michishita, January 2026)
218 ▪ AgenTRIM: Tool Risk Mitigation for Agentic AI (Betser et al., January 2026)
219 ▪ Zero-Permission Manipulation: Can We Trust Large Multimodal Model Powered GUI Agents? (Qian et al., January 2026)
220 ▪ Taming Various Privilege Escalation in LLM-Based Agent Systems: A Mandatory Access Control Framework (Ji et al., January 2026)
221 ▪ Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agents (Zhou et al., January 2026)
222 ▪ Too Helpful to Be Safe: User-Mediated Attacks on Planning and Web-Use Agents (Chen et al., January 2026)
223 ▪ AgentGuardian: Learning Access Control Policies to Govern AI Agent Behavior (Abaev et al., January 2026)
224 ▪ Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale (Liu et al., January 2026)
225 ▪ Sola-Visibility-ISPM: Benchmarking Agentic AI for Identity Security Posture Management Visibility (Engelberg et al., January 2026)
226 ▪ FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments (Yang et al., January 2026)
227 ▪ Agentic AI Microservice Framework for Deepfake and Document Fraud Detection in KYC Pipelines (Kubam, January 2026)
228 ▪ A Survey of Agentic AI and Cybersecurity: Challenges, Opportunities and Use-case Prototypes (Lazer et al., January 2026)
229 ▪ Integrating Multi-Agent Simulation, Behavioral Forensics, and Trust-Aware Machine Learning for Adaptive Insider Threat Detection (Kausar et al., January 2026)
230 ▪ Web Fraud Attacks Against LLM-Driven Multi-Agent Systems (Kong et al., January 2026)
231 ▪ SastBench: A Benchmark for Testing Agentic SAST Triage (Feiglin and Dar, January 2026)
232 ▪ MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI (Xing et al., January 2026)
233 ▪ Device-Native Autonomous Agents for Privacy-Preserving Negotiations (Roy, January 2026)
234 ▪ The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models (Canale and Thimmaraju, January 2026)
235 ▪ Security in the Age of AI Teammates: An Empirical Study of Agentic Pull Requests on GitHub (Siddiq et al., January 2026)
236 ▪ Zero-Trust Agentic Federated Learning for Secure IIoT Defense Systems (Singh, Roy, and So, January 2026)
237 ▪ Audited Skill-Graph Self-Improvement for Agentic LLMs via Verifiable Rewards, Experience Synthesis, and Continual Memory (Huang and Huang, January 2026)
238 ▪ Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation (Syed et al., December 2025)
239 ▪ Agentic AI for Cyber Resilience: A New Security Paradigm and Its System-Theoretic Foundations (Li and Zhu, December 2025)
240 ▪ Securing Agentic AI Systems -- A Multilayer Security Framework (Arora and Hastings, December 2025)
241 ▪ Binding Agent ID: Unleashing the Power of AI Agents with accountability and credibility (Lin et al., December 2025)
242 ▪ Biosecurity-Aware AI: Agentic Risk Auditing of Soft Prompt Attacks on ESM-Based Variant Predictors (Zhan, December 2025)
243 ▪ BashArena: A Control Setting for Highly Privileged AI Agents (Kaufman et al., December 2025)
244 ▪ Penetration Testing of Agentic AI: A Comparative Security Analysis Across Models and Frameworks (Nguyen and Husain, December 2025)
245 ▪ Factor(U,T): Controlling Untrusted AI by Monitoring their Plans (Lip et al., December 2025)
246 ▪ ceLLMate: Sandboxing Browser AI Agents (Meng et al., December 2025)
247 ▪ MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents (Zhu et al., December 2025)
248 ▪ Locus: Agentic Predicate Synthesis for Directed Fuzzing (Zhu et al., December 2025)
249 ▪ AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration (Karthikeyan et al., December 2025)
250 ▪ Agentic Artificial Intelligence for Ethical Cybersecurity in Uganda: A Reinforcement Learning Framework for Threat Detection in Resource-Constrained Environments (Adabara et al., December 2025)
251 ▪ SoK: Trust-Authorization Mismatch in LLM Agent Interactions (Shi et al., December 2025)
252 ▪ The Evolution of Agentic AI in Cybersecurity: From Single LLM Reasoners to Multi-Agent Systems and Autonomous Pipelines (Vinay, December 2025)
253 ▪ AgenticCyber: A GenAI-Powered Multi-Agent System for Multimodal Threat Detection and Adaptive Response in Cybersecurity (Roy, December 2025)
254 ▪ Please Don't Kill My Vibe: Empowering Agents with Data Flow Control (Summers, Mohammed, and Wu, December 2025)
255 ▪ ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications (Bandara et al., December 2025)
256 ▪ PBFuzz: Agentic Directed Fuzzing for PoV Generation (Zeng et al., December 2025)
257 ▪ Password-Activated Shutdown Protocols for Misaligned Frontier Agents (Williams, Subramani, and Ward, December 2025)
258 ▪ LeechHijack: Covert Computational Resource Exploitation in Intelligent Agent Systems (Zhang et al., December 2025)
259 ▪ Systems Security Foundations for Agentic Computing (Christodorescu et al., December 2025)
260 ▪ AgentShield: Make MAS more secure and efficient (Wang et al., December 2025)
261 ▪ A Safety and Security Framework for Real-World Agentic Systems (Ghosh et al., December 2025)
262 ▪ FHE-Agent: Automating CKKS Configuration for Practical Encrypted Inference via an LLM-Guided Agentic Framework (Xu et al., November 2025)
263 ▪ LLMs as Firmware Experts: A Runtime-Grown Tree-of-Agents Framework (Zhang et al., November 2025)
264 ▪ ASTRA: Agentic Steerability and Risk Assessment Framework (Hazan et al., November 2025)
265 ▪ Towards Automating Data Access Permissions in AI Agents (Wu et al., November 2025)
266 ▪ LLM-Agent-UMF: LLM-based Agent Unified Modeling Framework for Seamless Design of Multi Active/Passive Core-Agent Architectures (Hassouna, Chaari, and Belhaj, November 2025)
267 ▪ Hiding in the AI Traffic: Abusing MCP for LLM-Powered Agentic Red Teaming (Janjuesvic, Garcia, and Kazerounian, November 2025)
268 ▪ Secure Autonomous Agent Payments: Verifying Authenticity and Intent in a Trustless Environment (Acharya, November 2025)
269 ▪ MAIF: Enforcing AI Trust and Provenance with an Artifact-Centric Agentic Paradigm (Narajala et al., November 2025)
270 ▪ Value-Aligned Prompt Moderation via Zero-Shot Agentic Rewriting for Safe Image Generation (Zhao et al., November 2025)
271 ▪ Towards a Generalisable Cyber Defence Agent for Real-World Computer Networks (Dudman and Bull, November 2025)
272 ▪ 3D Guard-Layer: An Integrated Agentic AI Safety System for Edge Artificial Intelligence (Kurshan, Xie, and Franzon, November 2025)
273 ▪ From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection (Han et al., November 2025)
274 ▪ PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts (Andersson et al., November 2025)
275 ▪ Security Analysis of Agentic AI Communication Protocols: A Comparative Evaluation (Louck, Stulman, and Dvir, November 2025)
276 ▪ 1 PoCo: Agentic Proof-of-Concept Exploit Generation for Smart Contracts (Andersson et al., November 2025)
277 ▪ LAFA: Agentic LLM-Driven Federated Analytics over Decentralized Data Sources (Ji et al., November 2025)
278 ▪ AAGATE: A NIST AI RMF-Aligned Governance Platform for Agentic AI (Huang et al., October 2025)
279 ▪ Identity Management for Agentic AI: The new frontier of authorization, authentication, and security for an AI agent world (South et al., October 2025)
280 ▪ AgentCyTE: Leveraging Agentic AI to Generate Cybersecurity Training & Experimentation Scenarios (Rodriguez et al., October 2025)
281 ▪ Securing AI Agent Execution (B\"uhler et al., October 2025)
282 ▪ AI Agentic Vulnerability Injection And Transformation with Optimized Reasoning (Lbath et al., October 2025)
283 ▪ CLASP: Cost-Optimized LLM-based Agentic System for Phishing Detection (Trad and Chehab, October 2025)
284 ▪ The Trust Paradox in LLM-Based Multi-Agent Systems: When Collaboration Becomes a Security Vulnerability (Xu et al., October 2025)
285 ▪ Cybersecurity AI: Evaluating Agentic Cybersecurity in Attack/Defense CTFs (Balassone et al., October 2025)
286 ▪ Terrarium: Revisiting the Blackboard for Multi-Agent Safety, Privacy, and Security Studies (Nakamura et al., October 2025)
287 ▪ Formalizing the Safety, Security, and Functional Properties of Agentic AI Systems (Allegrini, Shreekumar, and Celik, October 2025)
288 ▪ A2AS: Agentic AI Runtime Security and Self-Defense (Neelou et al., October 2025)
289 ▪ AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement (Rosser and Foerster, October 2025)
290 ▪ A Vision for Access Control in LLM-based Agent Systems (Li et al., October 2025)
291 ▪ Uncertainty-Aware, Risk-Adaptive Access Control for Agentic Systems using an LLM-Judged TBAC Model (Fleming, Kundu, and Kompella, October 2025)
292 ▪ Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols (Terekhov et al., October 2025)
293 ▪ A Survey on Agentic Security: Applications, Threats and Defenses (Shahriar et al., October 2025)
294 ▪ Code Agent can be an End-to-end System Hacker: Benchmarking Real-world Threats of Computer-use Agent (Luo et al., October 2025)
295 ▪ AutoPentester: An LLM Agent-based Framework for Automated Pentesting (Ginige et al., October 2025)
296 ▪ Adapting Insider Risk mitigations for Agentic Misalignment: an empirical study (Gomez, October 2025)
297 ▪ Agentic Misalignment: How LLMs Could Be Insider Threats (Lynch et al., October 2025)
298 ▪ Autonomy Matters: A Study on Personalization-Privacy Dilemma in LLM Agents (Zhang et al., October 2025)
299 ▪ Quantifying Distributional Robustness of Agentic Tool-Selection (Yeon, Chaudhary, and Singh, October 2025)
300 ▪ PentestMCP: A Toolkit for Agentic Penetration Testing (Ezetta and Feng, October 2025)
301 ▪ MobiLLM: An Agentic AI Framework for Closed-Loop Threat Mitigation in 6G Open RANs (Sharma et al., October 2025)
302 ▪ ToolTweak: An Attack on Tool Selection in LLM-based Agents (Sneh et al., October 2025)
303 ▪ Agentic-AI Healthcare: Multilingual, Privacy-First Framework with MCP Agents (Shehab, October 2025)
304 ▪ FalseCrashReducer: Mitigating False Positive Crashes in OSS-Fuzz-Gen Using Agentic AI (Amusuo et al., October 2025)
305 ▪ Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness (Shayegani et al., October 2025)
306 ▪ Better Privilege Separation for Agents by Restricting Data Types (Jacob et al., October 2025)
307 ▪ Agentic Specification Generator for Move Programs (Fu, Xu, and Kim, September 2025)
308 ▪ LISA Technical Report: An Agentic Framework for Smart Contract Auditing (Sun, Tan, and Deng, September 2025)
309 ▪ DIRF: A Framework for Digital Identity Protection and Clone Governance in Agentic AI Systems (Atta et al., August 2025)
310 ▪ Measuring Harmfulness of Computer-Using Agents (Tian et al., August 2025)
311 ▪ Beyond DNS: Unlocking the Internet of AI Agents via the NANDA Index and Verified AgentFacts (Raskar et al., July 2025)
312 ▪ From Semantic Web and MAS to Agentic AI: A Unified Narrative of the Web of Agents (SnT et al., July 2025)
313 ▪ Game Theory Meets LLM and Agentic AI: Reimagining Cybersecurity for the Age of Intelligent Threats (Zhu, July 2025)
314 ▪ Giving AI Agents Access to Cryptocurrency and Smart Contracts Creates New Vectors of AI Harm (Marino and Juels, July 2025)
315 ▪ The Dark Side of LLMs: Agent-based Attacks for Complete Computer Takeover (Lupinacci et al., July 2025)
316 ▪ The Trust Fabric: Decentralized Interoperability and Economic Coordination for the Agentic Web (Balija et al., July 2025)
317 ▪ The Dark Side of LLMs Agent-based Attacks for Complete Computer Takeover (Lupinacci et al., July 2025)
318 ▪ A Systematization of Security Vulnerabilities in Computer Use Agents (Jones et al., July 2025)
319 ▪ CRAKEN: Cybersecurity LLM Agent with Knowledge-Based Execution<br><br> (Shao et al., May 2025)
320 ▪ The Odyssey of the Fittest: Can Agents Survive and Still Be Good?<br><br> (Waldner and Miikkulainen, May 2025)
321 ▪ Agency Is Frame-Dependent (Abel et al, Feb 2025)
322 ▪ Security of AI Agents (He et al, Dec 2024)
323 ▪ Dissecting Adversarial Robustness of Multimodal LM Agents (Wu et al, Dec 2024)
324 ▪ RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors (Bai et al, Dec 2024)
325 ▪ AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents (Debenedetti et al, Nov 2024)
326 ▪ AdvWeb: Controllable Black-box Attacks on VLM-powered Web Agents (Xu et al, Oct 2024)
327 ▪ Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation (Kwartler et al, Oct 2024)
328 ▪ Bayes-Nash Generative Privacy Protection Against Membership Inference Attacks (Zhang et al, Oct 2024)
329 ▪ Prompt Infection: LLM-to-LLM Prompt Injection within Multi-Agent Systems (Lee and Tiwari, Oct 2024)
330 ▪ Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents (Zhang et al, Oct 2024)
331 ▪ BreachSeek: A Multi-Agent Automated Penetration Tester (Alshehri et al, Sep 2024)
332 ▪ Safeguarding AI Agents: Developing and Analyzing Safety Architectures (Domkunwar and N S, Sep 2024)
333 ▪ Secret Collusion among Generative AI Agents (Motwani et al, Aug 2024)
334 ▪ Large Language Model Sentinel: LLM Agent for Adversarial Purification (Lin and Zhao, Aug 2024)
335 ▪ Compromising Embodied Agents with Contextual Backdoor Attacks (Liu et al, Aug 2024)
336 ▪ Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification (Zhang et al, Jul 2024)
Excessive Agency, Agentic Manipulation, Agentic Systems

>

<

‍

Copyright Infringement

Covers:

  • MITRE ATLAS Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Copyright Infringement

cybersecurity_tracker - Google Drive

2 ▪ From Compression to Accountability: Harmless Copyright Protection for Dataset Distillation (Liang et al., May 2026)
3 ▪ Position: No Retroactive Cure for Infringement during Training (Utsunomiya et al., April 2026)
4 ▪ MATRIX: Multi-Layer Code Watermarking via Dual-Channel Constrained Parity-Check Encoding (Nie et al., April 2026)
5 ▪ Geometry-Aware Localized Watermarking for Copyright Protection in Embedding-as-a-Service (Chen et al., April 2026)
6 ▪ Towards trustworthy management of AIGC copyright: blockchain-enabled full lifecycle recording and multi-party auditing approach (Jiang et al., April 2026)
7 ▪ Copyright Protection for Large Language Models: A Survey of Methods, Challenges, and Trends (Xu et al., April 2026)
8 ▪ A PUF-Based Approach for Copy Protection of Intellectual Property in Neural Network Models (Dorfmeister et al., March 2026)
9 ▪ Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment (Chen, Lou, and Wang, March 2026)
10 ▪ DWBench: Holistic Evaluation of Watermark for Dataset Copyright Auditing (Ren et al., February 2026)
11 ▪ Traceable Black-box Watermarks for Federated Learning (Xu et al., February 2026)
12 ▪ SEW: Strengthening Robustness of Black-box DNN Watermarking via Specificity Enhancement (Qiu et al., February 2026)
13 ▪ Lossless Copyright Protection via Intrinsic Model Fingerprinting (Chen et al., January 2026)
14 ▪ On the Evidentiary Limits of Membership Inference for Copyright Auditing (Ertan et al., January 2026)
15 ▪ DNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection (Xu et al., January 2026)
16 ▪ Multi-Agent Framework for Controllable and Protected Generative Content Creation: Addressing Copyright and Provenance in AI-Generated Media (Khan, Asif, and Asif, January 2026)
17 ▪ SRAF: Stealthy and Robust Adversarial Fingerprint for Copyright Verification of Large Language Models (Wang et al., January 2026)
18 ▪ SWaRL: Safeguard Code Watermarking via Reinforcement Learning (Javidnia et al., January 2026)
19 ▪ Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content? (Xu et al., December 2025)
20 ▪ Towards Dataset Copyright Evasion Attack against Personalized Text-to-Image Diffusion Models (Gao et al., December 2025)
21 ▪ From Essence to Defense: Adaptive Semantic-aware Watermarking for Embedding-as-a-Service Copyright Protection (Li et al., December 2025)
22 ▪ VICTOR: Dataset Copyright Auditing in Video Recognition Systems (Yuan et al., December 2025)
23 ▪ Blameless Users in a Clean Room: Defining Copyright Protection for Generative Models (Cohen, December 2025)
24 ▪ Copyright in AI Pre-Training Data Filtering: Regulatory Landscape and Mitigation Strategies (Kyrychenko, Mudryi, and Chaklosh, December 2025)
25 ▪ EnTruth: Enhancing the Traceability of Unauthorized Dataset Usage in Text-to-image Diffusion Models with Minimal and Robust Alterations (Ren et al., November 2025)
26 ▪ AuthenLoRA: Entangling Stylization with Imperceptible Watermarks for Copyright-Secure LoRA Adapters (Shi et al., November 2025)
27 ▪ PromptCOS: Towards Content-only System Prompt Copyright Auditing for LLMs (Yang et al., November 2025)
28 ▪ SoK: Large Language Model Copyright Auditing via Fingerprinting (Shao et al., November 2025)
29 ▪ RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright Protection (Yang et al., November 2025)
30 ▪ SWAP: Towards Copyright Auditing of Soft Prompts via Sequential Watermarking (Yang et al., November 2025)
31 ▪ SLIP: Securing LLMs IP Using Weights Decomposition (Refael et al., November 2025)
32 ▪ Large Language Models Are Effective Code Watermarkers (Xu et al., October 2025)
33 ▪ Protecting Copyrighted Material with Unique Identifiers in Large Language Model Training (Zhao et al., July 2025)
34 ▪ Content ARCs: Decentralized Content Rights in the Age of Generative AI (Balan, Gilbert, Collomosse, Mar 2025)
35 ▪ SoK: Dataset Copyright Auditing in Machine Learning Systems (Du et al, Oct 2024)
36 ▪ Protecting Copyright of Medical Pre-trained Language Models: Training-Free Backdoor Watermarking (Kong et al, Sep 2024)
37 ▪ Strong Copyright Protection for Language Models via Adaptive Model Fusion (Wang et al, Jul 2024)
Copyright Infringement

>

<

‍

Defenses

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Defenses

cybersecurity_tracker - Google Drive

2 ▪ Toward Securing AI Agents Like Operating Systems (Pirch et al., May 2026)
3 ▪ One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries (Zloczower et al., May 2026)
4 ▪ Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models (Meng et al., May 2026)
5 ▪ MemLineage: Lineage-Guided Enforcement for LLM Agent Memory (Ouyang and Hou, May 2026)
6 ▪ ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents (Lee and Brumley, May 2026)
7 ▪ Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution (Zhang et al., May 2026)
8 ▪ RACC: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., May 2026)
9 ▪ QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild (Safdar et al., May 2026)
10 ▪ Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing (Zhao et al., May 2026)
11 ▪ Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions (Susan, Arusoaie, and Lucanu, May 2026)
12 ▪ What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers (Zhang et al., May 2026)
13 ▪ Containment Verification: AI Safety Guarantees Independent of Alignment (Moon and Varshney, May 2026)
14 ▪ Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims (Konrad, Tanyel, and Ayvaz, May 2026)
15 ▪ Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning (Huff et al., May 2026)
16 ▪ MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring (Jotautait\.e et al., May 2026)
17 ▪ SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization (Liu et al., May 2026)
18 ▪ Seed Hijacking of LLM Sampling and Quantum Random Number Defense (You et al., May 2026)
19 ▪ Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks (Hu, May 2026)
20 ▪ GLiGuard: Schema-Conditioned Classification for LLM Safeguard (Zaratiana et al., May 2026)
21 ▪ FIT to Forget: Robust Continual Unlearning for Large Language Models (Xu et al., May 2026)
22 ▪ SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering (Lin et al., May 2026)
23 ▪ Information Theoretic Adversarial Training of Large Language Models (Zhang et al., May 2026)
24 ▪ Autonomous Adversary: Red-Teaming in the age of LLM (Mamun et al., May 2026)
25 ▪ Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks (Lu et al., May 2026)
26 ▪ SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety (Liu et al., May 2026)
27 ▪ GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy (Minko, Sadiekh, and Kokuykin, May 2026)
28 ▪ A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework (Chu, May 2026)
29 ▪ On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference (Li et al., May 2026)
30 ▪ Large Language Model assisted Hybrid Fuzzing (Meng, Duck, and Roychoudhury, May 2026)
31 ▪ MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents (Steinberg and Gal, May 2026)
32 ▪ ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection (Weng et al., May 2026)
33 ▪ MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory (Wang et al., May 2026)
34 ▪ Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses (Li et al., May 2026)
35 ▪ RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs (Asif and Amiri, May 2026)
36 ▪ Toward a Principled Framework for Agent Safety Measurement (Lin et al., May 2026)
37 ▪ When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems (Zhang, Zheng, and Chen, May 2026)
38 ▪ Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts (Hasan et al., May 2026)
39 ▪ ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models (Zhao et al., May 2026)
40 ▪ GAVEL: Towards Rule-Based Safety Through Activation Monitoring (Rozenfeld et al., May 2026)
41 ▪ Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures (Mahi et al., May 2026)
42 ▪ AdaBFL: Multi-Layer Defensive Adaptive Aggregation for Bzantine-Robust Federated Learning (Tang, Liu, and Huang, May 2026)
43 ▪ Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches (Wienczkowski, May 2026)
44 ▪ Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study (Xu and Chen, May 2026)
45 ▪ Toward Autonomous SOC Operations: End-to-End LLM Framework for Threat Detection, Query Generation, and Resolution in Security Operations (Saju and Azim, May 2026)
46 ▪ An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code (Elsayed, Fulton, and Yang, May 2026)
47 ▪ Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents (Dang, April 2026)
48 ▪ Large Language Models as Explainable Cyberattack Detectors for Energy Industrial Control Systems (Kong et al., April 2026)
49 ▪ A Comparative Evaluation of AI Agent Security Guardrails (Li et al., April 2026)
50 ▪ Mitigating Error Amplification in Fast Adversarial Training (Zhao et al., April 2026)
51 ▪ AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization (Ying et al., April 2026)
52 ▪ Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents (Li et al., April 2026)
53 ▪ Evaluation of Prompt Injection Defenses in Large Language Models (Deep et al., April 2026)
54 ▪ UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks (Yu et al., April 2026)
55 ▪ Secure eFPGA-Enabled Edge LLM Inference: Architectural and Hardware Countermeasures (Das et al., April 2026)
56 ▪ AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models (Gautam, Bahramali, and Atluri, April 2026)
57 ▪ SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios (Chen et al., April 2026)
58 ▪ Secure LLM Fine-Tuning via Safety-Aware Probing (Wu et al., April 2026)
59 ▪ Adaptive Instruction Composition for Automated LLM Red-Teaming (Zymet et al., April 2026)
60 ▪ Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks (Pallerla et al., April 2026)
61 ▪ Verification of Machine Unlearning is Fragile (Zhang et al., April 2026)
62 ▪ AVISE: Framework for Evaluating the Security of AI Systems (Lempinen, Kemppainen, and Raesalmi, April 2026)
63 ▪ Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing (Talluri, April 2026)
64 ▪ Towards Certified Malware Detection: Provable Guarantees Against Evasion Attacks (Giri et al., April 2026)
65 ▪ Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning (Ni, Christodorescu, and Jia, April 2026)
66 ▪ Benchmarking Misuse Mitigation Against Covert Adversaries (Brown et al., April 2026)
67 ▪ ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System (Liang et al., April 2026)
68 ▪ Malicious ML Model Detection by Learning Dynamic Behaviors (Nambiar, Pradhan, and Soremekun, April 2026)
69 ▪ SAGE: Signal-Amplified Guided Embeddings for LLM-based Vulnerability Detection (Shan et al., April 2026)
70 ▪ Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents (Xu, April 2026)
71 ▪ ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction (Wei, Wu, and Sun, April 2026)
72 ▪ SDLLMFuzz: Dynamic-static LLM-assisted greybox fuzzing for structured input programs (Zou et al., April 2026)
73 ▪ Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives (Son, April 2026)
74 ▪ enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways (Metere, April 2026)
75 ▪ SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models (Mohammad and Bayaz{\i}t, April 2026)
76 ▪ TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts (Chu, Wang, and Huang, April 2026)
77 ▪ Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility (Hong et al., April 2026)
78 ▪ Privacy-Preserving LLMs Routing (Wu et al., April 2026)
79 ▪ SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment (Lin et al., April 2026)
80 ▪ Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory (Fu and Wang, April 2026)
81 ▪ SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents (Begimher et al., April 2026)
82 ▪ Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns (Li et al., April 2026)
83 ▪ Towards Automated Pentesting with Large Language Models (Bessa et al., April 2026)
84 ▪ QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits (Azimi et al., April 2026)
85 ▪ PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification (Gong and Deng, April 2026)
86 ▪ Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark (Shao et al., April 2026)
87 ▪ Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction (Chen et al., April 2026)
88 ▪ Vulnerability Detection with Interprocedural Context in Multiple Languages: Assessing Effectiveness and Cost of Modern LLMs (Lira et al., April 2026)
89 ▪ Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions (Xu et al., April 2026)
90 ▪ Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models (Qi et al., April 2026)
91 ▪ MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security (Rostamzadeh et al., April 2026)
92 ▪ Can Drift-Adaptive Malware Detectors Be Made Robust? Attacks and Defenses Under White-Box and Black-Box Threats (Li, Akil, and Bertino, April 2026)
93 ▪ The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? (Bhatt et al., April 2026)
94 ▪ Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization (Maljkovic et al., April 2026)
95 ▪ SALLIE: Safeguarding Against Latent Language & Image Exploits (Azov, Rivlin, and Shtar, April 2026)
96 ▪ CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments (Keppler, Gst\"ur, and Hagenmeyer, April 2026)
97 ▪ A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms (Acharya and Gupta, April 2026)
98 ▪ Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts (Uenal, April 2026)
99 ▪ BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents (Ma, Wu, and Yan, April 2026)
100 ▪ Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code (Blain and Noiseux, April 2026)
101 ▪ SoSBench: Benchmarking Safety Alignment on Six Scientific Domains (Jiang et al., April 2026)
102 ▪ Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework (Zhou et al., April 2026)
103 ▪ Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning (Zhang, Goel, and Ahmad, April 2026)
104 ▪ CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks (Li et al., April 2026)
105 ▪ Automating Cloud Security and Forensics Through a Secure-by-Design Generative AI Framework (Alharthi and Garcia, April 2026)
106 ▪ SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization (Wang et al., April 2026)
107 ▪ Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI (Cheruvu et al., April 2026)
108 ▪ Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection (Halder et al., April 2026)
109 ▪ Assertain: Automated Security Assertion Generation Using Large Language Models (Tarek et al., April 2026)
110 ▪ Certifiably Robust RAG against Retrieval Corruption (Xiang et al., April 2026)
111 ▪ Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents (Ye et al., April 2026)
112 ▪ VibeGuard: A Security Gate Framework for AI-Generated Code (Xie, April 2026)
113 ▪ Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks (Sahu, Samanta, and Soosahabi, April 2026)
114 ▪ Quantum-Safe Code Auditing: LLM-Assisted Static Analysis and Quantum-Aware Risk Scoring for Post-Quantum Cryptography Migration (Shaw, April 2026)
115 ▪ SecureVibeBench: Evaluating Secure Coding Capabilities of Code Agents with Realistic Vulnerability Scenarios (Chen et al., April 2026)
116 ▪ Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks (Xiang et al., April 2026)
117 ▪ CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks (Patil, April 2026)
118 ▪ Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems (Cai et al., April 2026)
119 ▪ Attesting LLM Pipelines: Enforcing Verifiable Training and Release Claims (Tan, Singer, and Anagnostopoulos, April 2026)
120 ▪ GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models (Mia et al., April 2026)
121 ▪ SkillTester: Benchmarking Utility and Security of Agent Skills (Wang, Wang, and Xu, April 2026)
122 ▪ VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications (Zhu et al., March 2026)
123 ▪ Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries (Raz et al., March 2026)
124 ▪ A Systematic Taxonomy of Security Vulnerabilities in the OpenClaw AI Agent Framework (Suwansathit, Zhang, and Gu, March 2026)
125 ▪ Protecting User Prompts Via Character-Level Differential Privacy (Arachchige et al., March 2026)
126 ▪ Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models (Wang et al., March 2026)
127 ▪ Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage (Sahay et al., March 2026)
128 ▪ Agent Audit: A Security Analysis System for LLM Agent Applications (Zhang, Nian, and Zhao, March 2026)
129 ▪ Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval (Sabir et al., March 2026)
130 ▪ BioShield: A Context-Aware Firewall for Securing Bio-LLMs (Das et al., March 2026)
131 ▪ T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search (Lee et al., March 2026)
132 ▪ Agentproof: Static Verification of Agent Workflow Graphs (Xavier et al., March 2026)
133 ▪ SecureBreak -- A dataset towards safe and secure models (Arazzi, Kembu, and Nocera, March 2026)
134 ▪ Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks (Mu et al., March 2026)
135 ▪ DeepXplain: XAI-Guided Autonomous Defense Against Multi-Stage APT Campaigns (Phan and Bauschert, March 2026)
136 ▪ LISAA: A Framework for Large Language Model Information Security Awareness Assessment (Cohen et al., March 2026)
137 ▪ A Framework for Formalizing LLM Agent Security (Siu et al., March 2026)
138 ▪ The Autonomy Tax: Defense Training Breaks LLM Agents (Li and Zhao, March 2026)
139 ▪ Network and Device Level Cyber Deception for Contested Environments Using RL and LLMs (Sahu, Paul, and Macwan, March 2026)
140 ▪ Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety (Chen et al., March 2026)
141 ▪ Security awareness in LLM agents: the NDAI zone case (Bottazzi and Park, March 2026)
142 ▪ Toward Reliable, Safe, and Secure LLMs for Scientific Applications (Chaturvedi, Bergerson, and Mallick, March 2026)
143 ▪ Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows (Sigdel and Baral, March 2026)
144 ▪ Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs (Yang and Zhu, March 2026)
145 ▪ Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework (Onitiju and Vakilinia, March 2026)
146 ▪ DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns (Phan, Nguyen, and Bauschert, March 2026)
147 ▪ Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory (Zhang et al., March 2026)
148 ▪ Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models (Liu and Chen, March 2026)
149 ▪ TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems (Wang et al., March 2026)
150 ▪ SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration (Pan et al., March 2026)
151 ▪ Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats (Zhang et al., March 2026)
152 ▪ Governing Dynamic Capabilities: Cryptographic Binding and Reproducibility Verification for AI Agent Tool Use (Zhou, March 2026)
153 ▪ SecDTD: Dynamic Token Drop for Secure Transformers Inference (Cai et al., March 2026)
154 ▪ CTI-REALM: Benchmark to Evaluate Agent Performance on Security Detection Rule Generation Capabilities (Chakraborty et al., March 2026)
155 ▪ Why Neural Structural Obfuscation Can't Kill White-Box Watermarks for Good! (Jiang et al., March 2026)
156 ▪ Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw (Ying et al., March 2026)
157 ▪ AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents (Yuan, Su, and Zhao, March 2026)
158 ▪ Security Considerations for Artificial Intelligence Agents (Li et al., March 2026)
159 ▪ OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents (Li, March 2026)
160 ▪ Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms (Wendlinger et al., March 2026)
161 ▪ Differential Privacy in Machine Learning: A Survey from Symbolic AI to LLMs (Aguilera-Mart\'inez and Berzal, March 2026)
162 ▪ TOSSS: a CVE-based Software Security Benchmark for Large Language Models (Damie et al., March 2026)
163 ▪ Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial Attacks (Soltani et al., March 2026)
164 ▪ AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation (Li et al., March 2026)
165 ▪ SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement (Chen et al., March 2026)
166 ▪ DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation (Jiang, March 2026)
167 ▪ Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI (Jin et al., March 2026)
168 ▪ Where Do LLM-based Systems Break? A System-Level Security Framework for Risk Assessment and Treatment (Nagaraja and Bahsi, March 2026)
169 ▪ ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code (Filho, March 2026)
170 ▪ Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models (Wu et al., March 2026)
171 ▪ Secure human oversight of AI: Threat modeling in a socio-technical context (Ditz et al., March 2026)
172 ▪ EVMbench: Evaluating AI Agents on Smart Contract Security (Wang et al., March 2026)
173 ▪ Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks (Chu et al., March 2026)
174 ▪ Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study (Nagaraja and Bahsi, March 2026)
175 ▪ WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols (Maheri et al., March 2026)
176 ▪ LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities (Nakka et al., March 2026)
177 ▪ Contextualized Privacy Defense for LLM Agents (Wen et al., March 2026)
178 ▪ SOSecure: Safer Code Generation with RAG and StackOverflow Discussions (Mukherjee and Hellendoorn, March 2026)
179 ▪ ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline (Singh et al., March 2026)
180 ▪ Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision (Mukherjee and Hellendoorn, March 2026)
181 ▪ Token-level Data Selection for Safe LLM Fine-tuning (Li et al., March 2026)
182 ▪ MPU: Towards Secure and Privacy-Preserving Knowledge Unlearning for Large Language Models (Wang et al., March 2026)
183 ▪ Learning to Generate Secure Code via Token-Level Rewards (Quan et al., March 2026)
184 ▪ Automated Vulnerability Detection in Source Code Using Deep Representation Learning (Seas et al., February 2026)
185 ▪ A Lightweight Defense Mechanism against Next Generation of Phishing Emails using Distilled Attention-Augmented BiLSTM (Eskandarian et al., February 2026)
186 ▪ Secure Semantic Communications via AI Defenses: Fundamentals, Solutions, and Future Directions (Zhang et al., February 2026)
187 ▪ Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning (Garc\'ia-M\'arquez et al., February 2026)
188 ▪ A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications (Srivastava, Janardhan, and Jauhari, February 2026)
189 ▪ AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs (Zhang et al., February 2026)
190 ▪ Dynamic Probabilistic Noise Injection for Membership Inference Defense (Forough and Haddadi, February 2026)
191 ▪ LLM-enabled Applications Require System-Level Threat Monitoring (Zhang et al., February 2026)
192 ▪ CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents (Ba, Li, and Li, February 2026)
193 ▪ KUDA: Knowledge Unlearning by Deviating Representation for Large Language Models (Fang et al., February 2026)
194 ▪ MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs (Kan et al., February 2026)
195 ▪ DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing (Brinkhege and Menon, February 2026)
196 ▪ NeST: Neuron Selective Tuning for LLM Safety (Behrouzi et al., February 2026)
197 ▪ NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist (Bertram and Geiping, February 2026)
198 ▪ Secure Coding with AI -- From Detection to Repair (Belozerov, Barclay, and Sami, February 2026)
199 ▪ PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models (Yuan et al., February 2026)
200 ▪ Large Language Models for Secure Code Assessment: A Multi-Language Empirical Study (Dozono, Gasiba, and Stocco, February 2026)
201 ▪ ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI (Tong et al., February 2026)
202 ▪ MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents (Zhou et al., February 2026)
203 ▪ From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection (Li, Yu, and Wang, February 2026)
204 ▪ Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents (Li et al., February 2026)
205 ▪ Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks (Zafar et al., February 2026)
206 ▪ TensorCommitments: A Lightweight Verifiable Inference for Language Models (Baser et al., February 2026)
207 ▪ Sparse Autoencoders are Capable LLM Jailbreak Mitigators (Assogba et al., February 2026)
208 ▪ Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset (Kang et al., February 2026)
209 ▪ DeepSight: An All-in-One LM Safety Toolkit (Zhang et al., February 2026)
210 ▪ The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates (Willis, February 2026)
211 ▪ SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models (Thornton, February 2026)
212 ▪ GoodVibe: Security-by-Vibe for LLM-Based Code Generation (Thang et al., February 2026)
213 ▪ MAPS: A Multilingual Benchmark for Agent Performance and Security (Hofman et al., February 2026)
214 ▪ Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference (Chu et al., February 2026)
215 ▪ Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks (Dhabhi and Thimmaraju, February 2026)
216 ▪ Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks (Fu, Guo, and Luo, February 2026)
217 ▪ MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots (Wang et al., February 2026)
218 ▪ Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model (Wu et al., February 2026)
219 ▪ Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention (Singh et al., February 2026)
220 ▪ TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (Hossain et al., February 2026)
221 ▪ Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models (Lu et al., February 2026)
222 ▪ How Catastrophic is Your LLM? Certifying Risk in Conversation (Wang et al., February 2026)
223 ▪ Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection (Firouzi and Ghafari, February 2026)
224 ▪ Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening (Yu et al., February 2026)
225 ▪ Comparative Insights on Adversarial Machine Learning from Industry and Academia: A User-Study Approach (Kakkad et al., February 2026)
226 ▪ Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models (Ding et al., February 2026)
227 ▪ Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation (Marri, February 2026)
228 ▪ GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video (Zhu et al., February 2026)
229 ▪ TinyGuard:A lightweight Byzantine Defense for Resource-Constrained Federated Learning via Statistical Update Fingerprints (Mahdavi et al., February 2026)
230 ▪ To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack (Zhuo et al., February 2026)
231 ▪ MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS (Wang et al., February 2026)
232 ▪ When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models (Ou et al., February 2026)
233 ▪ RACA: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., February 2026)
234 ▪ Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents (He et al., February 2026)
235 ▪ SMCP: Secure Model Context Protocol (Hou et al., February 2026)
236 ▪ Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation (Jamshidi et al., February 2026)
237 ▪ Detecting Instruction Fine-tuning Attacks using Influence Function (Li, February 2026)
238 ▪ No More, No Less: Least-Privilege Language Models (Rauba et al., February 2026)
239 ▪ RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories (Wang et al., February 2026)
240 ▪ FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks (Xu et al., February 2026)
241 ▪ A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy (Correia et al., February 2026)
242 ▪ ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense (Ding, February 2026)
243 ▪ SafeSearch: Automated Red-Teaming of LLM-Based Search Agents (Dong et al., January 2026)
244 ▪ FIT: Defying Catastrophic Forgetting in Continual LLM Unlearning (Xu et al., January 2026)
245 ▪ RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing (Zhang et al., January 2026)
246 ▪ Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses (Hatami et al., January 2026)
247 ▪ Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications (Shahin and Alsmadi, January 2026)
248 ▪ Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges (Patel et al., January 2026)
249 ▪ GAVEL: Towards rule-based safety through activation monitoring (Rozenfeld et al., January 2026)
250 ▪ RvB: Automating AI System Hardening via Iterative Red-Blue Games (Huang et al., January 2026)
251 ▪ LLM-Assisted Authentication and Fraud Detection (Chan and Chan, January 2026)
252 ▪ SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks (Sivaroopan et al., January 2026)
253 ▪ Proactive Hardening of LLM Defenses with HASTE (Chen et al., January 2026)
254 ▪ Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models (Chu et al., January 2026)
255 ▪ $\alpha^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks (Ferrag, Lakas, and Debbah, January 2026)
256 ▪ Mitigating the OWASP Top 10 For Large Language Models Applications using Intelligent Agents (Fasha et al., January 2026)
257 ▪ PatchIsland: Orchestration of LLM Agents for Continuous Vulnerability Repair (Kim et al., January 2026)
258 ▪ ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMs (Fayyazi et al., January 2026)
259 ▪ Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks (Tsiolakis et al., January 2026)
260 ▪ PAL*M: Property Attestation for Large Generative Models (Chantasantitam et al., January 2026)
261 ▪ Introducing the Generative Application Firewall (GAF) (Farreny et al., January 2026)
262 ▪ The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses (G{\l}uch et al., January 2026)
263 ▪ Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM) (Biswas and Talukdar, January 2026)
264 ▪ Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification (Meng et al., January 2026)
265 ▪ HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation (Chen et al., January 2026)
266 ▪ Serverless AI Security: Attack Surface Analysis and Runtime Protection Mechanisms for FaaS-Based Machine Learning (Pathade et al., January 2026)
267 ▪ SecMLOps: A Comprehensive Framework for Integrating Security Throughout the MLOps Lifecycle (Zhang et al., January 2026)
268 ▪ Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay (Wang et al., January 2026)
269 ▪ Blue Teaming Function-Calling Agents (Dolcetti, Zizzo, and Maffeis, January 2026)
270 ▪ Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis (Song et al., January 2026)
271 ▪ AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation (Kumar, January 2026)
272 ▪ How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape (Kelley et al., January 2026)
273 ▪ BlindU: Blind Machine Unlearning without Revealing Erasing Data (Wang et al., January 2026)
274 ▪ Defenses Against Prompt Attacks Learn Surface Heuristics (Li et al., January 2026)
275 ▪ Safe-FedLLM: Delving into the Safety of Federated Large Language Models (Tao et al., January 2026)
276 ▪ United We Defend: Collaborative Membership Inference Defenses in Federated Learning (Bai et al., January 2026)
277 ▪ Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model (M et al., January 2026)
278 ▪ AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization (Wang et al., January 2026)
279 ▪ VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit (Lin et al., January 2026)
280 ▪ HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense (Li et al., January 2026)
281 ▪ AI-Driven Cybersecurity Threats: A Survey of Emerging Risks and Defensive Strategies (Erukude, Marella, and Veluru, January 2026)
282 ▪ Autonomous Threat Detection and Response in Cloud Security: A Comprehensive Survey of AI-Driven Strategies (Sarraf and Pal, January 2026)
283 ▪ Automated Post-Incident Policy Gap Analysis via Threat-Informed Evidence Mapping using Large Language Models (Oh et al., January 2026)
284 ▪ SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones (Wu and Cao, January 2026)
285 ▪ How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System (Zuo et al., January 2026)
286 ▪ Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment (Xie et al., January 2026)
287 ▪ Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models (Rosario, January 2026)
288 ▪ Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection (Iyer, January 2026)
289 ▪ One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models (Gu et al., January 2026)
290 ▪ Improving LLM-Assisted Secure Code Generation through Retrieval-Augmented-Generation and Multi-Tool Feedback (Sriram et al., January 2026)
291 ▪ PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices (Chattopadhyay et al., January 2026)
292 ▪ Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing (Bhatt et al., January 2026)
293 ▪ Towards Provably Secure Generative AI: Reliable Consensus Sampling (Cui et al., January 2026)
294 ▪ How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench (Sajadi, Damevski, and Chatterjee, December 2025)
295 ▪ Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses (Wang, December 2025)
296 ▪ Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems (Foundjem et al., December 2025)
297 ▪ Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management (Arora and Hastings, December 2025)
298 ▪ Assessing the Software Security Comprehension of Large Language Models (Siddiq et al., December 2025)
299 ▪ AIAuditTrack: A Framework for AI Security system (Luo et al., December 2025)
300 ▪ AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs (Wang et al., December 2025)
301 ▪ On the Effectiveness of Instruction-Tuning Local LLMs for Identifying Software Vulnerabilities (Park, Ko, and Cho, December 2025)
302 ▪ IoT-based Android Malware Detection Using Graph Neural Network With Adversarial Defense (Yumlembam et al., December 2025)
303 ▪ Certified Defense on the Fairness of Graph Neural Networks (Dong et al., December 2025)
304 ▪ DREAM: Dynamic Red-teaming across Environments for AI Models (Lu et al., December 2025)
305 ▪ Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection (Pan et al., December 2025)
306 ▪ AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs (Gaikwad, December 2025)
307 ▪ Prefix Probing: Lightweight Harmful Content Detection for Large Language Models (Yang et al., December 2025)
308 ▪ Adversarial Robustness in Financial Machine Learning: Defenses, Economic Impact, and Governance Evidence (Baviskar, December 2025)
309 ▪ PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy (Kundu, Ahmed, and Hoque, December 2025)
310 ▪ Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks (Shaheer et al., December 2025)
311 ▪ Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection (Song et al., December 2025)
312 ▪ Auto-Tuning Safety Guardrails for Black-Box Large Language Models (Abdulkadir, December 2025)
313 ▪ Quantifying Return on Security Controls in LLM Systems (Moulton, O'Brien, and Hastings, December 2025)
314 ▪ GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients (Razmjoo, Sharifian, and Shouraki, December 2025)
315 ▪ Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS (Ennaji, Benkhelifa, and Mancini, December 2025)
316 ▪ SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models (Afane et al., December 2025)
317 ▪ Toward Intelligent and Secure Cloud: Large Language Model Empowered Proactive Defense (Zhou et al., December 2025)
318 ▪ CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models (Hall, Ungaro, and Eiers, December 2025)
319 ▪ Adaptive Intrusion Detection System Leveraging Dynamic Neural Models with Adversarial Learning for 5G/6G Networks (Neha and Bhatia, December 2025)
320 ▪ From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection (Lu and Lagaisse, December 2025)
321 ▪ Chasing Shadows: Pitfalls in LLM Security Research (Evertz et al., December 2025)
322 ▪ AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models (Wang et al., December 2025)
323 ▪ Evaluating the robustness of adversarial defenses in malware detection systems (Jafari and Shameli-Sendi, December 2025)
324 ▪ Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples (Zhou et al., December 2025)
325 ▪ VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection (Nie et al., December 2025)
326 ▪ Securing the Model Context Protocol: Defending LLMs Against Tool Poisoning and Adversarial Attacks (Jamshidi et al., December 2025)
327 ▪ IF-GUIDE: Influence Function-Guided Detoxification of LLMs (Coalson et al., December 2025)
328 ▪ Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model (Liu et al., December 2025)
329 ▪ ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior (Lu et al., December 2025)
330 ▪ Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models (Cretu et al., December 2025)
331 ▪ AutoGuard: A Self-Healing Proactive Security Layer for DevSecOps Pipelines Using Reinforcement Learning (Anugula et al., December 2025)
332 ▪ Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs (Ma et al., December 2025)
333 ▪ Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks (Fu et al., December 2025)
334 ▪ VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs (Liao et al., December 2025)
335 ▪ OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning (Zhu et al., December 2025)
336 ▪ COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers (Wang et al., December 2025)
337 ▪ Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI (Sandoval and Rushing, December 2025)
338 ▪ Large Language Model based Smart Contract Auditing with LLMBugScanner (Yuan et al., December 2025)
339 ▪ Teleportation-Based Defenses for Privacy in Approximate Machine Unlearning (Maheri et al., December 2025)
340 ▪ Benchmarking and Understanding Safety Risks in AI Character Platforms (Wei, Zhang, and Tyson, December 2025)
341 ▪ Red Teaming Large Reasoning Models (Chen et al., December 2025)
342 ▪ An Empirical Study on the Security Vulnerabilities of GPTs (Wu, Wu, and Zheng, December 2025)
343 ▪ ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning (Chen, Kang, and Li, December 2025)
344 ▪ A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures (Kong et al., December 2025)
345 ▪ Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities (Garg et al., December 2025)
346 ▪ Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks (Young, December 2025)
347 ▪ Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance (Huwyler, December 2025)
348 ▪ GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision (Xiang et al., November 2025)
349 ▪ DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation (Pathak et al., November 2025)
350 ▪ VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., November 2025)
351 ▪ SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models (Alam et al., November 2025)
352 ▪ EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering (Gungor et al., November 2025)
353 ▪ Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization (Li et al., November 2025)
354 ▪ Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems (Wang et al., November 2025)
355 ▪ SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense (Krukowski et al., November 2025)
356 ▪ Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks (Ji et al., November 2025)
357 ▪ DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion (Liu et al., November 2025)
358 ▪ N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator (Lin et al., November 2025)
359 ▪ Certified but Fooled! Breaking Certified Defences with Ghost Certificates (Vo et al., November 2025)
360 ▪ ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning (Guan et al., November 2025)
361 ▪ AI Kill Switch for malicious web-based LLM agent (Lee and Park, November 2025)
362 ▪ DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt (Zhang et al., November 2025)
363 ▪ Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning (Zimmer and Karame, November 2025)
364 ▪ SGuard-v1: Safety Guardrail for Large Language Models (Lee et al., November 2025)
365 ▪ Defending Unauthorized Model Merging via Dual-Stage Weight Protection (Chen et al., November 2025)
366 ▪ On the Trade-Off Between Transparency and Security in Adversarial Machine Learning (Fenaux, Srinivasa, and Kerschbaum, November 2025)
367 ▪ TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone (Wang et al., November 2025)
368 ▪ InfoDecom: Decomposing Information for Defending against Privacy Leakage in Split Inference (Deng, Lu, and Duan, November 2025)
369 ▪ DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents (Zhang et al., November 2025)
370 ▪ SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization (Rong et al., November 2025)
371 ▪ AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework (Nathanson et al., November 2025)
372 ▪ VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization (Li, Yu, and Wang, November 2025)
373 ▪ Securing Generative AI in Healthcare: A Zero-Trust Architecture Powered by Confidential Computing on Google Cloud (Amanna and Shinde, November 2025)
374 ▪ PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities (Wei et al., November 2025)
375 ▪ Rethinking the Evaluation of Secure Code Generation (Dai, Xu, and Tao, November 2025)
376 ▪ AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness (Huang et al., November 2025)
377 ▪ iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification (Xiong et al., November 2025)
378 ▪ Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement (Ma et al., November 2025)
379 ▪ Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models (Zhang et al., November 2025)
380 ▪ Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models (Ahn and Jung, November 2025)
381 ▪ Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks (Chen et al., November 2025)
382 ▪ LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors (Lu et al., November 2025)
383 ▪ Efficient LLM Safety Evaluation through Multi-Agent Debate (Lin et al., November 2025)
384 ▪ Preserving security in a world with powerful AI Considerations for the future Defense Architecture (Generous, Cook, and Pruet, November 2025)
385 ▪ Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks (Waheed et al., November 2025)
386 ▪ ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (Gomaa, Salem, and Abdelnabi, November 2025)
387 ▪ Specification-Guided Vulnerability Detection with Large Language Models (Zhu et al., November 2025)
388 ▪ SME-TEAM: Leveraging Trust and Ethics for Secure and Responsible Use of AI and LLMs in SMEs (Sarker et al., November 2025)
389 ▪ LLM-Driven SAST-Genius: A Hybrid Static Analysis Framework for Comprehensive and Actionable Security (Agrawal and Ahi, November 2025)
390 ▪ SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories (Shen et al., November 2025)
391 ▪ LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models (Tahmasivand et al., November 2025)
392 ▪ RepoMark: A Data-Usage Auditing Framework for Code Large Language Models (Qu et al., November 2025)
393 ▪ AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding (Anwar and Liu, November 2025)
394 ▪ Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations (Okatan et al., November 2025)
395 ▪ DRIP: Defending Prompt Injection via De-instruction Training and Residual Fusion Model Architecture (Liu, Lin, and Dong, November 2025)
396 ▪ SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents (Yin et al., November 2025)
397 ▪ On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection (Hannan et al., November 2025)
398 ▪ SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation (Su et al., November 2025)
399 ▪ LLM-based Multi-class Attack Analysis and Mitigation Framework in IoT/IIoT Networks (Ikbarieh, Gupta, and Mahalal, November 2025)
400 ▪ The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs (Liu et al., October 2025)
401 ▪ Improving LLM Safety Alignment with Dual-Objective Optimization (Zhao et al., October 2025)
402 ▪ IRCopilot: Automated Incident Response with Large Language Models (Lin et al., October 2025)
403 ▪ ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models (Jin et al., October 2025)
404 ▪ Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control (Cai et al., October 2025)
405 ▪ SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning (Zhou et al., October 2025)
406 ▪ SoK: Honeypots & LLMs, More Than the Sum of Their Parts? (Bridges et al., October 2025)
407 ▪ OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models (Wang and Li, October 2025)
408 ▪ Secure Retrieval-Augmented Generation against Poisoning Attacks (Cheng et al., October 2025)
409 ▪ FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models (Nazari et al., October 2025)
410 ▪ SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability (Xu et al., October 2025)
411 ▪ Adversarially-Aware Architecture Design for Robust Medical AI Systems (Gerhart and Iyangar, October 2025)
412 ▪ Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents (Sanz-G\'omez et al., October 2025)
413 ▪ SAGE: A Generic Framework for LLM Safety Evaluation (Jindal et al., October 2025)
414 ▪ T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model (Zhang et al., October 2025)
415 ▪ SecureLearn - An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks (Paracha et al., October 2025)
416 ▪ Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs (Davies et al., October 2025)
417 ▪ SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots (Adebimpe, Neukirchen, and Welsh, October 2025)
418 ▪ FLAMES: Fine-tuning LLMs to Synthesize Invariants for Smart Contract Security (Eshghie et al., October 2025)
419 ▪ Soft Instruction De-escalation Defense (Walter et al., October 2025)
420 ▪ Towards Strong Certified Defense with Universal Asymmetric Randomization (Hong et al., October 2025)
421 ▪ Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses (Yichao et al., October 2025)
422 ▪ SAID: Empowering Large Language Models with Self-Activating Internal Defense (Chen et al., October 2025)
423 ▪ SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks (Lee et al., October 2025)
424 ▪ AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices (Zhan et al., October 2025)
425 ▪ Monitoring LLM-based Multi-Agent Systems Against Corruptions via Node Evaluation (Wu et al., October 2025)
426 ▪ Defending Against Prompt Injection with DataFilter (Wang et al., October 2025)
427 ▪ OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform (Wang and Li, October 2025)
428 ▪ Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization (Haddad et al., October 2025)
429 ▪ SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering (Lin et al., October 2025)
430 ▪ Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems (Jha et al., October 2025)
431 ▪ VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search (Li, Zhao, and Liu, October 2025)
432 ▪ Watermark Robustness and Radioactivity May Be at Odds in Federated Learning (Huang, Shao, and Baluta, October 2025)
433 ▪ Verifiable Fine-Tuning for LLMs: Zero-Knowledge Training Proofs Bound to Data Provenance and Policy (Akgul et al., October 2025)
434 ▪ SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection (Feng and Pan, October 2025)
435 ▪ Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers (Bertollo, Bodemir, and Burgess, October 2025)
436 ▪ A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks (Swanda et al., October 2025)
437 ▪ GuardReasoner: Towards Reasoning-based LLM Safeguards (Liu et al., October 2025)
438 ▪ Towards Proactive Defense Against Cyber Cognitive Attacks (Rushing, Umeokolo, and Xu, October 2025)
439 ▪ Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training (Wang et al., October 2025)
440 ▪ GRIDAI: Generating and Repairing Intrusion Detection Rules via Collaboration among Multiple LLM-based Agents (Li et al., October 2025)
441 ▪ Countermind: A Multi-Layered Security Architecture for Large Language Models (Schwarz, October 2025)
442 ▪ TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection (Liu et al., October 2025)
443 ▪ Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning (Liu et al., October 2025)
444 ▪ SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents (Ying et al., October 2025)
445 ▪ Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices (Patir et al., October 2025)
446 ▪ Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability (Mo et al., October 2025)
447 ▪ Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models (Meng et al., October 2025)
448 ▪ From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses (Meng et al., October 2025)
449 ▪ RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning (Horal et al., October 2025)
450 ▪ A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens (Dobre et al., October 2025)
451 ▪ DoomArena: A framework for Testing AI Agents Against Evolving Security Threats (Boisvert et al., October 2025)
452 ▪ VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation (Miculicich et al., October 2025)
453 ▪ Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling (Llewellyn et al., October 2025)
454 ▪ SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models (Qi et al., October 2025)
455 ▪ AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation (Liu and Anwar, October 2025)
456 ▪ Thought Purity: A Defense Framework For Chain-of-Thought Attack (Xue et al., October 2025)
457 ▪ SteerDiff: Steering towards Safe Text-to-Image Diffusion Models (Zhang, He, and Chen, October 2025)
458 ▪ Proactive defense against LLM Jailbreak (Zhao et al., October 2025)
459 ▪ MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection (Nguyen et al., October 2025)
460 ▪ SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition (Luo, Yau, and Song, October 2025)
461 ▪ Permissioned LLMs: Enforcing Access Control in Large Language Models (Jayaraman et al., October 2025)
462 ▪ A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory (Wei et al., October 2025)
463 ▪ Defend LLMs Through Self-Consciousness (Huang and Paula, October 2025)
464 ▪ UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models (Sun et al., October 2025)
465 ▪ TAIBOM: Bringing Trustworthiness to AI-Enabled Systems (Safronov et al., October 2025)
466 ▪ Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks (Palit, October 2025)
467 ▪ A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks (Hossain et al., October 2025)
468 ▪ A Call to Action for a Secure-by-Design Generative AI Paradigm (Alharthi and Garcia, October 2025)
469 ▪ MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement (Li et al., October 2025)
470 ▪ Detecting Instruction Fine-tuning Attacks on Language Models using Influence Function (Li, October 2025)
471 ▪ Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety (Zhang et al., October 2025)
472 ▪ Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation (Ndayipfukamiye et al., October 2025)
473 ▪ QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety (Lee et al., October 2025)
474 ▪ Federated Learning Resilient to Byzantine Attacks and Data Heterogeneity (Zuo et al., September 2025)
475 ▪ SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents (Dong et al., September 2025)
476 ▪ GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners (Li et al., September 2025)
477 ▪ Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting (Meng et al., September 2025)
478 ▪ ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search (Shen et al., September 2025)
479 ▪ Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment (Kim et al., September 2025)
480 ▪ Think Broad, Act Narrow: CWE Identification with Multi-Agent Large Language Models (Sayagh and Ghafari, August 2025)
481 ▪ ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models (Wang et al., August 2025)
482 ▪ Provably Secure Retrieval-Augmented Generation (Zhou, Feng, and Yang, August 2025)
483 ▪ FedGuard: A Diverse-Byzantine-Robust Mechanism for Federated Learning with Major Malicious Clients (Jiang et al., August 2025)
484 ▪ CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation (Yang et al., August 2025)
485 ▪ Bridging Privacy and Robustness for Trustworthy Machine Learning (Zhang and Chen, July 2025)
486 ▪ Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems (Sharshar et al., July 2025)
487 ▪ Strategic Deflection: Defending LLMs from Logit Manipulation (Rachidy et al., July 2025)
488 ▪ Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security (Dai et al., July 2025)
489 ▪ GUARD-CAN: Graph-Understanding and Recurrent Architecture for CAN Anomaly Detection (Kim and Kim, July 2025)
490 ▪ SDD: Self-Degraded Defense against Malicious Fine-tuning (Chen et al., July 2025)
491 ▪ OneShield -- the Next Generation of LLM Guardrails (DeLuca et al., July 2025)
492 ▪ Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards (Madhavan et al., July 2025)
493 ▪ Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs (Camporese and Massacci, July 2025)
494 ▪ PurpCode: Reasoning for Safer Code Generation (Liu et al., July 2025)
495 ▪ SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysis (Lingxiang et al., July 2025)
496 ▪ Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment (Li et al., July 2025)
497 ▪ CASCADE: LLM-Powered JavaScript Deobfuscator at Google (Jiang et al., July 2025)
498 ▪ LLM Meets the Sky: Heuristic Multi-Agent Reinforcement Learning for Secure Heterogeneous UAV Networks (Zheng et al., July 2025)
499 ▪ CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems (Hu et al., July 2025)
500 ▪ DP-TLDM: Differentially Private Tabular Latent Diffusion Model (Zhu et al., July 2025)
501 ▪ Recent Advances in Malware Detection: Graph Learning and Explainability (Shokouhinejad et al., July 2025)
502 ▪ "We Need a Standard": Toward an Expert-Informed Privacy Label for Differential Privacy (Dibia et al., July 2025)
503 ▪ ACFIX: Guiding LLMs with Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart Contracts (Zhang et al., July 2025)
504 ▪ OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting (Cheng et al., July 2025)
505 ▪ Defending Against Unforeseen Failure Modes with Latent Adversarial Training (Casper et al., July 2025)
506 ▪ FedStrategist: A Meta-Learning Framework for Adaptive and Robust Aggregation in Federated Learning (Haque, Kamal, and Hossain, July 2025)
507 ▪ PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants (Liu et al., July 2025)
508 ▪ PromptArmor: Simple yet Effective Prompt Injection Defenses (Shi et al., July 2025)
509 ▪ A Privacy-Centric Approach: Scalable and Secure Federated Learning Enabled by Hybrid Homomorphic Encryption (Nguyen, Khan, and Michalas, July 2025)
510 ▪ PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training (Du, July 2025)
511 ▪ Defense Against Prompt Injection Attack by Leveraging Attack Techniques (Chen et al., July 2025)
512 ▪ GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention (Abdalla et al., July 2025)
513 ▪ Risks of ignoring uncertainty propagation in AI-augmented security pipelines (Mezzi et al., July 2025)
514 ▪ JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model (Nian et al., July 2025)
515 ▪ SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks (Uddin et al., July 2025)
516 ▪ Safeguarding Federated Learning-based Road Condition Classification (Liu and Papadimitratos, July 2025)
517 ▪ Thought Purity: Defense Paradigm For Chain-of-Thought Attack (Xue et al., July 2025)
518 ▪ Expanding ML-Documentation Standards For Better Security (Appel, July 2025)
519 ▪ A Generative Approach to LLM Harmfulness Detection with Special Red Flag Tokens (Xhonneux et al., July 2025)
520 ▪ A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation (Trudslev et al., July 2025)
521 ▪ From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection (Sun et al., July 2025)
522 ▪ Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data (Alikhani and Kazemi, July 2025)
523 ▪ Defense-as-a-Service: Black-box Shielding against Backdoored Graph Models (Yang et al., July 2025)
524 ▪ Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators (Xue et al., July 2025)
525 ▪ SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework (Brito et al., July 2025)
526 ▪ AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions (Liu et al., July 2025)
527 ▪ Entangled Threats: A Unified Kill Chain Model for Quantum Machine Learning Security (Debus et al., July 2025)
528 ▪ ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware Detection (Alam, Piplai, and Rastogi, July 2025)
529 ▪ Beyond the Worst Case: Extending Differential Privacy Guarantees to Realistic Adversaries (Swanberg et al., July 2025)
530 ▪ A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning (Gluch and Goldwasser, July 2025)
531 ▪ Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences (Lyu et al., July 2025)
532 ▪ Adversarial Defenses via Vector Quantization (Dong and Mao, July 2025)
533 ▪ Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking (Wu et al., July 2025)
534 ▪ Defending Against Prompt Injection With a Few DefensiveTokens (Chen et al., July 2025)
535 ▪ Can Large Language Models Improve Phishing Defense? A Large-Scale Controlled Experiment on Warning Dialogue Explanations (Cau et al., July 2025)
536 ▪ Hybrid LLM-Enhanced Intrusion Detection for Zero-Day Threats in IoT Networks (Al-Hammouri et al., July 2025)
537 ▪ Saffron-1: Safety Inference Scaling (Qiu et al., July 2025)
538 ▪ LDP$^3$: An Extensible and Multi-Threaded Toolkit for Local Differential Privacy Protocols and Post-Processing Methods (Balioglu, Khodaie, and Gursoy, July 2025)
539 ▪ TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data (Cheruvu et al., July 2025)
540 ▪ PROTEAN: Federated Intrusion Detection in Non-IID Environments through Prototype-Based Knowledge Sharing (Chennoufi et al., July 2025)
541 ▪ MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamin<br><br> (Guo et al., May 2025)
542 ▪ Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration<br><br> (Tsmindashvili et al., May 2025)
543 ▪ LLM-Based Threat Detection and Prevention Framework for IoT Ecosystems<br><br> (Otoum, Asad, and Nayak, May 2025)
544 ▪ Securing RAG: A Risk Assessment and Mitigation Framework<br><br> (Ammann et al., May 2025)
545 ▪ Large Language Model Sentinel: LLM Agent for Adversarial Purification (Lin, Tanaka, and Zhao, Apr 2025)
546 ▪ JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks (Zhang et al, Mar 2025)
Defenses

>

<

‍

Threat to AI Models

General Approaches

We added this subsection to cover research that broadly looks at AI security.

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : General Approaches

cybersecurity_tracker - Google Drive

2 ▪ Talk is (Not) Cheap: A Taxonomy and Benchmark Coverage Audit for LLM Attacks (Iyer et al., May 2026)
3 ▪ From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World (Conde et al., May 2026)
4 ▪ Threat Modelling using Domain-Adapted Language Models: Empirical Evaluation and Insights (Pourhanifeh, AbdulGhaffar, and Matrawy, May 2026)
5 ▪ CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios (Lim et al., May 2026)
6 ▪ Evaluating the Reliability of Multiple Large Language Models in Risk Assessment: A CIS Controls Based Approach (Pinto, Labaki, and Miani, May 2026)
7 ▪ QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing (Ameen, Alam, and Islam, May 2026)
8 ▪ GAMMAF: A Common Framework for Graph-Based Anomaly Monitoring Benchmarking in LLM Multi-Agent Systems (Mateo-Torrej\'on and S\'anchez-Maci\'an, April 2026)
9 ▪ Training a General Purpose Automated Red Teaming Model (Padmakumar et al., April 2026)
10 ▪ CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents (Fu et al., April 2026)
11 ▪ Synthesizing Multi-Agent Harnesses for Vulnerability Discovery (Liu et al., April 2026)
12 ▪ WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks (Ramesh et al., April 2026)
13 ▪ SkillAttack: Automated Red Teaming of Agent Skills through Attack Path Refinement (Duan et al., April 2026)
14 ▪ Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities (Mouzouni, April 2026)
15 ▪ AI Security in the Foundation Model Era: A Comprehensive Survey from a Unified Perspective (Wang and Luan, March 2026)
16 ▪ Network- and Device-Level Cyber Deception for Contested Environments Using RL and LLMs (Sahu, Paul, and Macwan, March 2026)
17 ▪ ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models (Owiredu-Ashley, March 2026)
18 ▪ CyberThreat-Eval: Can Large Language Models Automate Real-World Threat Research? (Chen et al., March 2026)
19 ▪ SecureRAG-RTL: A Retrieval-Augmented, Multi-Agent, Zero-Shot LLM-Driven Framework for Hardware Vulnerability Detection (Hasan et al., March 2026)
20 ▪ From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architectures (Bonfanti et al., March 2026)
21 ▪ What Makes a Good LLM Agent for Real-world Penetration Testing? (Deng et al., February 2026)
22 ▪ Security Threat Modeling for Emerging AI-Agent Protocols: A Comparative Analysis of MCP, A2A, Agora, and ANP (Anbiaee et al., February 2026)
23 ▪ LLMs + Security = Trouble (Livshits, February 2026)
24 ▪ Efficient LLM Moderation with Multi-Layer Latent Prototypes (Chrab\k{a}szcz et al., February 2026)
25 ▪ TrajAD: Trajectory Anomaly Detection for Trustworthy LLM Agents (Liu et al., February 2026)
26 ▪ From Sands to Mansions: Towards Automated Cyberattack Emulation with Classical Planning and Large Language Models (Wang et al., February 2026)
27 ▪ Risk Assessment and Security Analysis of Large Language Models (Zhang, Lyu, and Li, February 2026)
28 ▪ LogicScan: An LLM-driven Framework for Detecting Business Logic Vulnerabilities in Smart Contracts (Gao et al., February 2026)
29 ▪ RedSage: A Cybersecurity Generalist LLM (Suryanto et al., January 2026)
30 ▪ Toward Risk Thresholds for AI-Enabled Cyber Threats: Enhancing Decision-Making Under Uncertainty with Bayesian Networks (Jackson et al., January 2026)
31 ▪ LLMs, You Can Evaluate It! Design of Multi-perspective Report Evaluation for Security Operation Centers (Okada, Oba, and Yanai, January 2026)
32 ▪ Enhancing Cloud Network Resilience via a Robust LLM-Empowered Multi-Agent Reinforcement Learning Framework (Peng et al., January 2026)
33 ▪ CyberLLM-FINDS 2025: Instruction-Tuned Fine-tuning of Domain-Specific LLMs with Retrieval-Augmented Generation and Graph Integration for MITRE Evaluation (Iyer, Bobadilla, and Iyengar, January 2026)
34 ▪ Cybersecurity AI: A Game-Theoretic AI for Guiding Attack and Defense (Mayoral-Vilches et al., January 2026)
35 ▪ Automated Red-Teaming Framework for Large Language Model Security Assessment: A Comprehensive Attack Generation and Detection System (Wei et al., December 2025)
36 ▪ Bounty Hunter: Autonomous, Comprehensive Emulation of Multi-Faceted Adversaries (Hackl\"ander-Jansen, Uetz, and Henze, December 2025)
37 ▪ PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design (Yang et al., December 2025)
38 ▪ Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing (Lin et al., December 2025)
39 ▪ SoK: a Comprehensive Causality Analysis Framework for Large Language Model Security (Zhao, Li, and Sun, December 2025)
40 ▪ Incalmo: An Autonomous LLM-assisted System for Red Teaming Multi-Host Networks (Singer et al., November 2025)
41 ▪ SoK: The Security-Safety Continuum of Multimodal Foundation Models through Information Flow and Global Game-Theoretic Analysis of Asymmetric Threats (Sun et al., November 2025)
42 ▪ Large Language Models for Cyber Security (Somani and Cherukuri, November 2025)
43 ▪ The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies (He et al., November 2025)
44 ▪ Adapting Large Language Models to Emerging Cybersecurity using Retrieval Augmented Generation (Borah, Alam, and Rastogi, November 2025)
45 ▪ Ask What Your Country Can Do For You: Towards a Public Red Teaming Model (Kennedy et al., October 2025)
46 ▪ SoK: Taxonomy and Evaluation of Prompt Security in Large Language Models (Hong et al., October 2025)
47 ▪ BlackIce: A Containerized Red Teaming Toolkit for AI Security Testing (Kaplan, Warnecke, and Archibald, October 2025)
48 ▪ Frontier AI's Impact on the Cybersecurity Landscape (Potter et al., October 2025)
49 ▪ PACEbench: A Framework for Evaluating Practical AI Cyber-Exploitation Capabilities (Liu et al., October 2025)
50 ▪ Living Off the LLM: How LLMs Will Change Adversary Tactics (Oesch et al., October 2025)
51 ▪ CREST-Search: Comprehensive Red-teaming for Evaluating Safety Threats in Large Language Models Powered by Web Search (Ou et al., October 2025)
52 ▪ MCPSecBench: A Systematic Security Benchmark and Playground for Testing Model Context Protocols (Yang, Wu, and Chen, October 2025)
53 ▪ Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions (Hou et al., October 2025)
54 ▪ Quantifying Risks in Multi-turn Conversation with Large Language Models (Wang et al., October 2025)
55 ▪ LLAMAFUZZ: Large Language Model Enhanced Greybox Fuzzing (Zhang et al., October 2025)
56 ▪ MALF: A Multi-Agent LLM Framework for Intelligent Fuzzing of Industrial Control Protocols (Ning, Zong, and He, October 2025)
57 ▪ Quant Fever, Reasoning Blackholes, Schrodinger's Compliance, and More: Probing GPT-OSS-20B (Lin et al., September 2025)
58 ▪ Generative AI-Empowered Secure Communications in Space-Air-Ground Integrated Networks: A Survey and Tutorial (Hu et al., August 2025)
59 ▪ AI-Driven Cybersecurity Threat Detection: Building Resilient Defense Systems Using Predictive Analytics (Das et al., August 2025)
60 ▪ BlockA2A: Towards Secure and Verifiable Agent-to-Agent Interoperability (Zou et al., August 2025)
61 ▪ From Cloud-Native to Trust-Native: A Protocol for Verifiable Multi-Agent Systems (Li, July 2025)
62 ▪ Prompt Optimization and Evaluation for LLM Automated Red Teaming (Freenor et al., July 2025)
63 ▪ Interpretable Anomaly-Based DDoS Detection in AI-RAN with XAI and LLMs (Chatzimiltis et al., July 2025)
64 ▪ Towards Unifying Quantitative Security Benchmarking for Multi Agent Systems (Sharma et al., July 2025)
65 ▪ Leveraging Trustworthy AI for Automotive Security in Multi-Domain Operations: Towards a Responsive Human-AI Multi-Domain Task Force for Cyber Social Security (Barletta et al., July 2025)
66 ▪ Security practices in AI development (Spelda and Stritecky, July 2025)
67 ▪ PRISM: A Personalized, Rapid, and Immersive Skill Mastery framework for personalizing experiential learning through Generative AI (Lin et al., July 2025)
68 ▪ Risks & Benefits of LLMs & GenAI for Platform Integrity, Healthcare Diagnostics, Financial Trust and Compliance, Cybersecurity, Privacy & AI Safety: A Comprehensive Survey, Roadmap & Implementation Blueprint (Ahi, July 2025)
69 ▪ A Large Language Model-Supported Threat Modeling Framework for Transportation Cyber-Physical Systems (Salek et al., July 2025)
70 ▪ Policy-Driven AI in Dataspaces: Taxonomy, Explainability, and Pathways for Compliant Innovation (Chandra and Navneet, July 2025)
71 ▪ Regression-aware Continual Learning for Android Malware Detection (Ghiani et al., July 2025)
72 ▪ Your ATs to Ts: MITRE ATT&CK Attack Technique to P-SSCRM Task Mapping (Hamer et al., July 2025)
73 ▪ Information Security Based on LLM Approaches: A Review (Gong, Li, and Li, July 2025)
74 ▪ Optimizing Privacy-Utility Trade-off in Decentralized Learning with Generalized Correlated Noise (Rodio, Chen, and Larsson, July 2025)
75 ▪ Enabling Cyber Security Education through Digital Twins and Generative AI (Barletta et al., July 2025)
76 ▪ Threshold-Protected Searchable Sharing: Privacy Preserving Aggregated-ANN Search for Collaborative RAG (Guo, July 2025)
77 ▪ Large Language Models are Autonomous Cyber Defenders (Castro et al., July 2025)
78 ▪ Enabling Efficient Attack Investigation via Human-in-the-Loop Security Analysis (Tsegai et al., July 2025)
79 ▪ Adaptive Network Security Policies via Belief Aggregation and Rollout (Hammar et al., July 2025)
80 ▪ Learning to Communicate in Multi-Agent Reinforcement Learning for Autonomous Cyber Defence (Contractor, Li, and Mallah, July 2025)
81 ▪ ExCyTIn-Bench: Evaluating LLM agents on Cyber Threat Investigation (Wu et al., July 2025)
82 ▪ Large Language Models in Cybersecurity: Applications, Vulnerabilities, and Defense Techniques (Jaffal, Alkhanafseh, and Mohaisen, July 2025)
83 ▪ Challenges in GenAI and Authentication: a scoping review (Bezerra, Bezerra, and Westphall, July 2025)
84 ▪ Towards Effective Complementary Security Analysis using Large Language Models (Wagner et al., July 2025)
85 ▪ PREAMBLE: Private and Efficient Aggregation via Block Sparse Vectors (Asi et al., July 2025)
86 ▪ Domain Borders Are There to Be Crossed With Federated Few-Shot Adaptation (R\"oder, Raab, and Schleif, July 2025)
87 ▪ ARBoids: Adaptive Residual Reinforcement Learning With Boids Model for Cooperative Multi-USV Target Defense (Tao et al., July 2025)
88 ▪ Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs) (Verma et al., July 2025)
89 ▪ BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems (Zhang et al., July 2025)
90 ▪ Autonomous AI-based Cybersecurity Framework for Critical Infrastructure: Real-Time Threat Mitigation (Paulraj et al., July 2025)
91 ▪ Automated Attack Testflow Extraction from Cyber Threat Report using BERT for Contextual Analysis (Ahmadou et al., July 2025)
92 ▪ Automated Reasoning for Vulnerability Management by Design (Shaked and Messe, July 2025)
93 ▪ LLMs on support of privacy and security of mobile apps: state of the art and research directions (Nguyen, Carminati, and Ferrari, July 2025)
94 ▪ Red Teaming AI Red Teaming (Majumdar, Pendleton, and Gupta, July 2025)
95 ▪ Taming Data Challenges in ML-based Security Tasks: Lessons from Integrating Generative AI (Kanchi et al., July 2025)
96 ▪ Phare: A Safety Probe for Large Language Models (Le Jeune et al, May 2025)
97 ▪ aiXamine: Simplified LLM Safety and Security (Deniz et al, Jun 2025)
98 ▪ Safety at Scale: A Comprehensive Survey of Large Model Safety (Ma et al, May 2025)
99 ▪ Emerging Security Challenges of Large Language Models (Debar et al, Dec 2024)
100 ▪ Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs (Pasini et al, Nov 2024)
101 ▪ Blockchain for Large Language Model Security and Safety: A Holistic Survey (Geren et al, Nov 2024)
102 ▪ One Prompt to Verify Your Models: Black-Box Text-to-Image Models Verification via Non-Transferable Adversarial Attacks (Guo et al, Nov 2024)
103 ▪ How (un)ethical are instruction-centric responses of LLMs? Unveiling the vulnerabilities of safety guardrails to harmful queries (Banerjee et al, Nov 2024)
104 ▪ Defending Large Language Models Against Attacks With Residual Stream Activation Analysis (Kawasaki, Davis, and Abbas, Nov 2024)
105 ▪ LProtector: An LLM-driven Vulnerability Detection System (Sheng et al, Nov 2024)
106 ▪ SECURE: Benchmarking Large Language Models for Cybersecurity (Bhusal et al, Oct 2024)
107 ▪ Insights and Current Gaps in Open-Source LLM Vulnerability Scanners: A Comparative Analysis (Brokman et al, Oct 2024)
108 ▪ Safety Layers in Aligned Large Language Models: The Key to LLM Security (Li et al, Oct 2024)
109 ▪ AttackBench: Evaluating Gradient-based Attacks for Adversarial Examples (Cinà et al, Oct 2024)
110 ▪ Advancing Cyber Incident Timeline Analysis Through Rule Based AI and Large Language Models (Loumachi and Ghanem, Sep 2024)
111 ▪ Real-world Adversarial Defense against Patch Attacks based on Diffusion Model (Wei et al, Sep 2024)
112 ▪ LLM Honeypot: Leveraging Large Language Models as Advanced Interactive Honeypot Systems (Otal and Canbaz, Sep 2024)
113 ▪ SECURE: Benchmarking Large Language Models for Cybersecurity Advisory (Bhusal et al, Sep 2024)
114 ▪ Continual Adversarial Defense (Wang et al, Aug 2024)
115 ▪ Audit-LLM: Multi-Agent Collaboration for Log-based Insider Threat Detection (Song et al, Aug 2024)
116 ▪ Attacks and Defenses for Generative Diffusion Models: A Comprehensive Survey (Truong, Dan, and Le, Aug 2024)
117 ▪ Blockchain for Large Language Model Security and Safety: A Holistic Survey (Geren et al, Jul 2024)
General Approaches

>

<

‍

Data Poisoning and Simulated Publication of Poisoned Public Datasets

Covers:

  • OWASP LLM 03: Training Data Poisoning
  • OWASP ML 02: Data Poisoning Attack
  • MITRE ATLAS Resource Development

We moved this subsection from 'Threats Using AI Models' to this section as poisoned data is a threat to AI.

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Data Poisoning and Simulated Publication of Poisoned Public Datasets

cybersecurity_tracker - Google Drive

2 ▪ DSTAN-Med: Dual-Channel Spatiotemporal Attention with Physiological Plausibility Filtering for False Data Injection Attack Detection in IoT-Based Medical Devices (Hasan, Islam, and Hossain, May 2026)
3 ▪ SoK: Unlearnability and Unlearning for Model Dememorization (Zhang et al., May 2026)
4 ▪ Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing (Mao et al., May 2026)
5 ▪ Generate "Normal", Edit Poisoned: Branding Injection via Hint Embedding in Image Editing (Sun et al., May 2026)
6 ▪ Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation (Yang et al., May 2026)
7 ▪ Oracle Poisoning: Corrupting Knowledge Graphs to Weaponise AI Agent Reasoning (Kereopa-Yorke et al., May 2026)
8 ▪ ShadowMerge: A Novel Poisoning Attack on Graph-Based Agent Memory via Relation-Channel Conflicts (Luo et al., May 2026)
9 ▪ When Routine Chats Turn Toxic: Unintended Long-Term State Poisoning in Personalized Agents (Xu et al., May 2026)
10 ▪ Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection (Zeng et al., May 2026)
11 ▪ Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms (Wang et al., May 2026)
12 ▪ LoopTrap: Termination Poisoning Attacks on LLM Agents (Xu et al., May 2026)
13 ▪ Architecture Matters: Comparing RAG Systems under Knowledge Base Poisoning (Korn, May 2026)
14 ▪ Gray-Box Poisoning of Continuous Malware Ingestion Pipelines (Dolej\v{s}, Jure\v{c}ek, and L\'orencz, May 2026)
15 ▪ MEMSAD: Gradient-Coupled Anomaly Detection for Memory Poisoning in Retrieval-Augmented Agents (California and Berkeley), May 2026)
16 ▪ Adversarial Update-Based Federated Unlearning for Poisoned Model Recovery (Zhao et al., May 2026)
17 ▪ Fight Poison with Poison: Enhancing Robustness in Few-shot Machine-Generated Text Detection with Adversarial Training (Duan, Zhou, and Li, May 2026)
18 ▪ Repurposing and Evaluating the (In)Feasibility of Dataset Poisoning enabled Watermarking for Contrastive Learning (Dai et al., May 2026)
19 ▪ Needle-in-RAG: Prompt-Conditioned Character-Level Traceback of Poisoned Spans in Retrieved Evidence (Cui and Liu, May 2026)
20 ▪ From Stealthy Data Fabrication to Unsafe Driving: Realistic Scenario Attacks on Collaborative Perception (Zhang, Zhang, and Mao, May 2026)
21 ▪ Repurposing Image Diffusion Models for Adversarial Synthetic Structured Data: A Case Study of Ground Truth Drift (Arthur and Schwartz, May 2026)
22 ▪ Defense against Poisoning Attacks under Shuffle-DP (Wang et al., May 2026)
23 ▪ CleanBase: Detecting Malicious Documents in RAG Knowledge Databases (Jin et al., May 2026)
24 ▪ SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation (Rezakhani et al., May 2026)
25 ▪ Poisoning Learned Index Structures: Static and Dynamic Adversarial Attacks on ALEX (Jue, April 2026)
26 ▪ RouteGuard: Internal-Signal Detection of Skill Poisoning in LLM Agents (Xiao et al., April 2026)
27 ▪ Train in Vain: Functionality-Preserving Poisoning to Prevent Unauthorized Use of Code Datasets (Xiao et al., April 2026)
28 ▪ CSC: Turning the Adversary's Poison against Itself (Shi et al., April 2026)
29 ▪ From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning (Xu et al., April 2026)
30 ▪ Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories (Bercovich et al., April 2026)
31 ▪ Visual Inception: Compromising Long-term Planning in Agentic Recommenders via Multimodal Memory Poisoning (Qian, April 2026)
32 ▪ Privacy-Aware Machine Unlearning with SISA for Reinforcement Learning-Based Ransomware Detection (Ferdous, Islam, and Islam, April 2026)
33 ▪ FedIDM: Achieving Fast and Stable Convergence in Byzantine Federated Learning through Iterative Distribution Matching (Yang et al., April 2026)
34 ▪ Robustness Analysis of Machine Learning Models for IoT Intrusion Detection Under Data Poisoning Attacks (Wulnye et al., April 2026)
35 ▪ PatchPoison: Poisoning Multi-View Datasets to Degrade 3D Reconstruction (Wadekar et al., April 2026)
36 ▪ Poisoning with A Pill: Circumventing Detection in Federated Learning (Guo et al., April 2026)
37 ▪ DuCodeMark: Dual-Purpose Code Dataset Watermarking via Style-Aware Watermark-Poison Design (Chen et al., April 2026)
38 ▪ XFED: Non-Collusive Model Poisoning Attack Against Byzantine-Robust Federated Classifiers (Mouri, Ridowan, and Adnan, April 2026)
39 ▪ One Shot Dominance: Knowledge Poisoning Attack on Retrieval-Augmented Generation Systems (Chang et al., April 2026)
40 ▪ TRUSTDESC: Preventing Tool Poisoning in LLM Applications via Trusted Description Generation (Ye et al., April 2026)
41 ▪ RefineRAG: Word-Level Poisoning Attacks via Retriever-Guided Text Refinement (Wang, Wang, and Wang, April 2026)
42 ▪ FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization (Zhu et al., April 2026)
43 ▪ Can You Trust the Vectors in Your Vector Database? Black-Hole Attack from Embedding Space Defects (Li et al., April 2026)
44 ▪ Poisoned Identifiers Survive LLM Deobfuscation: A Case Study on Claude Opus 4.6 (Lorenzo, April 2026)
45 ▪ Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents (Zou et al., April 2026)
46 ▪ Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation (Zarkadis and Douligeris, April 2026)
47 ▪ S-DAPT-2026: A Stage-Aware Synthetic Dataset for Advanced Persistent Threat Detection (Tijjani et al., April 2026)
48 ▪ Towards Explainable Privacy Preservation in Federated Learning via Shapley Value-Guided Noise Injection (Li, Gui, and Wu, April 2026)
49 ▪ RAGShield: Provenance-Verified Defense-in-Depth Against Knowledge Base Poisoning in Government Retrieval-Augmented Generation Systems (Patil, April 2026)
50 ▪ Unveiling the Security Risks of Federated Learning in the Wild: From Research to Practice (Chen et al., March 2026)
51 ▪ Memory poisoning and secure multi-agent systems (Torra and Bras-Amor\'os, March 2026)
52 ▪ A Model Consistency-Based Countermeasure to GAN-Based Data Poisoning Attack in Federated Learning (Sun et al., March 2026)
53 ▪ FedTrident: Resilient Road Condition Classification Against Poisoning Attacks in Federated Learning (Liu and Papadimitratos, March 2026)
54 ▪ Semantic Chameleon: Corpus-Dependent Poisoning Attacks and Defenses in RAG Systems (Thornton, March 2026)
55 ▪ Towards Unsupervised Adversarial Document Detection in Retrieval Augmented Generation Systems (Levi, March 2026)
56 ▪ Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented Scanning (Yan et al., March 2026)
57 ▪ KEPo: Knowledge Evolution Poison on Graph-based Retrieval-Augmented Generation (Chen et al., March 2026)
58 ▪ Adversarial Hubness Detector: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems (Habler et al., March 2026)
59 ▪ SPOILER: TEE-Shielded DNN Partitioning of On-Device Secure Inference with Poison Learning (Kang et al., March 2026)
60 ▪ SuperLocalMemory: Privacy-Preserving Multi-Agent Memory with Bayesian Trust Defense Against Memory Poisoning (Bhardwaj, March 2026)
61 ▪ Silent Sabotage During Fine-Tuning: Few-Shot Rationale Poisoning of Compact Medical LLMs (Xie et al., March 2026)
62 ▪ Aggressive or Imperceptible, or Both: Network Pruning Assisted Hybrid Byzantines in Federated Learning (Ozfatura et al., March 2026)
63 ▪ Sparsification Under Siege: Dual-Level Defense Against Poisoning in Communication-Efficient Federated Learning (Jin et al., March 2026)
64 ▪ Turning Black Box into White Box: Dataset Distillation Leaks (Chen et al., March 2026)
65 ▪ Hidden in the Metadata: Stealth Poisoning Attacks on Multimodal Retrieval-Augmented Generation (Edemacu and Shokri, March 2026)
66 ▪ Enhancing Continual Learning for Software Vulnerability Prediction: Addressing Catastrophic Forgetting via Hybrid-Confidence-Aware Selective Replay for Temporal LLM Fine-Tuning (Dou, Bahsi, and Guerra-Manzanares, March 2026)
67 ▪ HubScan: Detecting Hubness Poisoning in Retrieval-Augmented Generation Systems (Habler et al., February 2026)
68 ▪ Poisoned Acoustics (Dahme, February 2026)
69 ▪ PenTiDef: Enhancing Privacy and Robustness in Decentralized Federated Intrusion Detection Systems against Poisoning Attacks (Duy et al., February 2026)
70 ▪ Intent Laundering: AI Safety Datasets Are Not What They Seem (Golchin and Wetter, February 2026)
71 ▪ SRFed: Mitigating Poisoning Attacks in Privacy-Preserving Federated Learning with Heterogeneous Data (Lu, February 2026)
72 ▪ Efficient Semi-Supervised Adversarial Training via Latent Clustering-Based Data Reduction (Ghosh, Xu, and Zhang, February 2026)
73 ▪ Closing the Distribution Gap in Adversarial Training for LLMs (Hu et al., February 2026)
74 ▪ Towards Privacy-Guaranteed Label Unlearning in Vertical Federated Learning: Few-Shot Forgetting without Disclosure (Gu et al., February 2026)
75 ▪ One RNG to Rule Them All: How Randomness Becomes an Attack Vector in Machine Learning (Prabhu, Gan, and Ghodsi, February 2026)
76 ▪ Confundo: Learning to Generate Robust Poison for Practical RAG Systems (Hu et al., February 2026)
77 ▪ VENOMREC: Cross-Modal Interactive Poisoning for Targeted Promotion in Multimodal LLM Recommender Systems (Guan et al., February 2026)
78 ▪ ADCA: Attention-Driven Multi-Party Collusion Attack in Federated Self-Supervised Learning (Wang et al., February 2026)
79 ▪ Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoning (Draganov et al., February 2026)
80 ▪ Statistical MIA: Rethinking Membership Inference Attack for Reliable Unlearning Auditing (Sun et al., February 2026)
81 ▪ Spattack: Subgroup Poisoning Attacks on Federated Recommender Systems (Yan et al., February 2026)
82 ▪ Stealthy Poisoning Attacks Bypass Defenses in Regression Settings (Carnerero-Cano et al., February 2026)
83 ▪ On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression (Zhang et al., January 2026)
84 ▪ Robust Federated Learning for Malicious Clients using Loss Trend Deviation Detection (Bhaskar, B, and P, January 2026)
85 ▪ Thought-Transfer: Indirect Targeted Poisoning Attacks on Chain-of-Thought Reasoning Models (Chaudhari et al., January 2026)
86 ▪ Attacks on Approximate Caches in Text-to-Image Diffusion Models (Sun, Jie, and Liu, January 2026)
87 ▪ Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion Models (Liu et al., January 2026)
88 ▪ Less Is More -- Until It Breaks: Security Pitfalls of Vision Token Compression in Large Vision-Language Models (Zhang et al., January 2026)
89 ▪ Distributional Machine Unlearning via Selective Data Removal (Allouah, Guerraoui, and Koyejo, January 2026)
90 ▪ MindGuard: Intrinsic Decision Inspection for Securing LLM Agents Against Metadata Poisoning (Wang et al., January 2026)
91 ▪ MCP-ITP: An Automated Framework for Implicit Tool Poisoning in MCP (Li et al., January 2026)
92 ▪ Memory Poisoning Attack and Defense on Memory Based LLM-Agents (Sunil et al., January 2026)
93 ▪ Practical Poisoning Attacks against Retrieval-Augmented Generation (Zhang et al., January 2026)
94 ▪ Knowledge-to-Data: LLM-Driven Synthesis of Structured Network Traffic for Testbed-Free IDS Evaluation (Kampourakis et al., January 2026)
95 ▪ Shadow Unlearning: A Neuro-Semantic Approach to Fidelity-Preserving Faceless Forgetting in LLMs (P et al., January 2026)
96 ▪ VFEFL: Privacy-Preserving Federated Learning against Malicious Clients via Verifiable Functional Encryption (Cai, Han, and Meng, January 2026)
97 ▪ Quality Degradation Attack in Synthetic Data (Liu et al., January 2026)
98 ▪ Low Rank Comes with Low Security: Gradient Assembly Poisoning Attacks against Distributed LoRA-based LLM Systems (Dong et al., January 2026)
99 ▪ Learning from Negative Examples: Why Warning-Framed Training Data Teaches What It Warns Against (Enkhbayar, December 2025)
100 ▪ Certifying the Right to Be Forgotten: Primal-Dual Optimization for Sample and Label Unlearning in Vertical Federated Learning (Jiang et al., December 2025)
101 ▪ Look Twice before You Leap: A Rational Agent Framework for Localized Adversarial Anonymization (Duan et al., December 2025)
102 ▪ GShield: Mitigating Poisoning Attacks in Federated Learning (M. et al., December 2025)
103 ▪ The Erasure Illusion: Stress-Testing the Generalization of LLM Forgetting Evaluation (Jia et al., December 2025)
104 ▪ SecureCode v2.0: A Production-Grade Dataset for Training Security-Aware Code Generation Models (Thornton, December 2025)
105 ▪ A Certified Unlearning Approach without Access to Source Data (Basaran et al., December 2025)
106 ▪ Trust Me, I Know This Function: Hijacking LLM Static Analysis using Bias (Bernstein et al., December 2025)
107 ▪ Bits for Privacy: Evaluating Post-Training Quantization via Membership Inference (Zhang et al., December 2025)
108 ▪ Reasoning-Style Poisoning of LLM Agents via Stealthy Style Transfer: Process-Level Attacks and Runtime Monitoring in RSV Space (Zhou and Wang, December 2025)
109 ▪ Evaluating Adversarial Attacks on Federated Learning for Temperature Forecasting (Chichifoi, Merizzi, and Colajanni, December 2025)
110 ▪ CLOAK: Contrastive Guidance for Latent Diffusion-Based Data Obfuscation (Yang and Ardakanian, December 2025)
111 ▪ Deferred Poisoning: Making the Model More Vulnerable via Hessian Singularization (He et al., December 2025)
112 ▪ Hard Work Does Not Always Pay Off: Poisoning Attacks on Neural Architecture Search (Coalson et al., December 2025)
113 ▪ MIRAGE: Misleading Retrieval-Augmented Generation via Black-box and Query-agnostic Poisoning Attacks (Chen et al., December 2025)
114 ▪ Data Taggants: Dataset Ownership Verification via Harmless Targeted Data Poisoning (Bouaziz, Usunier, and El-Mhamdi, December 2025)
115 ▪ DEFEND: Poisoned Model Detection and Malicious Client Exclusion Mechanism for Secure Federated Learning-based Road Condition Classification (Liu and Papadimitratos, December 2025)
116 ▪ Adversarial Robustness of Traffic Classification under Resource Constraints: Input Structure Matters (Chehade et al., December 2025)
117 ▪ Bias Injection Attacks on RAG Databases and Sanitization Defenses (Wu and Saxena, December 2025)
118 ▪ Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models (Tice et al., December 2025)
119 ▪ Dataset Poisoning Attacks on Behavioral Cloning Policies (Kalra et al., November 2025)
120 ▪ Synthetic Data: AI's New Weapon Against Android Malware (Nogueira et al., November 2025)
121 ▪ FedPoisonTTP: A Threat Model and Poisoning Attack for Federated Test-Time Personalization (Iftee et al., November 2025)
122 ▪ Decoding Deception: Understanding Automatic Speech Recognition Vulnerabilities in Evasion and Poisoning Attacks (G, Govindarajulu, and Shah, November 2025)
123 ▪ One Pic is All it Takes: Poisoning Visual Document Retrieval Augmented Generation with a Single Image (Shereen et al., November 2025)
124 ▪ Robustness of LLM-enabled vehicle trajectory prediction under data security threats (Wang and Liu, November 2025)
125 ▪ Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System (He et al., November 2025)
126 ▪ CITADEL: A Semi-Supervised Active Learning Framework for Malware Detection Under Continuous Distribution Drift (Haque et al., November 2025)
127 ▪ AMUN: Adversarial Machine UNlearning (Ebrahimpour-Boroojeny, Sundaram, and Chandrasekaran, November 2025)
128 ▪ Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis (Abtahi et al., November 2025)
129 ▪ Taught by the Flawed: How Dataset Insecurity Breeds Vulnerable AI Code (Xia and Alalfi, November 2025)
130 ▪ DP-GENG : Differentially Private Dataset Distillation Guided by DP-Generated Data (Shi et al., November 2025)
131 ▪ Joint-GCG: Unified Gradient-Based Poisoning Attacks on Retrieval-Augmented Generation Systems (Wang et al., November 2025)
132 ▪ PrometheusFree: Concurrent Detection of Laser Fault Injection Attacks in Optical Neural Networks (Nishida et al., November 2025)
133 ▪ Privacy on the Fly: A Predictive Adversarial Transformation Network for Mobile Sensor Data (Song et al., November 2025)
134 ▪ Adversarial Node Placement in Decentralized Federated Learning: Maximum Spanning-Centrality Strategy and Performance Analysis (Piaseczny et al., November 2025)
135 ▪ IndirectAD: Practical Data Poisoning Attacks against Recommender Systems for Item Promotion (Wang et al., November 2025)
136 ▪ Retrieval-Augmented Review Generation for Poisoning Recommender Systems (Yang et al., November 2025)
137 ▪ Adaptive and Robust Data Poisoning Detection and Sanitization in Wearable IoT Systems using Large Language Models (Mithsara et al., November 2025)
138 ▪ On The Dangers of Poisoned LLMs In Security Automation (Karlsen and Eilertsen, November 2025)
139 ▪ Rescuing the Unpoisoned: Efficient Defense against Knowledge Corruption Attacks on RAG Systems (Kim, Lee, and Koo, November 2025)
140 ▪ Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning (Churina, Chebrolu, and Jaidka, November 2025)
141 ▪ VISAT: Benchmarking Adversarial and Distribution Shift Robustness in Traffic Sign Recognition with Visual Attributes (Yu et al., November 2025)
142 ▪ PEEL: A Poisoning-Exposing Encoding Theoretical Framework for Local Differential Privacy (Shuai et al., October 2025)
143 ▪ On the Fragility of Contribution Score Computation in Federated Learning (Pejo et al., October 2025)
144 ▪ RAGRank: Using PageRank to Counter Poisoning in CTI LLM Pipelines (Jia et al., October 2025)
145 ▪ The Black Tuesday Attack: how to crash the stock market with adversarial examples to financial forecasting models (Hofweber et al., October 2025)
146 ▪ Delta-Influence: Unlearning Poisons via Influence Functions (Li et al., October 2025)
147 ▪ Who Taught the Lie? Responsibility Attribution for Poisoned Knowledge in Retrieval-Augmented Generation (Zhang et al., October 2025)
148 ▪ Byzantine Failures Harm the Generalization of Robust Distributed Learning Algorithms More Than Data Poisoning (Boudou et al., October 2025)
149 ▪ ADMIT: Few-shot Knowledge Poisoning Attacks on RAG-based Fact Checking (Wu et al., October 2025)
150 ▪ Cascading Adversarial Bias from Injection to Distillation in Language Models (Chaudhari et al., October 2025)
151 ▪ Tracing Back the Malicious Clients in Poisoning Attacks to Federated Learning (Jia et al., October 2025)
152 ▪ Fairness-Constrained Optimization Attack in Federated Learning (Kasyap et al., October 2025)
153 ▪ GenoArmory: A Unified Evaluation Framework for Adversarial Attacks on Genomic Foundation Models (Luo et al., October 2025)
154 ▪ When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge? (Wang et al., October 2025)
155 ▪ When Vision Fails: Text Attacks Against ViT and OCR (Boucher et al., October 2025)
156 ▪ RAG-Pull: Imperceptible Attacks on RAG Systems for Code Generation (Stambolic, Dhar, and Cavigelli, October 2025)
157 ▪ How Secure is Forgetting? Linking Machine Unlearning to Machine Learning Attacks (P. et al., October 2025)
158 ▪ Provable Watermarking for Data Poisoning Attacks (Zhu, Yu, and Gao, October 2025)
159 ▪ MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks (Ha et al., October 2025)
160 ▪ A Statistical Hypothesis Testing Framework for Data Misappropriation Detection in Large Language Models (Cai, Li, and Zhang, October 2025)
161 ▪ Impact of Dataset Properties on Membership Inference Vulnerability of Deep Transfer Learning (Tobaben et al., October 2025)
162 ▪ From Theory to Practice: Evaluating Data Poisoning Attacks and Defenses in In-Context Learning on Social Media Health Discourse (Jhuma and Faisal, October 2025)
163 ▪ Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training (Yu et al., October 2025)
164 ▪ Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering Attractors (Chen et al., October 2025)
165 ▪ From Mean to Extreme: Formal Differential Privacy Bounds on the Success of Real-World Data Reconstruction Attacks (Riess et al., October 2025)
166 ▪ Fast Exact Unlearning for In-Context Learning Data for LLMs (Muresanu et al., October 2025)
167 ▪ The Impact of Scaling Training Data on Adversarial Robustness (Zimmerli et al., October 2025)
168 ▪ Can In-Context Reinforcement Learning Recover From Reward Poisoning Attacks? (Sasnauskas, Yalın, and Radanović, September 2025)
169 ▪ FuncPoison: Poisoning Function Library to Hijack Multi-agent Autonomous Driving Systems (Long and Li, September 2025)
170 ▪ Virus Infection Attack on LLMs: Your Poisoning Can Spread "VIA" Synthetic Data (Liang et al., September 2025)
171 ▪ AntiFLipper: A Secure and Efficient Defense Against Label-Flipping Attacks in Federated Learning (Rahman et al., September 2025)
172 ▪ Defending Against Beta Poisoning Attacks in Machine Learning Models (Gulciftci and Gursoy, August 2025)
173 ▪ UTrace: Poisoning Forensics for Private Collaborative Learning (Rose et al., August 2025)
174 ▪ Graph Representation-based Model Poisoning on Federated Large Language Models (Cai et al., August 2025)
175 ▪ Scalable contribution bounding to achieve privacy (Cohen-Addad et al., August 2025)
176 ▪ Privacy-Preserving Federated Learning Scheme with Mitigating Model Poisoning Attacks: Vulnerabilities and Countermeasures (Wu et al., July 2025)
177 ▪ Generalizable Targeted Data Poisoning against Varying Physical Objects (Chen et al., July 2025)
178 ▪ CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset Separation (Xu et al., July 2025)
179 ▪ PPFPL: Cross-silo Privacy-preserving Federated Prototype Learning Against Data Poisoning Attacks on Non-IID Data (Zhang et al., July 2025)
180 ▪ OnePath: Efficient and Privacy-Preserving Decision Tree Inference in the Cloud (Yuan et al., July 2025)
181 ▪ Sparsification Under Siege: Defending Against Poisoning Attacks in Communication-Efficient Federated Learning (Jin et al., July 2025)
182 ▪ Rethinking Data Protection in the (Generative) Artificial Intelligence Era (Li et al., July 2025)
183 ▪ A Bayesian Incentive Mechanism for Poison-Resilient Federated Learning (Commey et al., July 2025)
184 ▪ A Privacy-Preserving Framework for Advertising Personalization Incorporating Federated Learning and Differential Privacy (Li, Lin, and Zhang, July 2025)
185 ▪ Effective Fine-Tuning of Vision Transformers with Low-Rank Adaptation for Privacy-Preserving Image Classification (Lin, Imaizumi, and Kiya, July 2025)
186 ▪ Adaptive Federated Learning with Functional Encryption: A Comparison of Classical and Quantum-safe Options (Sorbera et al., July 2025)
187 ▪ When and Where do Data Poisons Attack Textual Inversion? (Styborski et al., July 2025)
188 ▪ Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts (Zarlenga et al., July 2025)
189 ▪ RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation (Zhao et al., July 2025)
190 ▪ Q-Detection: A Quantum-Classical Hybrid Poisoning Attack Detection Method (He et al., July 2025)
191 ▪ Phantom Subgroup Poisoning: Stealth Attacks on Federated Recommender Systems (Yan et al., July 2025)
192 ▪ Privacy-preserving Machine Learning in Internet of Vehicle Applications: Fundamentals, Recent Advances, and Future Direction (Islam and Zulkernine, July 2025)
193 ▪ The Impact of Event Data Partitioning on Privacy-aware Process Discovery (Lim et al., July 2025)
194 ▪ Asynchronous Event Error-Minimizing Noise for Safeguarding Event Dataset (Wang et al., July 2025)
195 ▪ DATABench: Evaluating Dataset Auditing in Deep Learning from an Adversarial Perspective (Shao et al., July 2025)
196 ▪ A Linear Approach to Data Poisoning (Granziol and Flynn, May 2025)
197 ▪ Traceback of Poisoning Attacks to Retrieval-Augmented Generation (Zhang et al, Apr 2025)
198 ▪ Machine Unlearning Fails to Remove Data Poisoning Attacks (Pawelczyk , Apr 2025)
199 ▪ Defending Against Sophisticated Poisoning Attacks with RL-based Aggregation in Federated Learning (Wang et al, Dec 2024)
200 ▪ Adversarial Poisoning Attack on Quantum Machine Learning Models (Kundu and Ghosh, Nov 2024)
201 ▪ Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws (Bowen et al, Oct 2024)
202 ▪ Inverting Gradient Attacks Naturally Makes Data Poisons: An Availability Attack on Neural Networks (Bouaziz, Mhamdi, and Usunier, Oct 2024)
203 ▪ Certified Robustness to Data Poisoning in Gradient-Based Training (Sosnin et al, Oct 2024)
204 ▪ Controlled Generation of Natural Adversarial Documents for Stealthy Retrieval Poisoning (Zhang et al, Oct 2024)
205 ▪ Securing Voice Authentication Applications Against Targeted Data Poisoning (Mohammadi et al, Oct 2024)
206 ▪ Data Poisoning-based Backdoor Attack Framework against Supervised Learning Rules of Spiking Neural Networks (Jin et al, Sep 2024)
207 ▪ Hiding Backdoors within Event Sequence Data via Poisoning Attacks (Ermilova et al, Aug 2024)
208 ▪ ConfusedPilot: Confused Deputy Risks in RAG-based LLMs (RoyChowdhury et al, Aug 2024)
209 ▪ PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models (Zou et al, Aug 2024)
210 ▪ Scaling Laws for Data Poisoning in LLMs (Bowen et al, Aug 2024)
211 ▪ Threats, Attacks, and Defenses in Machine Unlearning: A Survey (Liu et al, Aug 2024)
212 ▪ Debiased Graph Poisoning Attack via Contrastive Surrogate Objective (Yoon et al, Jul 2024)
Data Poisoning and Simulated Publication of Poisoned Public Datasets

>

<

‍

Model (Mis)Interpretability

Added this subsection to cover cybersecurity issues that arise from interpretability issues.

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model (Mis)Interpretability

cybersecurity_tracker - Google Drive

2 ▪ When the Ruler is Broken: Parsing-Induced Suppression in LLM-Based Security Log Evaluation (Garware and Zisad, May 2026)
3 ▪ How Code Representation Shapes False-Positive Dynamics in Cross-Language LLM Vulnerability Detection (Chen et al., May 2026)
4 ▪ Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models (Min, Pham, and Sun, April 2026)
5 ▪ Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways (Lin et al., April 2026)
6 ▪ Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs (Agarwal et al., April 2026)
7 ▪ Attacks Meet Interpretability (AmI) Evaluation and Findings (Ma, Ye, and Mehnaz, April 2026)
8 ▪ Attribution-Driven Explainable Intrusion Detection with Encoder-Based Large Language Models (Biswas et al., April 2026)
9 ▪ Systematic Integration of Digital Twins and Constrained LLMs for Interpretable Cyber-Physical Anomaly Detection (Kampourakis, Gkioulos, and Katsikas, April 2026)
10 ▪ Evaluating the Reliability and Fidelity of Automated Judgment Systems of Large Language Models (Biskupski and Kleber, March 2026)
11 ▪ Manifold of Failure: Behavioral Attraction Basins in Language Models (Munshi et al., February 2026)
12 ▪ Layer-Targeted Multilingual Knowledge Erasure in Large Language Models (Li, Chandrasekaran, and Yu, February 2026)
13 ▪ MARVEL: Multi-Agent RTL Vulnerability Extraction using Large Language Models (Collini et al., February 2026)
14 ▪ Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling (Srisumrith and Sodsee, February 2026)
15 ▪ LLM-FS: Zero-Shot Feature Selection for Effective and Interpretable Malware Detection (Gill, K, and D, February 2026)
16 ▪ Empirical Analysis of Adversarial Robustness and Explainability Drift in Cybersecurity Classifiers (Rajhans and Khawarey, February 2026)
17 ▪ Human-Centered Explainability in AI-Enhanced UI Security Interfaces: Designing Trustworthy Copilots for Cybersecurity Analysts (Rajhans, February 2026)
18 ▪ The Semantic Trap: Do Fine-tuned LLMs Learn Vulnerability Root Cause or Just Functional Pattern? (Huang et al., February 2026)
19 ▪ Unveiling the Black Box: A Multi-Layer Framework for Explaining Reinforcement Learning-Based Cyber Agents (Goel et al., January 2026)
20 ▪ Attesting Model Lineage by Consisted Knowledge Evolution with Fine-Tuning Trajectory (Shang et al., January 2026)
21 ▪ Verbatim Data Transcription Failures in LLM Code Generation: A State-Tracking Stress Test (Haque et al., January 2026)
22 ▪ Explain First, Trust Later: LLM-Augmented Explanations for Graph-Based Crypto Anomaly Detection (Watson, Richards, and Schiff, December 2025)
23 ▪ BEACON: A Unified Behavioral-Tactical Framework for Explainable Cybercrime Analysis with Large Language Models (Sachdeva et al., December 2025)
24 ▪ Can VLMs Detect and Localize Fine-Grained AI-Edited Images? (Sun et al., December 2025)
25 ▪ The Hidden DNA of LLM-Generated JavaScript: Structural Patterns Enable High-Accuracy Authorship Attribution (Tihanyi et al., December 2025)
26 ▪ LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge (Li et al., November 2025)
27 ▪ Interpretable Ransomware Detection Using Hybrid Large Language Models: A Comparative Analysis of BERT, RoBERTa, and DeBERTa Through LIME and SHAP (Ngoie et al., November 2025)
28 ▪ Interpretable LLM Guardrails via Sparse Representation Steering (He et al., November 2025)
29 ▪ RepoMark: A Code Usage Auditing Framework for Code Large Language Models (Qu et al., November 2025)
30 ▪ Model Provenance Testing for Large Language Models (Nikolic, Baluta, and Saxena, October 2025)
31 ▪ Explainable but Vulnerable: Adversarial Attacks on XAI Explanation in Cybersecurity Applications (Mia and Pritom, October 2025)
32 ▪ Reconstructing Trust Embeddings from Siamese Trust Scores: A Direct-Sum Approach with Fixed-Point Semantics (Alpay, Alpay, and Kilictas, August 2025)
33 ▪ Preliminary Investigation into Uncertainty-Aware Attack Stage Classification (Gaudenzi et al., August 2025)
34 ▪ Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems (Liu et al., August 2025)
35 ▪ CodableLLM: Automating Decompiled and Source Code Mapping for LLM Dataset Generation (Manuel and Rad, July 2025)
36 ▪ Breaking Obfuscation: Cluster-Aware Graph with LLM-Aided Recovery for Malicious JavaScript Detection (Liang et al., July 2025)
37 ▪ POLARIS: Explainable Artificial Intelligence for Mitigating Power Side-Channel Leakage (Mahfuz et al., July 2025)
38 ▪ Adversarial attacks and defenses in explainable artificial intelligence: A survey (Baniecki and Biecek, July 2025)
39 ▪ Out of Distribution, Out of Luck: How Well Can LLMs Trained on Vulnerability Datasets Detect Top 25 CWE Weaknesses? (Li et al., July 2025)
40 ▪ Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs (Pan et al., July 2025)
41 ▪ KGV: Integrating Large Language Models with Knowledge Graphs for Cyber Threat Intelligence Credibility Assessment (Wu et al., July 2025)
42 ▪ Scheduzz: Constraint-based Fuzz Driver Generation with Dual Scheduling (Li et al., July 2025)
43 ▪ Lower Bounds for Public-Private Learning under Distribution Shift (Setlur, Thaker, and Ullman, July 2025)
44 ▪ LLMxCPG: Context-Aware Vulnerability Detection Through Code Property Graph-Guided Large Language Models (Lekssays et al., July 2025)
45 ▪ Explainable Vulnerability Detection in C/C++ Using Edge-Aware Graph Attention Networks (Haque et al., July 2025)
46 ▪ Attacking interpretable NLP systems (Abdukhamidov et al., July 2025)
47 ▪ Too Much to Trust? Measuring the Security and Cognitive Impacts of Explainability in AI-Driven SOCs (Rastogi et al., July 2025)
48 ▪ Distributional Unlearning: Forgetting Distributions, Not Just Samples (Allouah, Guerraoui, and Koyejo, July 2025)
49 ▪ Breaking the Illusion of Security via Interpretation: Interpretable Vision Transformer Systems under Attack (Abdukhamidov et al., July 2025)
50 ▪ Multi-Granular Discretization for Interpretable Generalization in Precise Cyberattack Identification (Chung, Huang, and Pai, July 2025)
51 ▪ GPU-Accelerated Interpretable Generalization for Rapid Cyberattack Detection and Forensics (Huang, Chung, and Pai, July 2025)
52 ▪ SoK: Semantic Privacy in Large Language Models (Ma et al., July 2025)
53 ▪ What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift (Chang et al., July 2025)
54 ▪ Interpreting Differential Privacy in Terms of Disclosure Risk (Kazan et al., July 2025)
55 ▪ White-Basilisk: A Hybrid Model for Code Vulnerability Detection (Lamprou et al., July 2025)
56 ▪ Sparse Autoencoder as a Zero-Shot Classifier for Concept Erasing in Text-to-Image Diffusion Models (Tian et al., July 2025)
57 ▪ Protecting Classifiers From Attacks (Gallego et al., July 2025)
58 ▪ Open Problems in Mechanistic Interpretability (Sharkley et al, Jan 2025)
59 ▪ Fooling SHAP with Output Shuffling Attacks (Yuan and Dasgupta, Aug 2024)
60 ▪ Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability (García-Carrasco, Maté, and Trujillo, Jul 2024)
Model (Mis)Interpretability

>

<

‍

Model Collapse

Covers:

  • OWASP LLM 03: Training Data Poisoning
  • OWASP ML 02: Data Poisoning Attack

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Collapse

cybersecurity_tracker - Google Drive

2 ▪ Persona-Model Collapse in Emergent Misalignment (Costa and Vicente, May 2026)
3 ▪ Omission Constraints Decay While Commission Constraints Persist in Long-Context LLM Agents (Gamage, April 2026)
4 ▪ SpectralGuard: Detecting Memory Collapse Attacks in State Space Models (Bonetto, March 2026)
5 ▪ Self-Destructive Language Model (Wang, Zhu, and Wang, March 2026)
6 ▪ Forgetting-MarI: LLM Unlearning via Marginal Information Regularization (Xu et al., November 2025)
7 ▪ LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring (Li, Phuong, and Siegel, August 2025)
8 ▪ Threats, Attacks, and Defenses in Machine Unlearning: A Survey (Liu et al, Feb 2025)
9 ▪ Data Duplication: A Novel Multi-Purpose Attack Paradigm in Machine Unlearning (Ye et al, Jan 2025)
10 ▪ Understanding Implosion in Text-to-Image Generative Models (Ding et al, Sep 2024)
11 ▪ AI models collapse when trained on recursively generated data (Shumailov et al, Jul 2024)
Model Collapse

>

<

‍

Model Denial of Service and Chaff Data Spamming

Covers:

  • OWASP LLM 04: Model Denial of Service
  • MITRE ATLAS Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Denial of Service and Chaff Data Spamming

cybersecurity_tracker - Google Drive

2 ▪ Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Models (Wang et al., May 2026)
3 ▪ AESOP: Adversarial Execution-path Selection to Overload Deep Learning Pipelines (Li et al., May 2026)
4 ▪ Can a Single Message Paralyze the AI Infrastructure? The Rise of AbO-DDoS Attacks through Targeted Mobius Injection (Liang et al., May 2026)
5 ▪ FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework (B\'ejar, Romera-Paredes, and Hern\'andez-Ramos, May 2026)
6 ▪ Position Paper: Denial-of-Service against Multi-Round Transaction Simulation (Tang et al., April 2026)
7 ▪ Semantic Denial of Service in LLM-controlled robots (Steinberg and Gal, April 2026)
8 ▪ Position Paper: Denial-of-Service Against Multi-Round Transaction Simulation (Tang et al., April 2026)
9 ▪ Bit-Flip Vulnerability of Shared KV-Cache Blocks in LLM Serving Systems (Yamamoto and Matsuura, April 2026)
10 ▪ Detecting and Mitigating DDoS Attacks with AI: A Survey (Apostu et al., March 2026)
11 ▪ RECUR: Resource Exhaustion Attack via Recursive-Entropy Guided Counterfactual Utilization and Reflection (Wang et al., February 2026)
12 ▪ Rethinking Latency Denial-of-Service: Attacking the LLM Serving Framework, Not the Model (Wang et al., February 2026)
13 ▪ OverThink: Slowdown Attacks on Reasoning LLMs (Kumar et al., February 2026)
14 ▪ ReasoningBomb: A Stealthy Denial-of-Service Attack by Inducing Pathologically Long Reasoning in Large Reasoning Models (Liu et al., February 2026)
15 ▪ Rethinking On-Device LLM Reasoning: Why Analogical Mapping Outperforms Abstract Thinking for IoT DDoS Detection (Pan et al., January 2026)
16 ▪ RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress (Huang et al., January 2026)
17 ▪ Prompt-Induced Over-Generation as Denial-of-Service: A Black-Box Attack-Side Benchmark (Manu et al., January 2026)
18 ▪ ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking (Li et al., December 2025)
19 ▪ RemedyGS: Defend 3D Gaussian Splatting against Computation Cost Attacks (Li et al., December 2025)
20 ▪ LoopLLM: Transferable Energy-Latency Attacks in LLMs via Repetitive Generation (Li et al., November 2025)
21 ▪ Automated and Explainable Denial of Service Analysis for AI-Driven Intrusion Detection Systems (Yakubu et al., November 2025)
22 ▪ Proactive DDoS Detection and Mitigation in Decentralized Software-Defined Networking via Port-Level Monitoring and Zero-Training Large Language Models (Swileh and Zhang, November 2025)
23 ▪ AdaDoS: Adaptive DoS Attack via Deep Adversarial Reinforcement Learning in SDN (Shao et al., October 2025)
24 ▪ One Token Embedding Is Enough to Deadlock Your Large Reasoning Model (Zhang et al., October 2025)
25 ▪ From Description to Detection: LLM based Extendable O-RAN Compliant Blind DoS Detection in 5G and Beyond (Dayaratne et al., October 2025)
26 ▪ BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models (Yan et al., September 2025)
27 ▪ RECALLED: An Unbounded Resource Consumption Attack on Large Vision-Language Models (Gao et al., July 2025)
28 ▪ Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World (Sadasivan et al., July 2025)
29 ▪ Per-Row Activation Counting on Real Hardware: Demystifying Performance Overheads (Kim et al., July 2025)
30 ▪ Impact of White-Box Adversarial Attacks on Convolutional Neural Networks (Podder and Ghosh, Oct 2024)
31 ▪ DrLLM: Prompt-Enhanced Distributed Denial-of-Service Resistance Method with Large Language Models (Yin, Liu and Xu, Sep 2024)
32 ▪ Bergeron: Combating Adversarial Attacks through a Conscience-Based Alignment Framework (Pisano et al, Aug 2024)
33 ▪ Self-Evaluation as a Defense Against Adversarial Attacks on LLMs (Brown et al, Aug 2024)
34 ▪ Vulnerabilities in AI-generated Image Detection: The Challenge of Adversarial Attacks (Diao et al, Jul 2024)
35 ▪ Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent (Waghela, Sen, and Rakshit, Jul 2024)
Model Denial of Service and Chaff Data Spamming

>

<

‍

Model Modifications

We added this subsection to include security issues that arise from post-hoc model modifications such as fine-tuning, quantization.

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Modifications

cybersecurity_tracker - Google Drive

2 ▪ CachePrune: Teaching LLMs What Not to Follow via KV-Cache Editing (Wang et al., May 2026)
3 ▪ Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs (Cai et al., May 2026)
4 ▪ BitFlipScope: Scalable Fault Localization and Recovery for Bit-Flip Corruptions in LLMs (Karamat, Saif, and Garcia, April 2026)
5 ▪ Fine-Tuning Integrity for Modern Neural Networks: Structured Drift Proofs via Norm, Rank, and Sparsity Certificates (Shang and Chen, April 2026)
6 ▪ Locket: Robust Feature-Locking Technique for Language Models (He, Duddu, and Asokan, March 2026)
7 ▪ CNT: Safety-oriented Function Reuse across LLMs via Cross-Model Neuron Transfer (Zhao et al., March 2026)
8 ▪ CtrlAttack: A Unified Attack on World-Model Control in Diffusion Models (Xu et al., March 2026)
9 ▪ Reverse-Engineering Model Editing on Language Models (Sun et al., February 2026)
10 ▪ Making Models Unmergeable via Scaling-Sensitive Loss Landscape (Jang et al., January 2026)
11 ▪ FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing (Wang et al., October 2025)
12 ▪ Are Robust LLM Fingerprints Adversarially Robust? (Nasery et al., October 2025)
13 ▪ Empirical Evaluation of Concept Drift in ML-Based Android Malware Detection (Sabbah et al., July 2025)
14 ▪ Generating Adversarial Point Clouds Using Diffusion Model (Zhao et al., July 2025)
15 ▪ Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning (Li et al., July 2025)
16 ▪ FORTA: Byzantine-Resilient FL Aggregation via DFT-Guided Krum (Shahul and Harshan, July 2025)
17 ▪ Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities (Che et al., July 2025)
18 ▪ LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents (Yan and Zhang, July 2025)
19 ▪ Score Attack: A Lower Bound Technique for Optimal Differentially Private Learning (Cai, Wang, and Zhang, July 2025)
20 ▪ Differentially Private Federated Low Rank Adaptation Beyond Fixed-Matrix (Wen et al., July 2025)
21 ▪ Quantum Properties Trojans (QuPTs) for Attacking Quantum Neural Networks (Bhowmik, Humble, and Thapliyal, July 2025)
22 ▪ Adaptive Diffusion Denoised Smoothing : Certified Robustness via Randomized Smoothing with Differentially Private Guided Denoising Diffusion (Shpilevskiy et al., July 2025)
23 ▪ Decomposition-Based Optimal Bounds for Privacy Amplification via Shuffling (Su, Cheng, and Wang, July 2025)
24 ▪ Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice (Zhang et al., July 2025)
25 ▪ Harry Potter is Still Here! Probing Knowledge Leakage in Targeted Unlearned Large Language Models via Automated Adversarial Prompting<br><br> (To and Le, May 2025)
26 ▪ Finetuning-Activated Backdoors in LLMs (Gloaguen, May 2025)
27 ▪ The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning (Xiang et al, Apr 2025)
28 ▪ Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey (Huang et al, Oct 2024)
29 ▪ Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks (Poppi et al, Oct 2024)
30 ▪ Unveiling the Vulnerability of Private Fine-Tuning in Split-Based Frameworks for Large Language Models: A Bidirectionally Enhanced Attack (Chen et al, Sep 2024)
31 ▪ The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs (Chen et al, Sep 2024)
32 ▪ BadMerging: Backdoor Attacks Against Model Merging (Zhang et al, Sep 2024)
33 ▪ Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning (Huang et al, Sep 2024)
34 ▪ RLCP: A Reinforcement Learning-based Copyright Protection Method for Text-to-Image Diffusion Model (Shi et al, Sep 2024)
35 ▪ Large Language Models as Carriers of Hidden Messages (Hoscilowicz et al, Aug 2024)
36 ▪ Fight Perturbations with Perturbations: Defending Adversarial Attacks via Neuron Influence (Chen et al, Aug 2024)
37 ▪ Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning (Huang et al, Aug 2024)
38 ▪ Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data (Baumgärtner et al, Aug 2024)
39 ▪ Fine-Tuning, Quantization, and LLMs: Navigating Unintended Outcomes (Kumar et al, Jul 2024)
40 ▪ DeepBaR: Fault Backdoor Attack on Deep Neural Network Layers (Martínez-Mejía et al, Jul 2024)
41 ▪ Resilience and Security of Deep Neural Networks Against Intentional and Unintentional Perturbations: Survey and Research Challenges (Sayyed et al, Jul 2024)
Model Modifications

>

<

‍

Inadequate AI Alignment

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Inadequate AI Alignment

cybersecurity_tracker - Google Drive

2 ▪ Ghost in the Context: Measuring Policy-Carriage Failures in Decision-Time Assembly (Santos-Grueiro, May 2026)
3 ▪ No More, No Less: Task Alignment in Terminal Agents (Mavali et al., May 2026)
4 ▪ Leveraging RAG for Training-Free Alignment of LLMs (Halloran, May 2026)
5 ▪ You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation (Arazzi et al., May 2026)
6 ▪ When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models (Hossain et al., May 2026)
7 ▪ Dependency-Aware Privacy for Multi-turn Agents (Anshumaan et al., May 2026)
8 ▪ Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness (Li and Liu, May 2026)
9 ▪ Tatemae: Detecting Alignment Faking via Tool Selection in LLMs (Leonesi et al., April 2026)
10 ▪ Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers (Dubi\'nski et al., April 2026)
11 ▪ Beyond Context: Large Language Models' Failure to Grasp Users' Intent (Hussain and Salahuddin, April 2026)
12 ▪ Beyond Single-Agent Alignment: Preventing Context-Fragmented Violations in Multi-Agent Systems (Wu and Gong, April 2026)
13 ▪ SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs (Pan, Wu, and Yao, April 2026)
14 ▪ Sensitivity Uncertainty Alignment in Large Language Models (Hiremath and Hiremath, April 2026)
15 ▪ Reasoning Hijacking: The Fragility of Reasoning Alignment in Large Language Models (Liu, Tang, and Tun, April 2026)
16 ▪ Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning (Hui et al., April 2026)
17 ▪ Uncovering Logit Suppression Vulnerabilities in LLM Safety Alignment (Li et al., April 2026)
18 ▪ Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs (Roh and Houmansadr, April 2026)
19 ▪ Into the Gray Zone: Domain Contexts Can Blur LLM Safety Boundaries (Hung et al., April 2026)
20 ▪ Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design (Eshuijs, Wang, and Fokkens, April 2026)
21 ▪ The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training (Zhang et al., April 2026)
22 ▪ FreakOut-LLM: The Effect of Emotional Stimuli on Safety Alignment (Kuznetsov et al., April 2026)
23 ▪ Understanding the Effects of Safety Unalignment on Large Language Models (Halloran, April 2026)
24 ▪ PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality (Li et al., April 2026)
25 ▪ Safety, Security, and Cognitive Risks in World Models (Parmar, April 2026)
26 ▪ UK AISI Alignment Evaluation Case-Study (Souly et al., April 2026)
27 ▪ Security in LLM-as-a-Judge: A Comprehensive SoK (Almasoud et al., April 2026)
28 ▪ Colluding LoRA: A Compositional Vulnerability in LLM Safety Alignment (Ding, March 2026)
29 ▪ SafetyDrift: Predicting When AI Agents Cross the Line Before They Actually Do (Dhodapkar and Pishori, March 2026)
30 ▪ Internal Safety Collapse in Frontier Large Language Models (Wu et al., March 2026)
31 ▪ Silent Commitment Failure in Instruction-Tuned Language Models: Evidence of Governability Divergence Across Architectures (Ruddell, March 2026)
32 ▪ Structured Visual Narratives Undermine Safety Alignment in Multimodal Large Language Models (Tan, Hu, and Lee, March 2026)
33 ▪ MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment (Dou and Yang, March 2026)
34 ▪ Beyond Static Pattern Matching? Rethinking Automatic Cryptographic API Misuse Detection in the Era of LLMs (Xia et al., March 2026)
35 ▪ State-Dependent Safety Failures in Multi-Turn Language Model Interaction (Li et al., March 2026)
36 ▪ Understanding LLM Behavior When Encountering User-Supplied Harmful Content in Harmless Tasks (Chu et al., March 2026)
37 ▪ Proof-of-Guardrail in AI Agents and What (Not) to Trust from It (Jin et al., March 2026)
38 ▪ When Safety Becomes a Vulnerability: Exploiting LLM Alignment Homogeneity for Transferable Blocking in RAG (Li et al., March 2026)
39 ▪ Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders (Campbell et al., March 2026)
40 ▪ Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems (Gao et al., February 2026)
41 ▪ Fail-Closed Alignment for Large Language Models (Coalson et al., February 2026)
42 ▪ A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Models (Segal et al., February 2026)
43 ▪ Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning (Wang et al., February 2026)
44 ▪ CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis (Eswaran et al., February 2026)
45 ▪ Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment (Wang et al., February 2026)
46 ▪ Is Reasoning Capability Enough for Safety in Long-Context Language Models? (Fu et al., February 2026)
47 ▪ When Evaluation Becomes a Side Channel: Regime Leakage and Structural Mitigations for Alignment Assessment (Santos-Grueiro, February 2026)
48 ▪ When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents (Jones et al., February 2026)
49 ▪ RASA: Routing-Aware Safety Alignment for Mixture-of-Experts Models (Liang et al., February 2026)
50 ▪ Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs (Chen et al., February 2026)
51 ▪ Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures (Su et al., February 2026)
52 ▪ SafeThinker: Reasoning about Risk to Deepen Safety Beyond Shallow Alignment (Fang et al., January 2026)
53 ▪ LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions (Hu et al., January 2026)
54 ▪ Between a Rock and a Hard Place: The Tension Between Ethical Reasoning and Safety Alignment in LLMs (Chua et al., January 2026)
55 ▪ Are LLMs Vulnerable to Preference-Undermining Attacks (PUA)? A Factorial Analysis Methodology for Diagnosing the Trade-off between Preference Alignment and Real-World Validity (An et al., January 2026)
56 ▪ Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs (Zhang et al., January 2026)
57 ▪ What Matters For Safety Alignment? (Li et al., January 2026)
58 ▪ Beyond Context: Large Language Models Failure to Grasp Users Intent (Hussain, Salahuddin, and Papadimitratos, December 2025)
59 ▪ Large Language Models as a (Bad) Security Norm in the Context of Regulation and Compliance (Ludvigsen, December 2025)
60 ▪ PROPS: Progressively Private Self-alignment of Large Language Models (Teku et al., December 2025)
61 ▪ Cognitive Control Architecture (CCA): A Lifecycle Supervision Framework for Robustly Aligned AI Agents (Liang et al., December 2025)
62 ▪ Matching Ranks Over Probability Yields Truly Deep Safety Alignment (Vega and Singh, December 2025)
63 ▪ Where to Start Alignment? Diffusion Large Language Model May Demand a Distinct Position (Xie, Song, and Luo, November 2025)
64 ▪ Can LLMs Make (Personalized) Access Control Decisions? (Groschupp et al., November 2025)
65 ▪ Can LLMs Threaten Human Survival? Benchmarking Potential Existential Threats from LLMs via Prefix Completion (Cui et al., November 2025)
66 ▪ Understanding and Mitigating Over-refusal for Large Language Models via Safety Representation (Zhang et al., November 2025)
67 ▪ EnchTable: Unified Safety Alignment Transfer in Fine-tuned Large Language Models (Wu et al., November 2025)
68 ▪ A Self-Improving Architecture for Dynamic Safety in Large Language Models (Slater, November 2025)
69 ▪ HLPD: Aligning LLMs to Human Language Preference for Machine-Revised Text Detection (Dai, Jiang, and Deng, November 2025)
70 ▪ EASE: Practical and Efficient Safety Alignment for Small Language Models (Shi et al., November 2025)
71 ▪ XBreaking: Understanding how LLMs security alignment can be broken (Arazzi et al., November 2025)
72 ▪ Reimagining Safety Alignment with An Image (Xia et al., November 2025)
73 ▪ Unvalidated Trust: Cross-Stage Vulnerabilities in Large Language Model Architectures (Schwarz, November 2025)
74 ▪ When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents? (Peng et al., October 2025)
75 ▪ HarmRLVR: Weaponizing Verifiable Rewards for Harmful LLM Alignment (Liu et al., October 2025)
76 ▪ Cross-Modal Safety Alignment: Is textual unlearning all you need? (Chakraborty et al., October 2025)
77 ▪ A Biosecurity Agent for Lifecycle LLM Biosecurity Alignment (Meng and Zhang, October 2025)
78 ▪ LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions (Hu et al., October 2025)
79 ▪ Refusal Falls off a Cliff: How Safety Alignment Fails in Reasoning? (Yin et al., October 2025)
80 ▪ Superficial Safety Alignment Hypothesis (Li and Kim, October 2025)
81 ▪ AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning (Zhang et al., July 2025)
82 ▪ Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework (Dassanayake et al., July 2025)
83 ▪ ARMOR: Aligning Secure and Safe Large Language Models via Meticulous Reasoning (Zhao et al., July 2025)
84 ▪ Agent Safety Alignment via Reinforcement Learning (Sha et al., July 2025)
85 ▪ On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment (Ball et al., July 2025)
86 ▪ Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework (Krishna et al., July 2025)
87 ▪ Emergent misalignment as prompt sensitivity: A research note (Wyse et al., July 2025)
Inadequate AI Alignment

>

<

‍

‍

Discover ML Model Family and Ontology/Model Extraction

Added model extraction to this as it did not have its own category.

Covers:

  • MITRE ATLAS Discovery

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Discover ML Model Family and Ontology/Model Extraction

cybersecurity_tracker - Google Drive

2 ▪ Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies (Cotti et al., April 2026)
3 ▪ Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs (Liu, Evans, and Xiong, April 2026)
4 ▪ CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models (Lee, Koo, and Kwak, April 2026)
5 ▪ AttnDiff: Attention-based Differential Fingerprinting for Large Language Models (Zhang et al., April 2026)
6 ▪ Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test (Zhu et al., March 2026)
7 ▪ Navigating the Deep: End-to-End Extraction on Deep Neural Networks (Liu et al., February 2026)
8 ▪ A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors (Xu and Sheng, February 2026)
9 ▪ FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting (Fu et al., January 2026)
10 ▪ Identifying Models Behind Text-to-Image Leaderboards (Naseh et al., January 2026)
11 ▪ Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures (Wang et al., December 2025)
12 ▪ A Fingerprint for Large Language Models (Yang and Wu, December 2025)
13 ▪ SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting (Zhang and Zheng, December 2025)
14 ▪ A Systematic Study of Model Extraction Attacks on Graph Foundation Models (Xu et al., November 2025)
15 ▪ Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution Approach (Li et al., November 2025)
16 ▪ Ghost in the Transformer: Tracing LLM Lineage with SVD-Fingerprint (Wang et al., November 2025)
17 ▪ Structuring Security: A Survey of Cybersecurity Ontologies, Semantic Log Processing, and LLMs Application (Louren\c{c}o et al., October 2025)
18 ▪ MalCVE: Malware Detection and CVE Association Using Large Language Models (Cristea, Molnes, and Li, October 2025)
19 ▪ Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation (Shao et al., October 2025)
20 ▪ SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From (Tong et al., October 2025)
21 ▪ LLM-Assisted Model-Based Fuzzing of Protocol Implementations (Huang, Wang, and Zhou, August 2025)
22 ▪ PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models (Gasmi et al., July 2025)
23 ▪ Evaluating Ensemble and Deep Learning Models for Static Malware Detection with Dimensionality Reduction Using the EMBER Dataset (Abedin and Mehrub, July 2025)
24 ▪ Revisiting Pre-trained Language Models for Vulnerability Detection (Li et al., July 2025)
25 ▪ Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications (H\"arer, July 2025)
26 ▪ TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph (Wang et al., July 2025)
27 ▪ Toward an Intent-Based and Ontology-Driven Autonomic Security Response in Security Orchestration Automation and Response (Huang et al., July 2025)
28 ▪ SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents (Masoor, July 2025)
29 ▪ UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification (Cai et al., July 2025)
30 ▪ AICrypto: A Comprehensive Benchmark For Evaluating Cryptography Capabilities of Large Language Models (Wang et al., July 2025)
31 ▪ BISON: Blind Identification with Stateless scOped pseudoNyms (Heher, More, and Heimberger, July 2025)
32 ▪ Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms (Gasmi et al., July 2025)
33 ▪ From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks (Khouna, Ferry, and Vidal, July 2025)
34 ▪ One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs<br><br> (Li et al., May 2025)
35 ▪ Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries (Wang et al, May 2025)
36 ▪ How to Backdoor the Knowledge Distillation (Wu et al, Apr 2025)
37 ▪ Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations (Ezzeddine, Ayoub, and Giordano, Oct 2024)
38 ▪ Efficient and Effective Model Extraction (Zhu et al, Sep 2024)
39 ▪ CaBaGe: Data-Free Model Extraction using ClAss BAlanced Generator Ensemble (Rosenthal et al, Sep 2024)
40 ▪ Alignment-Aware Model Extraction Attacks on Large Language Models (Liang et al, Sep 2024)
Discover ML Model Family and Ontology/Model Extraction

>

<

‍

Improper Error Handling

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Improper Error Handling

cybersecurity_tracker - Google Drive

Improper Error Handling

>

<

‍

Robust Multi-Prompt and Multi-Model Attacks

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Robust Multi-Prompt and Multi-Model Attacks

cybersecurity_tracker - Google Drive

2 ▪ Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting (Chin et al., May 2026)
3 ▪ When LLMs Team Up: A Coordinated Attack Framework for Automated Cyber Intrusions (Qi et al., May 2026)
4 ▪ Don't Trust Your Upstream: Exploiting LLM Multi-Agent System via Topology-Guided Adversarial Propagation (Liang et al., May 2026)
5 ▪ Semantic Intent Fragmentation: A Single-Shot Compositional Attack on Multi-Agent AI Pipelines (Ahad et al., April 2026)
6 ▪ TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories (Chen et al., April 2026)
7 ▪ When Safe Models Merge into Danger: Exploiting Latent Vulnerabilities in LLM Fusion (Li et al., April 2026)
8 ▪ ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents (Wu et al., March 2026)
9 ▪ Multi-Stream Perturbation Attack: Breaking Safety Alignment of Thinking LLMs Through Concurrent Task Interference (Yang, March 2026)
10 ▪ OpenRT: An Open-Source Red Teaming Framework for Multimodal LLMs (Wang et al., January 2026)
11 ▪ Tipping the Dominos: Topology-Aware Multi-Hop Attacks on LLM-Based Multi-Agent Systems (Liang et al., December 2025)
12 ▪ Multi-Faceted Attack: Exposing Cross-Model Vulnerabilities in Defense-Equipped Vision-Language Models (Yang et al., November 2025)
13 ▪ RAG-targeted Adversarial Attack on LLM-based Threat Detection and Mitigation Framework (Ikbarieh, Aryal, and Gupta, November 2025)
14 ▪ Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks (Chang et al., November 2025)
15 ▪ MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agents (Zhang et al., October 2025)
16 ▪ Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment (Shahid et al., October 2025)
Robust Multi-Prompt and Multi-Model Attacks

>

<

‍

Multi-Modal Attacks

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Multi-Modal Attacks

cybersecurity_tracker - Google Drive

2 ▪ Adversarial Hubness in Multi-Modal Retrieval (Zhang et al., May 2026)
3 ▪ Still Camouflage, Moving Illusion: View-Induced Trajectory Manipulation in Autonomous Driving (Ju et al., May 2026)
4 ▪ FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence (Yan et al., May 2026)
5 ▪ STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack (Mao et al., May 2026)
6 ▪ Cross-Modal Phantom: Coordinated Camera-LiDAR Spoofing Against Multi-Sensor Fusion in Autonomous Vehicles (Khan and Hasan, April 2026)
7 ▪ SoK: The Next Frontier in AV Security: Systematizing Perception Attacks and the Emerging Threat of Multi-Sensor Fusion (Khan, Islam, and Hasan, April 2026)
8 ▪ Text Steganography with Dynamic Codebook and Multimodal Large Language Model (Gao, Lei, and Peng, April 2026)
9 ▪ ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety (Wang et al., April 2026)
10 ▪ Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmark (Zhang et al., April 2026)
11 ▪ From Incomplete Architecture to Quantified Risk: Multimodal LLM-Driven Security Assessment for Cyber-Physical Systems (Huang, Poskitt, and Shar, April 2026)
12 ▪ See No Evil: Adversarial Attacks Against Linguistic-Visual Association in Referring Multi-Object Tracking Systems (Bouzidi, Liu, and Faruque, March 2026)
13 ▪ Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey (Jain, Ar{\i}k, and Thakur, March 2026)
14 ▪ Improving Generalization on Cybersecurity Tasks with Multi-Modal Contrastive Learning (Huang et al., March 2026)
15 ▪ REFORGE: Multi-modal Attacks Reveal Vulnerable Concept Unlearning in Image Generation Models (Zou et al., March 2026)
16 ▪ Co-Evolutionary Multi-Modal Alignment via Structured Adversarial Evolution (Shi et al., March 2026)
17 ▪ MemoPhishAgent: Memory-Augmented Multi-Modal LLM Agent for Phishing URL Detection (Chen et al., February 2026)
18 ▪ Dynamic Deception: When Pedestrians Team Up to Fool Autonomous Cars (Tehrani et al., February 2026)
19 ▪ SecureScan: An AI-Driven Multi-Layer Framework for Malware and Phishing Detection Using Logistic Regression and Threat Intelligence Integration (Firdos and Dangi, February 2026)
20 ▪ Universal Anti-forensics Attack against Image Forgery Detection via Multi-modal Guidance (Li et al., February 2026)
21 ▪ Now You Hear Me: Audio Narrative Attacks Against Large Audio-Language Models (Yu et al., February 2026)
22 ▪ FOCA: Multimodal Malware Classification via Hyperbolic Cross-Attention (Choudhury et al., January 2026)
23 ▪ Integrating APK Image and Text Data for Enhanced Threat Detection: A Multimodal Deep Learning Approach to Android Malware (Arifin, Rahman, and Eisty, January 2026)
24 ▪ Failure Analysis of Safety Controllers in Autonomous Vehicles Under Object-Based LiDAR Attacks (Ganiuly, Bolatbek, and Smaiyl, December 2025)
25 ▪ FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning (Khalil and Hoque, December 2025)
26 ▪ Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities (Xiong et al., December 2025)
27 ▪ Vulnerability-Aware Robust Multimodal Adversarial Training (Zhang et al., November 2025)
28 ▪ Medusa: Cross-Modal Transferable Adversarial Attacks on Multimodal Medical Retrieval-Augmented Generation (Shang et al., November 2025)
29 ▪ Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security (Zhao et al., November 2025)
30 ▪ Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video Approach (Huang et al., November 2025)
31 ▪ MIP against Agent: Malicious Image Patches Hijacking Multimodal OS Agents (Aichberger et al., November 2025)
32 ▪ Security Risk of Misalignment between Text and Image in Multi-modal Model (Wang, Ge, and Wang, October 2025)
33 ▪ DeepTx: Real-Time Transaction Risk Analysis via Multi-Modal Features and LLM Reasoning (Liu, Li, and Li, October 2025)
34 ▪ CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks (Zhang, Li, and Lu, October 2025)
35 ▪ Hierarchical Multi-Modal Threat Intelligence Fusion Without Aligned Data: A Practical Framework for Real-World Security Operations (Doppalapudi, October 2025)
36 ▪ IP-Augmented Multi-Modal Malicious URL Detection Via Token-Contrastive Representation Enhancement and Multi-Granularity Fusion (Tian et al., October 2025)
37 ▪ Cross-Modal Content Optimization for Steering Web Agent Preferences (Jiang et al., October 2025)
38 ▪ Evil Vizier: Vulnerabilities of LLM-Integrated XR Systems (Zhang et al., September 2025)
39 ▪ Multi-Stage Prompt Inference Attacks on Enterprise LLM Systems (Balashov, Ponomarova, and Zhai, July 2025)
40 ▪ Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems (Ko et al., July 2025)
41 ▪ Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities (Qu, Backes, and Zhang, July 2025)
42 ▪ LLM-Stackelberg Games: Conjectural Reasoning Equilibria and Their Applications to Spearphishing (Zhu, July 2025)
43 ▪ The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents (Wang et al., July 2025)
44 ▪ CLIProv: A Contrastive Log-to-Intelligence Multimodal Approach for Threat Detection and Provenance Analysis (Li et al., July 2025)
45 ▪ FreqCross: A Multi-Modal Frequency-Spatial Fusion Network for Robust Detection of Stable Diffusion 3.5 Generated Images (Yang, July 2025)
46 ▪ JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models<br><br> (Peng et al., May 2025)
47 ▪ Vision-LLMs Can Fool Themselves with Self-Generated Typographic Attacks (Qraitem et al, Feb 2025)
48 ▪ Detecting Malicious Concepts Without Image Generation in AIGC (Xu et al, Feb 2025)
49 ▪ Typographic Attacks in a Multi-Image Setting (Wang, Zhao, and Larson, Feb 2025)
50 ▪ T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models (Miao et al, Sep 2024)
Multi-Modal Attacks

>

<

‍

LLM Data Leakage and ML Artifact Collection

  • MITRE ATLAS Exfiltration & Collection

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : LLM Data Leakage and ML Artifact Collection

cybersecurity_tracker - Google Drive

2 ▪ Identifying AI Web Scrapers Using Canary Tokens (Seiden et al., May 2026)
3 ▪ LeakDojo: Decoding the Leakage Threats of RAG Systems (Zhang et al., May 2026)
4 ▪ SecureMCP: A Policy-Enforced LLM Data Access Framework for AIoT Systems via Model Context Protocol (Kim and Yoo, May 2026)
5 ▪ Do Multimodal RAG Systems Leak Data? A Comprehensive Evaluation of Membership Inference and Image Caption Retrieval Attacks (Al-Lawati and Wang, May 2026)
6 ▪ On the Privacy of LLMs: An Ablation Study (Makhlouf et al., May 2026)
7 ▪ Quantamination: Dynamic Quantization Leaks Your Data Across the Batch (Foerster et al., April 2026)
8 ▪ OpenSOC-AI: Democratizing Security Operations with Parameter Efficient LLM Log Analysis (Garware and Zisad, April 2026)
9 ▪ LLM-CEG: Extending the Classification Error Gauge Framework for Privacy Auditing of Large Language Models (Mivule, April 2026)
10 ▪ Understanding Secret Leakage Risks in Code LLMs: A Tokenization Perspective (Chen et al., April 2026)
11 ▪ Data Leakage in Automotive Perception: Practitioners' Insights (Babu et al., April 2026)
12 ▪ LLM-Enabled Open-Source Systems in the Wild: An Empirical Study of Vulnerabilities in GitHub Security Advisories (Shifat et al., April 2026)
13 ▪ Expert Selections In MoE Models Reveal (Almost) As Much As Text (Nuriyev and Kulp, March 2026)
14 ▪ You Told Me to Do It: Measuring Instructional Text-induced Private Data Leakage in LLM Agents (Kao et al., March 2026)
15 ▪ Detecting Cryptographically Relevant Software Packages with Collaborative LLMs (Hirsch et al., March 2026)
16 ▪ The Silent Spill: Measuring Sensitive Data Leaks Across Public URL Repositories (Ramadan et al., February 2026)
17 ▪ From Data Leak to Secret Misses: The Impact of Data Leakage on Secret Detection Models (Soltaniani and Ghafari, February 2026)
18 ▪ Can Large Language Models Really Recognize Your Name? (Pham et al., January 2026)
19 ▪ A Systemic Evaluation of Multimodal RAG Privacy (Al-Lawati and Wang, January 2026)
20 ▪ Network-Level Prompt and Trait Leakage in Local Research Agents (Jeong et al., January 2026)
21 ▪ Burn-After-Use for Preventing Data Leakage through a Secure Multi-Tenant Architecture in Enterprise LLM (Zhang et al., January 2026)
22 ▪ SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use (Desai et al., January 2026)
23 ▪ Beyond BeautifulSoup: Benchmarking LLM-Powered Web Scraping for Everyday Users (Bhardwaj, Diwan, and Wang, January 2026)
24 ▪ Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation (Qiao et al., December 2025)
25 ▪ ContextLeak: Auditing Leakage in Private In-Context Learning Methods (Choi et al., December 2025)
26 ▪ ExpShield: Safeguarding Web Text from Unauthorized Crawling and LLM Exploitation (Liu et al., December 2025)
27 ▪ Topology Matters: Measuring Memory Leakage in Multi-Agent LLMs (Liu et al., December 2025)
28 ▪ Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models (Dang and Mohaisen, November 2025)
29 ▪ Structured Extraction of Vulnerabilities in OpenVAS and Tenable WAS Reports Using LLMs (Machado et al., November 2025)
30 ▪ MCP-RiskCue: Can LLM Infer Risk Information From MCP Server System Logs? (Fu and Sun, November 2025)
31 ▪ MCP-RiskCue: Can LLM infer risk information from MCP server System Logs? (Fu and Sun, November 2025)
32 ▪ Security Logs to ATT&CK Insights: Leveraging LLMs for High-Level Threat Understanding and Cognitive Trait Inference (Hans et al., October 2025)
33 ▪ Unlearned but Not Forgotten: Data Extraction after Exact Unlearning in LLM (Wu et al., October 2025)
34 ▪ Bits Leaked per Query: Information-Theoretic Bounds on Adversarial Attacks against LLMs (Kaneko and Baldwin, October 2025)
35 ▪ Exposing LLM User Privacy via Traffic Fingerprint Analysis: A Study of Privacy Risks in LLM Agent Interactions (Zhang et al., October 2025)
36 ▪ AttackSeqBench: Benchmarking Large Language Models in Analyzing Attack Sequences within Cyber Threat Intelligence (Ma et al., October 2025)
37 ▪ You Have Been LaTeXpOsEd: A Systematic Analysis of Information Leakage in Preprint Archives Using Large Language Models (Dubniczky, Borsos, and Norbert, October 2025)
38 ▪ External Data Extraction Attacks against Retrieval-Augmented Large Language Models (He et al., October 2025)
39 ▪ Sentry: Authenticating Machine Learning Artifacts on the Fly (Gan and Ghodsi, October 2025)
40 ▪ Uncovering Vulnerabilities of LLM-Assisted Cyber Threat Intelligence (Meng et al., September 2025)
41 ▪ LeakyCLIP: Extracting Training Data from CLIP (Chen et al., August 2025)
42 ▪ Trivial Trojans: How Minimal MCP Servers Enable Cross-Tool Exfiltration of Sensitive Data (Croce and South, July 2025)
43 ▪ LoRA-Leak: Membership Inference Attacks Against LoRA Fine-tuned Language Models (Ran et al., July 2025)
44 ▪ Optimizing Canaries for Privacy Auditing with Metagradient Descent (Boglioni et al., July 2025)
45 ▪ Sharing is CAIRing: Characterizing Principles and Assessing Properties of Universal Privacy Evaluation for Synthetic Tabular Data (Hyrup et al., July 2025)
46 ▪ Training Set Reconstruction from Differentially Private Forests: How Effective is DP? (Gorg\'e et al., July 2025)
47 ▪ Targeting Alignment: Extracting Safety Classifiers of Aligned LLMs (Ferrand et al, Jan 2025)
48 ▪ From Prompt Injections to SQL Injection Attacks: How Protected is Your LLM-Integrated Web Application? (Pedro et al, Jan 2025)
49 ▪ RAG-Thief: Scalable Extraction of Private Data from Retrieval-Augmented Generation Applications with Agent-based Attacks (Jiang et al, Nov 2024)
50 ▪ Towards More Realistic Extraction Attacks: An Adversarial Perspective (More, Ganesh, Farnadi, Nov 2024)
51 ▪ Stealing User Prompts from Mixture of Experts (Yona et al, Oct 2024)
52 ▪ Breach By A Thousand Leaks: Unsafe Information Leakage in `Safe' AI Responses (Glukhov et al, Oct 2024)
53 ▪ CodeCloak: A Method for Evaluating and Mitigating Code Leakage by LLM Code Assistants (Finkman Noah et al, Oct 2024)
54 ▪ Towards a Theoretical Understanding of Memorization in Diffusion Models (Chen et al, Oct 2024)
55 ▪ Extracting Memorized Training Data via Decomposition (Su et al, Sep 2024)
56 ▪ Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships (Gyo Oh et al, Sep 2024)
57 ▪ Adaptive Pre-training Data Detection for Large Language Models via Surprising Tokens (Zhang and Wu, Jul 2024)
LLM Data Leakage and ML Artifact Collection

>

<

‍

Evade ML Model

Covers:

  • MITRE ATLAS Defense Evasion & Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Evade ML Model

cybersecurity_tracker - Google Drive

2 ▪ The Role of Learning in Attacking ML-based Network Intrusion Detection (Domico, Ferrand, and McDaniel, May 2026)
3 ▪ WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation (Liu et al., May 2026)
4 ▪ Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption (Morales, Pastrana, and Tapiador, May 2026)
5 ▪ Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, May 2026)
6 ▪ MalPurifier: Enhancing Android Malware Detection with Adversarial Purification against Evasion Attacks (Zhou et al., May 2026)
7 ▪ Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting (Ediga and Chattopadhyay, May 2026)
8 ▪ Trident: Improving Malware Detection with LLMs and Behavioral Features (Saul et al., May 2026)
9 ▪ AsmRAG: LLM-Driven Malware Detection by Retrieving Functionally Similar Assembly Code (Karbab, April 2026)
10 ▪ Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations (Acharya and Zhang, April 2026)
11 ▪ Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing (Peng et al., April 2026)
12 ▪ Evasion Adversarial Attacks Remain Impractical Against ML-based Network Intrusion Detection Systems, Especially Dynamic Ones (elShehaby and Matrawy, March 2026)
13 ▪ ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks (Elnawawy et al., March 2026)
14 ▪ Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation (Chen et al., March 2026)
15 ▪ Targeted Adversarial Traffic Generation : Black-box Approach to Evade Intrusion Detection Systems in IoT Networks (Debicha et al., March 2026)
16 ▪ Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents (Aonzo et al., March 2026)
17 ▪ FuzzySQL: Uncovering Hidden Vulnerabilities in DBMS Special Features with LLM-Driven Fuzzing (Chen et al., February 2026)
18 ▪ Anticipating Adversary Behavior in DevSecOps Scenarios through Large Language Models (Caballero et al., February 2026)
19 ▪ SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization (Zhang et al., February 2026)
20 ▪ Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach (Ali et al., February 2026)
21 ▪ StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors (Ranganath and Ramesh, February 2026)
22 ▪ DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems (Ou et al., February 2026)
23 ▪ Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks (Zhang and Zhang, February 2026)
24 ▪ "Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval (Li et al., February 2026)
25 ▪ Semantics-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, February 2026)
26 ▪ In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement (Shetty, Joshi, and Kanhere, February 2026)
27 ▪ AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises (Tung et al., February 2026)
28 ▪ ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses (He et al., January 2026)
29 ▪ CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation (Zhang et al., January 2026)
30 ▪ A Decompilation-Driven Framework for Malware Detection with Large Language Models (Chawla and Prasad, January 2026)
31 ▪ MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization (Gu, Li, and Hu, January 2026)
32 ▪ VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference (Qi et al., January 2026)
33 ▪ Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools? (Quantrill et al., January 2026)
34 ▪ Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks (Xu et al., December 2025)
35 ▪ Automated Penetration Testing with LLM Agents and Classical Planning (Wang et al., December 2025)
36 ▪ NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration (Wang et al., December 2025)
37 ▪ ExtendAttack: Attacking Servers of LRMs via Extending Reasoning (Zhu et al., November 2025)
38 ▪ Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning (Domico et al., November 2025)
39 ▪ MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification (Luo et al., November 2025)
40 ▪ Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment (Zhang and sun, November 2025)
41 ▪ GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors (Meng et al., October 2025)
42 ▪ The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems (Choi et al., October 2025)
43 ▪ Adversarial Pre-Padding: Generating Evasive Network Traffic Against Transformer-Based Classifiers (Jing et al., October 2025)
44 ▪ Detecting Various DeFi Price Manipulations with LLM Reasoning (Zhong et al., October 2025)
45 ▪ Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks (Feng and Sakurai, October 2025)
46 ▪ Black-Box Evasion Attacks on Data-Driven Open RAN Apps: Tailored Design and Experimental Evaluation (Gajjar et al., October 2025)
47 ▪ From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15 (Rehman et al., October 2025)
48 ▪ SoK: Adversarial Evasion Attacks Practicality in NIDS Domain and the Impact of Dynamic Learning (elShehaby and Matrawy, October 2025)
49 ▪ A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems (Liu et al., October 2025)
50 ▪ VulSolver: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., October 2025)
51 ▪ Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework (Han et al., October 2025)
52 ▪ Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks (Engineering et al., October 2025)
53 ▪ Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection (Creo, August 2025)
54 ▪ AdVAR-DNN: Adversarial Misclassification Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, August 2025)
55 ▪ ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models (Yook et al., July 2025)
56 ▪ Hierarchical Graph Neural Network for Compressed Speech Steganalysis (Hemis et al., July 2025)
57 ▪ Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms (Sato, Ohki, and Nishigaki, July 2025)
58 ▪ GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI (Kalyanapu et al., July 2025)
59 ▪ BandFuzz: An ML-powered Collaborative Fuzzing Framework (Shi et al., July 2025)
60 ▪ PotentRegion4MalDetect: Advanced Features from Potential Malicious Regions for Malware Detection (Koppanati, Santra, and Peddoju, July 2025)
61 ▪ Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors (Wang et al, Oct 2024)
Evade ML Model

>

<

‍

Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access

Covers:

  • OWASP LLM 10: Model Theft
  • OWASP ML 05: Model Theft
  • MITRE ATLAS Exfiltration and ML Model Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access

cybersecurity_tracker - Google Drive

2 ▪ GraphIP-Bench: How Hard Is It to Steal a Graph Neural Network, and Can We Stop It? (Zhao et al., May 2026)
3 ▪ Protecting the Trace: A Principled Black-Box Approach Against Distillation Attacks (Hartman et al., April 2026)
4 ▪ Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model! (Yoon et al., April 2026)
5 ▪ Black-Box Skill Stealing Attack from Proprietary LLM Agents: An Empirical Study (Wang et al., April 2026)
6 ▪ TrEEStealer: Stealing Decision Trees via Enclave Side Channels (Sander et al., April 2026)
7 ▪ AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection (Zhao et al., April 2026)
8 ▪ Automated Malware Family Classification using Weighted Hierarchical Ensembles of Large Language Models (Bai et al., April 2026)
9 ▪ LiteGuard: Efficient Task-Agnostic Model Fingerprinting with Enhanced Generalization (Yang et al., March 2026)
10 ▪ Fingerprinting Deep Neural Networks for Ownership Protection: An Analytical Approach (Yang et al., March 2026)
11 ▪ Towards Model Extraction Attacks in GAN-Based Image Translation via Domain Shift Mitigation (Mi et al., March 2026)
12 ▪ Good-Enough LLM Obfuscation (GELO) (Belikov and Fedotov, March 2026)
13 ▪ Osmosis Distillation: Model Hijacking with the Fewest Samples (Shi et al., March 2026)
14 ▪ Few-shot Model Extraction Attacks against Sequential Recommender Systems (Zhang and Liu, March 2026)
15 ▪ DuFFin: A Dual-Level Fingerprinting Framework for LLMs IP Protection (Yan et al., February 2026)
16 ▪ Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious Packages (Guo et al., January 2026)
17 ▪ KinGuard: Hierarchical Kinship-Aware Fingerprinting to Defend Against Large Language Model Stealing (Xu et al., January 2026)
18 ▪ Deep Dive into the Abuse of DL APIs To Create Malicious AI Models and How to Detect Them (Nabeel and Starov, January 2026)
19 ▪ Aggressive Compression Enables LLM Weight Theft (Brown et al., January 2026)
20 ▪ Making Theft Useless: Adulteration-Based Protection of Proprietary Knowledge Graphs in GraphRAG Systems (Wang et al., January 2026)
21 ▪ DivQAT: Enhancing Robustness of Quantized Convolutional Neural Networks against Model Extraction Attacks (Khaled, Magalh\~aes, and Nicolescu, January 2026)
22 ▪ To See or Not to See -- Fingerprinting Devices in Adversarial Environments Amid Advanced Machine Learning (Feng, Haddad, and Sehatbakhsh, December 2025)
23 ▪ A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection (Li et al., December 2025)
24 ▪ AED: Automatic Discovery of Effective and Diverse Vulnerabilities for Autonomous Driving Policy with Large Language Models (Qiu et al., December 2025)
25 ▪ On Stealing Graph Neural Network Models (Podhajski et al., November 2025)
26 ▪ Quantifying the Risk of Transferred Black Box Attacks (Cox and Bunzel, November 2025)
27 ▪ SLIP-SEC: Formalizing Secure Protocols for Model IP Protection (Jain et al., October 2025)
28 ▪ $\delta$-STEAL: LLM Stealing Attack with Local Differential Privacy (Dang et al., October 2025)
29 ▪ Black Box Absorption: LLMs Undermining Innovative Ideas (Cao, October 2025)
30 ▪ When Intelligence Fails: An Empirical Study on Why LLMs Struggle with Password Cracking (Rehman et al., October 2025)
31 ▪ DistilLock: Safeguarding LLMs from Unauthorized Knowledge Distillation on the Edge (Mohanty et al., October 2025)
32 ▪ Rotation, Scale, and Translation Resilient Black-box Fingerprinting for Intellectual Property Protection of EaaS Models (Zhang et al., October 2025)
33 ▪ CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment (Li et al., October 2025)
34 ▪ LLMAtKGE: Large Language Models as Explainable Attackers against Knowledge Graph Embeddings (Li et al., October 2025)
35 ▪ Is the Hard-Label Cryptanalytic Model Extraction Really Polynomial? (Ito, Miura, and Todo, October 2025)
36 ▪ Real-VulLLM: An LLM Based Assessment Framework in the Wild (Safdar et al., October 2025)
37 ▪ From Trace to Line: LLM Agent for Real-World OSS Vulnerability Localization (Xi et al., October 2025)
38 ▪ POLAR: Automating Cyber Threat Prioritization through LLM-Powered Assessment (Tang et al., October 2025)
39 ▪ Stealing AI Model Weights Through Covert Communication Channels (Barbaza et al., October 2025)
40 ▪ Model Extraction Attacks Revisited (Liang et al., October 2025)
41 ▪ Scalable Fingerprinting of Large Language Models (Nasery et al., October 2025)
42 ▪ Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs (Cai et al., September 2025)
43 ▪ StolenLoRA: Exploring LoRA Extraction Attacks via Synthetic Data (Wang et al., September 2025)
44 ▪ PhishParrot: LLM-Driven Adaptive Crawling to Unveil Cloaked Phishing Sites (Nakano, Koide, and Chiba, August 2025)
45 ▪ RouteMark: A Fingerprint for Intellectual Property Attribution in Routing-based Model Merging (He et al., August 2025)
46 ▪ "Energon": Unveiling Transformers from GPU Power and Thermal Side-Channels (Chaudhuri et al., August 2025)
47 ▪ Leveraging Machine Learning for Botnet Attack Detection in Edge-Computing Assisted IoT Networks (Rupanetti and Kaabouch, August 2025)
48 ▪ Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report (Weerawardhena et al., August 2025)
49 ▪ Safe machine learning model release from Trusted Research Environments: The SACRO-ML package (Smith et al., August 2025)
50 ▪ Theoretically Unmasking Inference Attacks Against LDP-Protected Clients in Federated Vision Models (Nguyen et al., August 2025)
51 ▪ Medical Image De-Identification Benchmark Challenge (Pei et al., August 2025)
52 ▪ LLM-Based Identification of Infostealer Infection Vectors from Screenshots: The Case of Aurora (Ruellan, Clay, and Ascoli, August 2025)
53 ▪ Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems via Knowledge Asymmetry Exploitation (Chen et al., August 2025)
54 ▪ The Impact of Train-Test Leakage on Machine Learning-based Android Malware Detection (Liu et al., July 2025)
55 ▪ Cascading and Proxy Membership Inference Attacks (Du et al., July 2025)
56 ▪ Radio Adversarial Attacks on EMG-based Gesture Recognition Networks (Xie, July 2025)
57 ▪ Learning-based Privacy-Preserving Graph Publishing Against Sensitive Link Inference Attacks (Wu et al., July 2025)
58 ▪ Privacy-Preserving AI for Encrypted Medical Imaging: A Framework for Secure Diagnosis and Learning (Siam and Shohan, July 2025)
59 ▪ Guard-GBDT: Efficient Privacy-Preserving Approximated GBDT Training on Vertical Dataset (Song et al., July 2025)
60 ▪ Encrypted-State Quantum Compilation Scheme Based on Quantum Circuit Obfuscation (Zhang, Shang, and Guo, July 2025)
61 ▪ CompLeak: Deep Learning Model Compression Exacerbates Privacy Leakage (Li et al., July 2025)
62 ▪ SVAgent: AI Agent for Hardware Security Verification Assertion (Guo et al., July 2025)
63 ▪ DP2Guard: A Lightweight and Byzantine-Robust Privacy-Preserving Federated Learning Scheme for Industrial IoT (Han et al., July 2025)
64 ▪ zkFL: Zero-Knowledge Proof-based Gradient Aggregation for Federated Learning (Wang et al., July 2025)
65 ▪ Detecting Benchmark Contamination Through Watermarking (Sander et al., July 2025)
66 ▪ Frame-level Temporal Difference Learning for Partial Deepfake Speech Detection (Li, Zhang, and Zhao, July 2025)
67 ▪ VMask: Tunable Label Privacy Protection for Vertical Federated Learning via Layer Masking (Tan et al., July 2025)
68 ▪ Towards Efficient Privacy-Preserving Machine Learning: A Systematic Review from Protocol, Model, and System Perspectives (Zeng et al., July 2025)
69 ▪ FuSeFL: Fully Secure and Scalable Cross-Silo Federated Learning (Ghinani and Sadredini, July 2025)
70 ▪ A Privacy-Preserving Semantic-Segmentation Method Using Domain-Adaptation Technique (Sueyoshi, Nishikawa, and Kiya, July 2025)
71 ▪ A Crowdsensing Intrusion Detection Dataset For Decentralized Federated Learning Models (Feng et al., July 2025)
72 ▪ Privacy Against Agnostic Inference Attacks in Vertical Federated Learning (Varasteh, July 2025)
73 ▪ FacialMotionID: Identifying Users of Mixed Reality Headsets using Abstract Facial Motion Representations (Castro et al., July 2025)
74 ▪ AdRo-FL: Informed and Secure Client Selection for Federated Learning in the Presence of Adversarial Aggregator (Hossain et al., July 2025)
75 ▪ TimberStrike: Dataset Reconstruction Attack Revealing Privacy Leakage in Federated Tree-Based Systems (Gennaro et al., July 2025)
76 ▪ Split Happens: Combating Advanced Threats with Split Learning and Function Secret Sharing (Khan, Budzys, and Michalas, July 2025)
77 ▪ Secure and Efficient UAV-Based Face Detection via Homomorphic Encryption and Edge Computing (Duc et al., July 2025)
78 ▪ Efficient Private Inference Based on Helper-Assisted Malicious Security Dishonest Majority MPC (Wang et al., July 2025)
79 ▪ Invariant-based Robust Weights Watermark for Large Language Models (Guo et al., July 2025)
80 ▪ Research on Data Right Confirmation Mechanism of Federated Learning based on Blockchain (Cheng and Guo, July 2025)
81 ▪ FedP3E: Privacy-Preserving Prototype Exchange for Non-IID IoT Malware Detection in Cross-Silo Federated Learning (Darwish et al., July 2025)
82 ▪ A Blockchain Solution for Collaborative Machine Learning over IoT (Beis-Penedo et al., July 2025)
83 ▪ ZKTorch: Compiling ML Inference to Zero-Knowledge Proofs via Parallel Proof Accumulation (Chen, Tang, and Kang, July 2025)
84 ▪ BarkBeetle: Stealing Decision Tree Models with Fault Injection (Wang et al., July 2025)
85 ▪ Fundamental Limits of Hierarchical Secure Aggregation with Cyclic User Association (Zhang et al., July 2025)
86 ▪ Learning Federated Neural Graph Databases for Answering Complex Queries from Distributed Knowledge Graphs (Hu et al., July 2025)
87 ▪ TT-TFHE: a Torus Fully Homomorphic Encryption-Friendly Neural Network Architecture (Benamira et al., July 2025)
88 ▪ A Model Stealing Attack Against Multi-Exit Networks (Pan et al, Mar 2025)
89 ▪ Model Stealing Attack against Graph Classification with Authenticity, Uncertainty and Diversity (Zhu et al, Aug 2024)
Model Theft, Data Leakage, ML-Enabled Product or Service, and API Access

>

<

‍

Model Inversion Attack

Covers:

  • OWASP ML 03: Model Inversion Attack
  • MITRE ATLAS Exfiltration

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Inversion Attack

cybersecurity_tracker - Google Drive

2 ▪ FERMI: Exploiting Relations for Membership Inference Against Tabular Diffusion Models (Mahyar et al., May 2026)
3 ▪ Auditing Data Membership in Reinforcement Learning With Verifiable Rewards (Liu et al., May 2026)
4 ▪ Membership Inference Attacks on Vision-Language-Action Models (Peng et al., May 2026)
5 ▪ SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing (Sun et al., May 2026)
6 ▪ Pop Quiz Attack: Black-box Membership Inference Attacks Against Large Language Models (Chen et al., May 2026)
7 ▪ Membership Inference Attacks for Retrieval Based In-Context Learning for Document Question Answering (Kulkarni, Koskela, and Zumot, May 2026)
8 ▪ Membership Inference Attacks Against Video Large Language Models (Song et al., May 2026)
9 ▪ A Data-Free Membership Inference Attack on Federated Learning in Hardware Assurance (Lee et al., April 2026)
10 ▪ No More Guessing: a Verifiable Gradient Inversion Attack in Federated Learning (Diana et al., April 2026)
11 ▪ Label Leakage Attacks in Machine Unlearning: A Parameter and Inversion-Based Approach (Zheng et al., April 2026)
12 ▪ FedSpy-LLM: Towards Scalable and Generalizable Data Reconstruction Attacks from Gradients on LLMs (Meerza, Wang, and Liu, April 2026)
13 ▪ ReproMIA: A Comprehensive Analysis of Model Reprogramming for Proactive Membership Inference Attacks (Huang, Wang, and Wang, April 2026)
14 ▪ A Divide-and-Conquer Strategy for Hard-Label Extraction of Deep Neural Networks via Side-Channel Attacks (Coqueret et al., April 2026)
15 ▪ SERSEM: Selective Entropy-Weighted Scoring for Membership Inference in Code Language Models (Dikici et al., April 2026)
16 ▪ \texttt{ReproMIA}: A Comprehensive Analysis of Model Reprogramming for Proactive Membership Inference Attacks (Huang, Wang, and Wang, April 2026)
17 ▪ Automated Membership Inference Attacks: Discovering MIA Signal Computations using LLM Agents (Tran, Kotevska, and Xiong, March 2026)
18 ▪ ARES: Scalable and Practical Gradient Inversion Attack in Federated Learning through Activation Recovery (Gong et al., March 2026)
19 ▪ Exponential-Family Membership Inference: From LiRA and RMIA to BaVarIA (Br\"annvall, March 2026)
20 ▪ Revisiting the LiRA Membership Inference Attack Under Realistic Assumptions (Jebreel et al., March 2026)
21 ▪ How to Steal Reasoning Without Reasoning Traces (Zhang, Morris, and Shmatikov, March 2026)
22 ▪ Protection against Source Inference Attacks in Federated Learning (Athanasiou, Jung, and Palamidessi, March 2026)
23 ▪ No Caption, No Problem: Caption-Free Membership Inference via Model-Fitted Embeddings (Jeon et al., February 2026)
24 ▪ ImpMIA: Leveraging Implicit Bias for Membership Inference Attack (Golbari et al., February 2026)
25 ▪ LoMime: Query-Efficient Membership Inference using Model Extraction in Label-Only Settings (Oksuz, Halimi, and Ayday, February 2026)
26 ▪ Sequential Membership Inference Attacks (Michel, Basu, and Kaufmann, February 2026)
27 ▪ The Role of Learning in Attacking Intrusion Detection Systems (Domico, Ferrand, and McDaniel, February 2026)
28 ▪ Practical Feasibility of Gradient Inversion Attacks in Federated Learning (Valadi et al., February 2026)
29 ▪ Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks (Tsai et al., February 2026)
30 ▪ Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software (Kordonsky et al., February 2026)
31 ▪ Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models (Wu et al., February 2026)
32 ▪ Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective (Ganev and Cristofaro, February 2026)
33 ▪ Noise as a Probe: Membership Inference Attacks on Diffusion Models Leveraging Initial Noise (Lian et al., January 2026)
34 ▪ What Hard Tokens Reveal: Exploiting Low-confidence Tokens for Membership Inference Attacks against Large Language Models (Jawad, Xiao, and Wu, January 2026)
35 ▪ UnlearnShield: Shielding Forgotten Privacy against Unlearning Inversion (Xue et al., January 2026)
36 ▪ VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks (Tsaprazlis et al., January 2026)
37 ▪ How does Graph Structure Modulate Membership-Inference Risk for Graph Neural Networks? (Khosla, January 2026)
38 ▪ Reconstructing Training Data from Adapter-based Federated Large Language Models (Chen et al., January 2026)
39 ▪ Res-MIA: A Training-Free Resolution-Based Membership Inference Attack on Federated Learning Models (Zare and Shamsinejadbabaki, January 2026)
40 ▪ Data-Free Privacy-Preserving for LLMs via Model Inversion and Selective Unlearning (Zhou et al., January 2026)
41 ▪ Powerful Training-Free Membership Inference Against Autoregressive Language Models (Ili\'c, Stanojevi\'c, and Cvejoski, January 2026)
42 ▪ When Reasoning Leaks Membership: Membership Inference Attack on Black-box Large Reasoning Models (Hu et al., January 2026)
43 ▪ Exploring the Vulnerabilities of Federated Learning: A Deep Dive into Gradient Inversion Attacks (Guo et al., January 2026)
44 ▪ DeepLeak: Privacy Enhancing Hardening of Model Explanations Against Membership Leakage (Hmida et al., January 2026)
45 ▪ Window-based Membership Inference Attacks Against Fine-tuned Large Language Models (Chen et al., January 2026)
46 ▪ Assessing the Effectiveness of Membership Inference on Generative Music (Chow et al., December 2025)
47 ▪ Membership Inference Attack with Partial Features (Wang et al., December 2025)
48 ▪ Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks (Fan et al., December 2025)
49 ▪ In-Context Probing for Membership Inference in Fine-Tuned Language Models (Lu et al., December 2025)
50 ▪ How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code? (Yang et al., December 2025)
51 ▪ An Efficient Gradient-Based Inference Attack for Federated Learning (Monta\~na-Fern\'andez and Ortega-Fernandez, December 2025)
52 ▪ IntentMiner: Intent Inversion Attack via Tool Call Analysis in the Model Context Protocol (Yao et al., December 2025)
53 ▪ Non-Linear Trajectory Modeling for Multi-Step Gradient Inversion Attacks in Federated Learning (Xia et al., December 2025)
54 ▪ On the Effectiveness of Membership Inference in Targeted Data Extraction from Large Language Models (Sahili, Chehab, and Tajeddine, December 2025)
55 ▪ Imitative Membership Inference Attack (Du et al., December 2025)
56 ▪ Reference Recommendation based Membership Inference Attack against Hybrid-based Recommender Systems (Chi et al., December 2025)
57 ▪ Unlearning Inversion Attacks for Graph Neural Networks (Zhang et al., December 2025)
58 ▪ Lost in Modality: Evaluating the Effectiveness of Text-Based Membership Inference Attacks on Large Multimodal Models (Tong, Sun, and Nguyen, December 2025)
59 ▪ ICAS: Detecting Training Data from Autoregressive Image Generative Models (Yu et al., December 2025)
60 ▪ Rank Matters: Understanding and Defending Model Inversion Attacks via Low-Rank Feature Filtering (Yu et al., December 2025)
61 ▪ Ghosting Your LLM: Without The Knowledge of Your Gradient and Data (Almalky et al., December 2025)
62 ▪ Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage? (Amebley and Dibbo, November 2025)
63 ▪ Do Spikes Protect Privacy? Investigating Black-Box Model Inversion Attacks in Spiking Neural Networks (Poursiami, Moshruba, and Parsa, November 2025)
64 ▪ Eguard: Defending LLM Embeddings Against Inversion Attacks via Text Mutual Information Optimization (Liu et al., November 2025)
65 ▪ GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards (Liu et al., November 2025)
66 ▪ GraphToxin: Reconstructing Full Unlearned Graphs from Graph Unlearning (Song and Palanisamy, November 2025)
67 ▪ On the Detectability of Active Gradient Inversion Attacks in Federated Learning (Carletti et al., November 2025)
68 ▪ Enhanced Privacy Leakage from Noise-Perturbed Gradients via Gradient-Guided Conditional Diffusion Models (Meng et al., November 2025)
69 ▪ Safeguarding Graph Neural Networks against Topology Inference Attacks (Fu et al., November 2025)
70 ▪ Biologically-Informed Hybrid Membership Inference Attacks on Generative Genomic Models (Belfiore, Passerat-Palmbach, and Usynin, November 2025)
71 ▪ P-MIA: A Profiled-Based Membership Inference Attack on Cognitive Diagnosis Models (Hou et al., November 2025)
72 ▪ Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing (Yin et al., November 2025)
73 ▪ TextCrafter: Optimization-Calibrated Noise for Defending Against Text Embedding Inversion (Tang, Jiang, and Niu, October 2025)
74 ▪ Model Inversion Attacks Meet Cryptographic Fuzzy Extractors (Prabhakar, Xu, and Saxena, October 2025)
75 ▪ Practical Bayes-Optimal Membership Inference Attacks (Lassila et al., October 2025)
76 ▪ SPEAR++: Scaling Gradient Inversion via Sparsely-Used Dictionary Learning (Bakarsky et al., October 2025)
77 ▪ Membership Inference Attacks for Unseen Classes (Thaker et al., October 2025)
78 ▪ Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models (Li, Yu, and Xu, October 2025)
79 ▪ Fast-MIA: Efficient and Scalable Membership Inference for LLMs (Takahashi and Ishihara, October 2025)
80 ▪ GUIDE: Enhancing Gradient Inversion Attacks in Federated Learning with Denoising Models (Carletti et al., October 2025)
81 ▪ Detecting Adversarial Fine-tuning with Auditing Agents (Egler, Schulman, and Carlini, October 2025)
82 ▪ Toward Efficient Inference Attacks: Shadow Model Sharing via Mixture-of-Experts (Bai et al., October 2025)
83 ▪ ImpMIA: Leveraging Implicit Bias for Membership Inference Attack under Realistic Scenarios (Golbari et al., October 2025)
84 ▪ DiffMI: Breaking Face Recognition Privacy via Diffusion-Driven Training-Free Model Inversion (Wang et al., October 2025)
85 ▪ Diffusion-aided Task-oriented Semantic Communications with Model Inversion Attack (Wang et al., October 2025)
86 ▪ Uncovering Privacy Vulnerabilities through Analytical Gradient Inversion Attacks (Eltaras et al., September 2025)
87 ▪ Unveiling Impact of Frequency Components on Membership Inference Attacks for Diffusion Models (Lian et al., September 2025)
88 ▪ Accurate Latent Inversion for Generative Image Steganography via Rectified Flow (Qian et al., August 2025)
89 ▪ An Inversion-based Measure of Memorization for Diffusion Models (Ma et al., August 2025)
90 ▪ MASQUE: A Text-Guided Diffusion-Based Framework for Localized and Customized Adversarial Makeup (Kwon and Zhang, July 2025)
91 ▪ Noise Contrastive Estimation-based Matching Framework for Low-Resource Security Attack Pattern Recognition (Nguyen, \v{S}rndi\'c, and Neth, July 2025)
92 ▪ Boosting Ray Search Procedure of Hard-label Attacks with Transfer-based Priors (Ma et al., July 2025)
93 ▪ Tab-MIA: A Benchmark Dataset for Membership Inference Attacks on Tabular Data in LLMs (German et al., July 2025)
94 ▪ Depth Gives a False Sense of Privacy: LLM Internal States Inversion (Dong et al., July 2025)
95 ▪ Blackbox Dataset Inference for LLM (Zhou et al., July 2025)
96 ▪ Exploiting Context-dependent Duration Features for Voice Anonymization Attack Systems (Tomashenko, Vincent, and Tommasi, July 2025)
97 ▪ REAL-IoT: Characterizing GNN Intrusion Detection Robustness under Practical Adversarial Attack (Zhan, Zhou, and Haddadi, July 2025)
98 ▪ Random Erasing vs. Model Inversion: A Promising Defense or a False Hope? (Tran et al., July 2025)
99 ▪ AdvGrasp: Adversarial Attacks on Robotic Grasping from a Physical Perspective (Wang et al., July 2025)
100 ▪ Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy (Kulynych et al., July 2025)
101 ▪ Detection of Intelligent Tampering in Wireless Electrocardiogram Signals Using Hybrid Machine Learning (Deshpande, Getnet, and Dargie, July 2025)
102 ▪ Deep Learning Model Inversion Attacks and Defenses: A Comprehensive Survey (Yang et al, Apr 2025)
103 ▪ Trap-MID: Trapdoor-based Defense against Model Inversion Attacks (Liu and Chen, Nov 2024)
104 ▪ Geminio: Language-Guided Gradient Inversion Attacks in Federated Learning (Shan et al, Nov 2024)
105 ▪ MIBench: A Comprehensive Benchmark for Model Inversion Attack and Defense (Qiu et al, Oct 2024)
106 ▪ Defending against Model Inversion Attacks via Random Erasing (Tran et al, Sep 2024)
107 ▪ Improving Robustness to Model Inversion Attacks via Sparse Coding Architectures (Dibbo et al, Aug 2024)
Model Inversion Attack

>

<

‍

Exfiltration via Cyber Means

Covers:

  • MITRE ATLAS Exfiltration

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Exfiltration via Cyber Means

cybersecurity_tracker - Google Drive

2 ▪ VectorSmuggle: Steganographic Exfiltration in Embedding Stores and a Cryptographic Provenance Defense (Wanger, May 2026)
3 ▪ DECIFR: Domain-Aware Exfiltration of Circuit Information from Federated Gradient Reconstruction (Lee et al., April 2026)
4 ▪ Improving DNS Exfiltration Detection via Transformer Pretraining (Tomi\'c, Cvetanovi\'c, and Tadi\'c, April 2026)
5 ▪ FLARE: A Wireless Side-Channel Fingerprinting Attack on Federated Learning (Shuvo et al., December 2025)
6 ▪ Malicious GenAI Chrome Extensions: Unpacking Data Exfiltration and Malicious Behaviours (Seetharam, Nabeel, and Melicher, December 2025)
7 ▪ Data Exfiltration by Compression Attack: Definition and Evaluation on Medical Image Data (Li, Ayache, and Delingette, November 2025)
8 ▪ Exploiting Web Search Tools of AI Agents for Data Exfiltration (Rall et al., October 2025)
Exfiltration via Cyber Means

>

<

‍

Model Skewing Attack

Covers:

  • OWASP ML 08: Model Skewing

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Model Skewing Attack

cybersecurity_tracker - Google Drive

2 ▪ Hiding in Plain Sight: Detectability-Aware Antidistillation of Reasoning Models (Hartman et al., May 2026)
3 ▪ Conflicts Make Large Reasoning Models Vulnerable to Attacks (Liu et al., April 2026)
4 ▪ Robustness Over Time: Understanding Adversarial Examples' Effectiveness on Longitudinal Versions of Large Language Models (Liu et al., March 2026)
5 ▪ Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacks (Struppek, Gleave, and Pelrine, February 2026)
6 ▪ Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges (Ding et al., February 2026)
7 ▪ Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs (Jiang et al., February 2026)
8 ▪ From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching (Zhang et al., February 2026)
9 ▪ Beyond Denial-of-Service: The Puppeteer's Attack for Fine-Grained Control in Ranking-Based Federated Learning (Chen et al., January 2026)
10 ▪ Breaking Diffusion with Cache: Exploiting Approximate Caches in Diffusion Models (Sun, Jie, and Liu, January 2026)
11 ▪ SafeCOMM: A Study on Safety Degradation in Fine-Tuned Telecom Large Language Models (Djuhera et al., October 2025)
12 ▪ SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification (Lai et al., October 2025)
13 ▪ Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence (Gong et al., September 2025)
14 ▪ Understanding Concept Drift with Deprecated Permissions in Android Malware Detection (Sabbah et al., July 2025)
15 ▪ A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census (Abowd et al., July 2025)
16 ▪ HASSLE: A Self-Supervised Learning Enhanced Hijacking Attack on Vertical Federated Learning (He and Chang, July 2025)
17 ▪ SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents (Kutasov et al., July 2025)
18 ▪ False Alarms, Real Damage: Adversarial Attacks Using LLM-based Models on Text-based Cyber Threat Intelligence Systems (Shafee, Bessani, and Ferreira, July 2025)
Model Skewing Attack

>

<

‍

Evade ML Model

Covers:

  • MITRE ATLAS Initial Access
  • MITRE ATLAS Reconnaissance

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Evade ML Model

cybersecurity_tracker - Google Drive

2 ▪ The Role of Learning in Attacking ML-based Network Intrusion Detection (Domico, Ferrand, and McDaniel, May 2026)
3 ▪ WebTrap: Stealthy Mid-Task Hijacking of Browser Agents During Navigation (Liu et al., May 2026)
4 ▪ Benchmarking Large Language Models for IoC Recovery under Adversarial Code Obfuscation and Encryption (Morales, Pastrana, and Tapiador, May 2026)
5 ▪ Syntax- and Compilation-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, May 2026)
6 ▪ MalPurifier: Enhancing Android Malware Detection with Adversarial Purification against Evasion Attacks (Zhou et al., May 2026)
7 ▪ Trace: Unmasking AI Attack Agents Through Terminal Behavior Fingerprinting (Ediga and Chattopadhyay, May 2026)
8 ▪ Trident: Improving Malware Detection with LLMs and Behavioral Features (Saul et al., May 2026)
9 ▪ AsmRAG: LLM-Driven Malware Detection by Retrieving Functionally Similar Assembly Code (Karbab, April 2026)
10 ▪ Adversarial Evasion in Non-Stationary Malware Detection: Minimizing Drift Signals through Similarity-Constrained Perturbations (Acharya and Zhang, April 2026)
11 ▪ Hackers or Hallucinators? A Comprehensive Analysis of LLM-Based Automated Penetration Testing (Peng et al., April 2026)
12 ▪ Evasion Adversarial Attacks Remain Impractical Against ML-based Network Intrusion Detection Systems, Especially Dynamic Ones (elShehaby and Matrawy, March 2026)
13 ▪ ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks (Elnawawy et al., March 2026)
14 ▪ Unveiling the Resilience of LLM-Enhanced Search Engines against Black-Hat SEO Manipulation (Chen et al., March 2026)
15 ▪ Targeted Adversarial Traffic Generation : Black-box Approach to Evade Intrusion Detection Systems in IoT Networks (Debicha et al., March 2026)
16 ▪ Evasive Intelligence: Lessons from Malware Analysis for Evaluating AI Agents (Aonzo et al., March 2026)
17 ▪ FuzzySQL: Uncovering Hidden Vulnerabilities in DBMS Special Features with LLM-Driven Fuzzing (Chen et al., February 2026)
18 ▪ Anticipating Adversary Behavior in DevSecOps Scenarios through Large Language Models (Caballero et al., February 2026)
19 ▪ SmartGuard: Leveraging Large Language Models for Network Attack Detection through Audit Log Analysis and Summarization (Zhang et al., February 2026)
20 ▪ Unknown Attack Detection in IoT Networks using Large Language Models: A Robust, Data-efficient Approach (Ali et al., February 2026)
21 ▪ StealthRL: Reinforcement Learning Paraphrase Attacks for Multi-Detector Evasion of AI-Text Detectors (Ranganath and Ramesh, February 2026)
22 ▪ DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems (Ou et al., February 2026)
23 ▪ Beyond Imprecise Distance Metrics: Trace-Guided Directed Greybox Fuzzing via LLM-Predicted Call Stacks (Zhang and Zhang, February 2026)
24 ▪ "Someone Hid It": Query-Agnostic Black-Box Attacks on LLM-Based Retrieval (Li et al., February 2026)
25 ▪ Semantics-Preserving Evasion of LLM Vulnerability Detectors (Sun, Oprea, and Wong, February 2026)
26 ▪ In Vino Veritas and Vulnerabilities: Examining LLM Safety via Drunk Language Inducement (Shetty, Joshi, and Kanhere, February 2026)
27 ▪ AEGIS: White-Box Attack Path Generation using LLMs and Training Effectiveness Evaluation for Large-Scale Cyber Defence Exercises (Tung et al., February 2026)
28 ▪ ICL-EVADER: Zero-Query Black-Box Evasion Attacks on In-Context Learning and Their Defenses (He et al., January 2026)
29 ▪ CODE: A Contradiction-Based Deliberation Extension Framework for Overthinking Attacks on Retrieval-Augmented Generation (Zhang et al., January 2026)
30 ▪ A Decompilation-Driven Framework for Malware Detection with Large Language Models (Chawla and Prasad, January 2026)
31 ▪ MASH: Evading Black-Box AI-Generated Text Detectors via Style Humanization (Gu, Li, and Hu, January 2026)
32 ▪ VIPER Strike: Defeating Visual Reasoning CAPTCHAs via Structured Vision-Language Inference (Qi et al., January 2026)
33 ▪ Cracking IoT Security: Can LLMs Outsmart Static Analysis Tools? (Quantrill et al., January 2026)
34 ▪ Improving the Convergence Rate of Ray Search Optimization for Query-Efficient Hard-Label Attacks (Xu et al., December 2025)
35 ▪ Automated Penetration Testing with LLM Agents and Classical Planning (Wang et al., December 2025)
36 ▪ NetDeTox: Adversarial and Efficient Evasion of Hardware-Security GNNs via RL-LLM Orchestration (Wang et al., December 2025)
37 ▪ ExtendAttack: Attacking Servers of LRMs via Extending Reasoning (Zhu et al., November 2025)
38 ▪ Adversarial Agents: Black-Box Evasion Attacks with Reinforcement Learning (Domico et al., November 2025)
39 ▪ MalRAG: A Retrieval-Augmented LLM Framework for Open-set Malicious Traffic Identification (Luo et al., November 2025)
40 ▪ Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment (Zhang and sun, November 2025)
41 ▪ GradEscape: A Gradient-Based Evader Against AI-Generated Text Detectors (Meng et al., October 2025)
42 ▪ The RAG Paradox: A Black-Box Attack Exploiting Unintentional Vulnerabilities in Retrieval-Augmented Generation Systems (Choi et al., October 2025)
43 ▪ Adversarial Pre-Padding: Generating Evasive Network Traffic Against Transformer-Based Classifiers (Jing et al., October 2025)
44 ▪ Detecting Various DeFi Price Manipulations with LLM Reasoning (Zhong et al., October 2025)
45 ▪ Network Intrusion Detection: Evolution from Conventional Approaches to LLM Collaboration and Emerging Risks (Feng and Sakurai, October 2025)
46 ▪ Black-Box Evasion Attacks on Data-Driven Open RAN Apps: Tailored Design and Experimental Evaluation (Gajjar et al., October 2025)
47 ▪ From Flows to Words: Can Zero-/Few-Shot LLMs Detect Network Intrusions? A Grammar-Constrained, Calibrated Evaluation on UNSW-NB15 (Rehman et al., October 2025)
48 ▪ SoK: Adversarial Evasion Attacks Practicality in NIDS Domain and the Impact of Dynamic Learning (elShehaby and Matrawy, October 2025)
49 ▪ A Hard-Label Black-Box Evasion Attack against ML-based Malicious Traffic Detection Systems (Liu et al., October 2025)
50 ▪ VulSolver: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., October 2025)
51 ▪ Can LLMs Handle WebShell Detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework (Han et al., October 2025)
52 ▪ Enhancing GraphQL Security by Detecting Malicious Queries Using Large Language Models, Sentence Transformers, and Convolutional Neural Networks (Engineering et al., October 2025)
53 ▪ Complete Evasion, Zero Modification: PDF Attacks on AI Text Detection (Creo, August 2025)
54 ▪ AdVAR-DNN: Adversarial Misclassification Attack on Collaborative DNN Inference (Yousefi, Mounesan, and Debroy, August 2025)
55 ▪ ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models (Yook et al., July 2025)
56 ▪ Hierarchical Graph Neural Network for Compressed Speech Steganalysis (Hemis et al., July 2025)
57 ▪ Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms (Sato, Ohki, and Nishigaki, July 2025)
58 ▪ GATEBLEED: Exploiting On-Core Accelerator Power Gating for High Performance & Stealthy Attacks on AI (Kalyanapu et al., July 2025)
59 ▪ BandFuzz: An ML-powered Collaborative Fuzzing Framework (Shi et al., July 2025)
60 ▪ PotentRegion4MalDetect: Advanced Features from Potential Malicious Regions for Malware Detection (Koppanati, Santra, and Peddoju, July 2025)
61 ▪ Humanizing the Machine: Proxy Attacks to Mislead LLM Detectors (Wang et al, Oct 2024)
Evade ML Model

>

<

‍

Discover ML Artifacts, Data from Information Repositories and Local System, and Acquire Public ML Artifacts

Covers:

  • MITRE ATLAS Resource Development, Discovery, and Collection

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Discover ML Model Family and Ontology/Model Extraction

cybersecurity_tracker - Google Drive

2 ▪ Enabling Transparent Cyber Threat Intelligence Combining Large Language Models and Domain Ontologies (Cotti et al., April 2026)
3 ▪ Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs (Liu, Evans, and Xiong, April 2026)
4 ▪ CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models (Lee, Koo, and Kwak, April 2026)
5 ▪ AttnDiff: Attention-based Differential Fingerprinting for Large Language Models (Zhang et al., April 2026)
6 ▪ Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test (Zhu et al., March 2026)
7 ▪ Navigating the Deep: End-to-End Extraction on Deep Neural Networks (Liu et al., February 2026)
8 ▪ A Behavioral Fingerprint for Large Language Models: Provenance Tracking via Refusal Vectors (Xu and Sheng, February 2026)
9 ▪ FDLLM: A Dedicated Detector for Black-Box LLMs Fingerprinting (Fu et al., January 2026)
10 ▪ Identifying Models Behind Text-to-Image Leaderboards (Naseh et al., January 2026)
11 ▪ Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral Signatures (Wang et al., December 2025)
12 ▪ A Fingerprint for Large Language Models (Yang and Wu, December 2025)
13 ▪ SELF: A Robust Singular Value and Eigenvalue Approach for LLM Fingerprinting (Zhang and Zheng, December 2025)
14 ▪ A Systematic Study of Model Extraction Attacks on Graph Foundation Models (Xu et al., November 2025)
15 ▪ Uncovering Pretraining Code in LLMs: A Syntax-Aware Attribution Approach (Li et al., November 2025)
16 ▪ Ghost in the Transformer: Tracing LLM Lineage with SVD-Fingerprint (Wang et al., November 2025)
17 ▪ Structuring Security: A Survey of Cybersecurity Ontologies, Semantic Log Processing, and LLMs Application (Louren\c{c}o et al., October 2025)
18 ▪ MalCVE: Malware Detection and CVE Association Using Large Language Models (Cristea, Molnes, and Li, October 2025)
19 ▪ Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient Estimation (Shao et al., October 2025)
20 ▪ SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From (Tong et al., October 2025)
21 ▪ LLM-Assisted Model-Based Fuzzing of Protocol Implementations (Huang, Wang, and Zhou, August 2025)
22 ▪ PrompTrend: Continuous Community-Driven Vulnerability Discovery and Assessment for Large Language Models (Gasmi et al., July 2025)
23 ▪ Evaluating Ensemble and Deep Learning Models for Static Malware Detection with Dimensionality Reduction Using the EMBER Dataset (Abedin and Mehrub, July 2025)
24 ▪ Revisiting Pre-trained Language Models for Vulnerability Detection (Li et al., July 2025)
25 ▪ Specification and Evaluation of Multi-Agent LLM Systems -- Prototype and Cybersecurity Applications (H\"arer, July 2025)
26 ▪ TBDetector:Transformer-Based Detector for Advanced Persistent Threats with Provenance Graph (Wang et al., July 2025)
27 ▪ Toward an Intent-Based and Ontology-Driven Autonomic Security Response in Security Orchestration Automation and Response (Huang et al., July 2025)
28 ▪ SAMEP: A Secure Protocol for Persistent Context Sharing Across AI Agents (Masoor, July 2025)
29 ▪ UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification (Cai et al., July 2025)
30 ▪ AICrypto: A Comprehensive Benchmark For Evaluating Cryptography Capabilities of Large Language Models (Wang et al., July 2025)
31 ▪ BISON: Blind Identification with Stateless scOped pseudoNyms (Heher, More, and Heimberger, July 2025)
32 ▪ Bridging AI and Software Security: A Comparative Vulnerability Assessment of LLM Agent Deployment Paradigms (Gasmi et al., July 2025)
33 ▪ From Counterfactuals to Trees: Competitive Analysis of Model Extraction Attacks (Khouna, Ferry, and Vidal, July 2025)
34 ▪ One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs<br><br> (Li et al., May 2025)
35 ▪ Silent Leaks: Implicit Knowledge Extraction Attack on RAG Systems through Benign Queries (Wang et al, May 2025)
36 ▪ How to Backdoor the Knowledge Distillation (Wu et al, Apr 2025)
37 ▪ Knowledge Distillation-Based Model Extraction Attack using GAN-based Private Counterfactual Explanations (Ezzeddine, Ayoub, and Giordano, Oct 2024)
38 ▪ Efficient and Effective Model Extraction (Zhu et al, Sep 2024)
39 ▪ CaBaGe: Data-Free Model Extraction using ClAss BAlanced Generator Ensemble (Rosenthal et al, Sep 2024)
40 ▪ Alignment-Aware Model Extraction Attacks on Large Language Models (Liang et al, Sep 2024)
Discover ML Model Family and Ontology/Model Extraction

>

<

‍

User Execution, Command and Scripting Interpreter

Covers:

  • MITRE ATLAS Execution

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : User Execution, Command and Scripting Interpreter

cybersecurity_tracker - Google Drive

2 ▪ Post-Training Local LLM Agents for Linux Privilege Escalation with Verifiable Rewards (Normann et al., March 2026)
3 ▪ RapidPen: Fully Automated IP-to-Shell Penetration Testing with LLM-based Agents (Nakatani, February 2026)
4 ▪ CHAI: Command Hijacking against embodied AI (Burbano et al., October 2025)
User Execution, Command and Scripting Interpreter

>

<

‍

Physical Model Access and Full Model Access

Covers:

  • MITRE ATLAS ML Model Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Physical Model Access and Full Model Access

cybersecurity_tracker - Google Drive

2 ▪ Shape and Substance: Dual-Layer Side-Channel Attacks on Local Vision-Language Models (Hadad and Guri, March 2026)
3 ▪ T2I-Based Physical-World Appearance Attack against Traffic Sign Recognition Systems in Autonomous Driving (Ma et al., November 2025)
4 ▪ SleepWalk: Exploiting Context Switching and Residual Power for Physical Side-Channel Attacks (Sanjaya, Jayasena, and Mishra, July 2025)
5 ▪ Rainbow Artifacts from Electromagnetic Signal Injection Attacks on Image Sensors (Zhang et al., July 2025)
Physical Model Access and Full Model Access

>

<

‍

Valid Accounts

Covers:

  • MITRE ATLAS Initial Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Valid Accounts

cybersecurity_tracker - Google Drive

Valid Accounts

>

<

‍

Exploit Public Facing Application

Covers:

  • MITRE ATLAS Initial Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Exploit Public Facing Application

cybersecurity_tracker - Google Drive

2 ▪ An Empirical Study on Remote Code Execution in Machine Learning Model Hosting Ecosystems (Siddiq et al., January 2026)
Exploit Public Facing Application

>

<

‍

Threats from AI Model

Misinformation

Covers:

  • MITRE ATLAS Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Misinformation

cybersecurity_tracker - Google Drive

2 ▪ Day-to-Day Traffic Network Modeling under Route-Guidance Misinformation: Endogenous Trust and Resilience in CAV Environments (Ka and Ukkusuri, May 2026)
3 ▪ CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-Bunking of Online Misinformation (Loth, Kappes, and Pahl, April 2026)
4 ▪ The Synthetic Media Shift: Tracking the Rise, Virality, and Detectability of AI-Generated Multimodal Misinformation (Chrysidis, Papadopoulos, and Papadopoulos, April 2026)
5 ▪ Toward Accountable AI-Generated Content on Social Platforms: Steganographic Attribution and Multimodal Harm Detection (Guan et al., April 2026)
6 ▪ "That's another doom I haven't thought about": A User Study on AI Labels as a Safeguard Against Image-Based Misinformation (H\"oltervennhoff et al., March 2026)
7 ▪ FakeZero: Real-Time, Privacy-Preserving Misinformation Detection for Facebook and X (Essahli et al., October 2025)
8 ▪ Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks (Aldahoul and Zaki, October 2025)
9 ▪ Fake or Real: The Impostor Hunt in Texts for Space Operations (Kaczmarek et al., July 2025)
Misinformation

>

<

‍

Over Reliance on LLM Outputs and External (Social) Harms

Covers:

  • OWASP LLM 09: Overreliance
  • MITRE ATLAS Impact

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Over Reliance on LLM Outputs and External (Social) Harms

cybersecurity_tracker - Google Drive

2 ▪ False Security Confidence in Benign LLM Code Generation (Ren, April 2026)
3 ▪ Measuring and Exploiting Confirmation Bias in LLM-Assisted Security Code Review (Mitropoulos et al., March 2026)
4 ▪ Can Developers rely on LLMs for Secure IaC Development? (Firouzi, Bhatt, and Ghafari, February 2026)
5 ▪ Exploring the Secondary Risks of Large Language Models (Chen et al., January 2026)
6 ▪ Large Language Models Are Unreliable for Cyber Threat Intelligence (Mezzi, Massacci, and Tuma, November 2025)
7 ▪ Learned, Lagged, LLM-splained: LLM Responses to End User Security Questions (Prakash et al., October 2025)
8 ▪ Cyber-Zero: Training Cybersecurity Agents without Runtime (Zhuo et al., August 2025)
9 ▪ Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions (Qu et al., July 2025)
10 ▪ Verification Cost Asymmetry in Cognitive Warfare: A Complexity-Theoretic Framework (Luberisse, July 2025)
11 ▪ Decentralized AI-driven IoT Architecture for Privacy-Preserving and Latency-Optimized Healthcare in Pandemic and Critical Care Scenarios (Sammangi et al., July 2025)
12 ▪ Differential Privacy in Kernelized Contextual Bandits via Random Projections (Pavlovic, Salgia, and Zhao, July 2025)
13 ▪ Large Language Models are Unreliable for Cyber Threat Intelligence (Mezzi, Massacci, and Tuma, July 2025)
14 ▪ X Hacking: The Threat of Misguided AutoML (Sharma et al., July 2025)
15 ▪ Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025 (Ying et al., July 2025)
16 ▪ Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress? (Ren et al, Jul 2024)
Over Reliance on LLM Outputs and External (Social) Harms

>

<

‍

Fake Resources and Phishing

Covers:

  • MITRE ATLAS Initial Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Fake Resources and Phishing

cybersecurity_tracker - Google Drive

2 ▪ Phishing the Phishers with SpecularNet: Hierarchical Graph Autoencoding for Reference-Free Web Phishing Detection (Song, Casas, and Meo, March 2026)
3 ▪ Assessing Spear-Phishing Website Generation in Large Language Model Coding Agents (Malloy and Bissyande, February 2026)
4 ▪ CIC-Trap4Phish: A Unified Multi-Format Dataset for Phishing and Quishing Attachment Detection (Nejati et al., February 2026)
5 ▪ User-Centric Phishing Detection: A RAG and LLM-Based Approach (Barwani, Korba, and Anwar, January 2026)
6 ▪ PhishLumos: An Adaptive Multi-Agent System for Proactive Phishing Campaign Mitigation (Chiba, Nakano, and Koide, January 2026)
7 ▪ Phishing Email Detection Using Large Language Models (Hasan et al., December 2025)
8 ▪ LLM-PEA: Leveraging Large Language Models Against Phishing Email Attacks (Hassan et al., December 2025)
9 ▪ Deep Reinforcement Learning for Phishing Detection with Transformer-Based Semantic Features (Faisal, December 2025)
10 ▪ Constructing and Benchmarking: a Labeled Email Dataset for Text-Based Phishing and Spam Detection Framework (Toth, Bisztray, and Dubniczky, November 2025)
11 ▪ Small Language Models for Phishing Website Detection: Cost, Performance, and Privacy Trade-Offs (Goldenits et al., November 2025)
12 ▪ Can MLLMs Detect Phishing? A Comprehensive Security Benchmark Suite Focusing on Dynamic Threats and Multimodal Evaluation in Academic Environments (Zhou, November 2025)
13 ▪ How Can We Effectively Use LLMs for Phishing Detection?: Evaluating the Effectiveness of Large Language Model-based Phishing Detection Models (Ji and Kim, November 2025)
14 ▪ MeAJOR Corpus: A Multi-Source Dataset for Phishing Email Detection ((GECAD et al., July 2025)
15 ▪ A Login Page Transparency and Visual Similarity Based Zero Day Phishing Defense Protocol (Varshney et al., July 2025)
16 ▪ From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks (Kulkarni et al, Jul 2024)
Fake Resources and Phishing

>

<

‍

Social Manipulation

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Social Manipulation

cybersecurity_tracker - Google Drive

2 ▪ Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces (Zhang et al., May 2026)
3 ▪ A Synthetic Conversational Smishing Dataset for Social Engineering Detection (Lochstampfor and Roy, April 2026)
4 ▪ Synthetic Trust Attacks: Modeling How Generative AI Manipulates Human Decisions in Social Engineering Fraud (Ashraf, April 2026)
5 ▪ Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems (Huang et al., March 2026)
6 ▪ When Bots Take the Bait: Exposing and Mitigating the Emerging Social Engineering Attack in Web Automation Agent (Wu et al., January 2026)
7 ▪ Love, Lies, and Language Models: Investigating AI's Role in Romance-Baiting Scams (Gressel et al., December 2025)
8 ▪ AI-in-the-Loop: Privacy Preserving Real-Time Scam Detection and Conversational Scambaiting by Leveraging LLMs and Federated Learning (Hossain et al., November 2025)
9 ▪ MURMUR: Using cross-user chatter to break collaborative language agents in groups (Patlan et al., November 2025)
10 ▪ Investigating the Impact of Dark Patterns on LLM-Based Web Agents (Ersoy et al., October 2025)
11 ▪ Exploring User Security and Privacy Attitudes and Concerns Toward the Use of General-Purpose LLM Chatbots for Mental Health (Kwesi et al., July 2025)
12 ▪ Real AI Agents with Fake Memories: Fatal Context Manipulation Attacks on Web3 Agents (Patlan et al., July 2025)
13 ▪ PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety (Zhang et al, Aug 2024)
Social Manipulation

>

<

‍

Deep Fakes, Content Provenance, and Watermarking

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Deep Fakes, Content Provenance, and Watermarking

cybersecurity_tracker - Google Drive

2 ▪ RLCracker: Evaluating the Worst-Case Vulnerability of LLM Watermarks with Adaptive RL Attacks (Huang et al., May 2026)
3 ▪ Watermarking Game-Playing Agents in Perfect-Information Extensive-Form Games (Kim, Fang, and Sandholm, May 2026)
4 ▪ DeePen: Penetration Testing for Audio Deepfake Detection (M\"uller et al., May 2026)
5 ▪ Watermarking Should Be Treated as a Monitoring Primitive (Aremu, Lukas, and Zhang, May 2026)
6 ▪ TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection (Sander et al., May 2026)
7 ▪ The Deepfakes We Missed: We Built Detectors for a Threat That Didn't Arrive (Raza, May 2026)
8 ▪ Every Bit, Everywhere, All at Once: A Binomial Multibit LLM Watermark (Gloaguen et al., May 2026)
9 ▪ Sequential Behavioral Watermarking for LLM Agents (An et al., May 2026)
10 ▪ PASA: A Principled Embedding-Space Watermarking Approach for LLM-Generated Text under Semantic-Invariant Attacks (Ai and He, May 2026)
11 ▪ Majority Bit-Aware Watermarking For Large Language Models (Xu et al., May 2026)
12 ▪ Robust Spectral Watermark for Synthetic Tabular Data (Zhao et al., May 2026)
13 ▪ Evidence-based Decision Modeling for Synthetic Face Detection with Uncertainty-driven Active Learning (Jiang et al., May 2026)
14 ▪ "Training robust watermarking model may hurt authentication!'' Exploring and Mitigating the Identity Leakage in Robust Watermarking (Zhang et al., May 2026)
15 ▪ Removing the Watermark Is Not Enough: Forensic Stealth in Generative-AI Watermark Removal (Goonatilake and Ateniese, May 2026)
16 ▪ MirrorMark: Generalizable Mirrored Sampling for Multi-bit LLM Watermarking (Jiang et al., May 2026)
17 ▪ Vaporizer: Breaking Watermarking Schemes for Large Language Model Outputs (Ng, Ngo, and Chattopadhyay, May 2026)
18 ▪ Guidance Watermarking for Diffusion Models (Gesny et al., May 2026)
19 ▪ Secure Seed-Based Multi-bit Watermarking for Diffusion Models from First Principles (Gesny and Giboulot, May 2026)
20 ▪ SWAN: Semantic Watermarking with Abstract Meaning Representation (Ye et al., May 2026)
21 ▪ SPRINT: Robust Model Attribution of Generated Images via Secret Pixel Reconstruction (Yao and Juarez, May 2026)
22 ▪ MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech (Jin et al., May 2026)
23 ▪ VertMark: A Unified Training-Free Robust Watermarking Framework for Vertical Domain Pre-trained Language Models (Kong et al., May 2026)
24 ▪ Selfie-Capture Dynamics as an Auxiliary Signal Against Deepfakes and Injection Attacks for Mobile Identity Verification (Rantahalvari et al., May 2026)
25 ▪ Tell-Tale Watermarks for Explanatory Reasoning in Synthetic Media Forensics (Chang and Echizen, May 2026)
26 ▪ VOW: Verifiable and Oblivious Watermark Detection for Large Language Models (Luan et al., May 2026)
27 ▪ R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models (Zhang et al., April 2026)
28 ▪ DeepSignature: Digitally Signed, Content-Encoding Watermarks for Robust and Transparent Image Authentication (Graf et al., April 2026)
29 ▪ PoLO: Proof-of-Learning and Proof-of-Ownership at Once with Chained Watermarking (Deng et al., April 2026)
30 ▪ ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encoders (Jiang et al., April 2026)
31 ▪ SSG: Logit-Balanced Vocabulary Partitioning for LLM Watermarking (Gu, Du, and Grundy, April 2026)
32 ▪ Dual-Guard: Dual-Channel Latent Watermarking for Provenance and Tamper Localization in Diffusion Images (Xie et al., April 2026)
33 ▪ CLASP: Training-Free LLM-Assisted Source Code Watermarking via Semantic-Preserving Transformations (Xu et al., April 2026)
34 ▪ Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking (Nemecek et al., April 2026)
35 ▪ TimeMark: A Trustworthy Time Watermarking Framework for Exact Generation-Time Recovery from AIGC (Che, Du, and Gao, April 2026)
36 ▪ Can we Watermark Low-Entropy LLM Outputs? (Mazor, Morgan, and Pass, April 2026)
37 ▪ On the Robustness of Watermarking for Autoregressive Image Generation (M\"uller et al., April 2026)
38 ▪ RLSpoofer: A Lightweight Evaluator for LLM Watermark Spoofing Resilience (Huang et al., April 2026)
39 ▪ Beyond A Fixed Seal: Adaptive Stealing Watermark in Large Language Models (Zhang et al., April 2026)
40 ▪ SEED: A Large-Scale Benchmark for Provenance Tracing in Sequential Deepfake Facial Edits (Hoi et al., April 2026)
41 ▪ Towards Better Statistical Understanding of Watermarking LLMs (Cai et al., April 2026)
42 ▪ XMark: Reliable Multi-Bit Watermarking for LLM-Generated Texts (Xu et al., April 2026)
43 ▪ An End-to-End Model for Logits-Based Large Language Models Watermarking (Wong et al., April 2026)
44 ▪ Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors (Stan\v{e}k et al., April 2026)
45 ▪ Refined Detection for Gumbel Watermarking (Lattimore, April 2026)
46 ▪ SHIFT: Stochastic Hidden-Trajectory Deflection for Removing Diffusion-based Watermark (Bao et al., April 2026)
47 ▪ Gaussian Shannon: High-Precision Diffusion Model Watermarking Based on Communication (Zhang, Huang, and Zhang, March 2026)
48 ▪ NOWA: Null-space Optical Watermark for Invisible Capture Fingerprinting and Tamper Localization (Vargas et al., March 2026)
49 ▪ Robust Safety Monitoring of Language Models via Activation Watermarking (Aremu et al., March 2026)
50 ▪ Functional Subspace Watermarking for Large Language Models (Ding et al., March 2026)
51 ▪ Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing (Chen et al., March 2026)
52 ▪ Rethinking LLM Watermark Detection in Black-Box Settings: A Non-Intrusive Third-Party Framework (Wang et al., March 2026)
53 ▪ TableMark: A Multi-bit Watermark for Synthetic Tabular Data (Xia et al., March 2026)
54 ▪ Editing Away the Evidence: Diffusion-Based Image Manipulation and the Failure Modes of Robust Watermarking (Qi et al., March 2026)
55 ▪ SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking (Gao et al., March 2026)
56 ▪ EmbTracker: Traceable Black-box Watermarking for Federated Language Models (Zhao et al., March 2026)
57 ▪ Cluster-Aware Attacks on Graph Watermarks (Nemecek, Yilmaz, and Ayday, March 2026)
58 ▪ Na\"ive Exposure of Generative AI Capabilities Undermines Deepfake Detection (Kim et al., March 2026)
59 ▪ The Orthogonal Vulnerabilities of Generative AI Watermarks: A Comparative Empirical Benchmark of Spatial and Latent Provenance (Yu and Wei, March 2026)
60 ▪ ShapeMark: Robust and Diversity-Preserving Watermarking for Diffusion Models (Qian et al., March 2026)
61 ▪ mAVE: A Watermark for Joint Audio-Visual Generation Models (Si, Pan, and Wen, March 2026)
62 ▪ When Denoising Becomes Unsigning: Theoretical and Empirical Analysis of Watermark Fragility Under Diffusion-Based Image Editing (Gu et al., March 2026)
63 ▪ How Effective Are Publicly Accessible Deepfake Detection Tools? A Comparative Evaluation of Open-Source and Free-to-Use Platforms (Rettinger et al., March 2026)
64 ▪ On Google's SynthID-Text LLM Watermarking System: Theoretical Analysis and Empirical Validation (Omidi, Dong, and Wang, March 2026)
65 ▪ Watermarking Without Standards Is Not AI Governance (Nemecek, Jiang, and Ayday, March 2026)
66 ▪ Topic-Based Watermarks for Large Language Models (Nemecek, Jiang, and Ayday, March 2026)
67 ▪ Scores Know Bobs Voice: Speaker Impersonation Attack (Hwang et al., March 2026)
68 ▪ Authenticated Contradictions from Desynchronized Provenance and Watermarking (Nemecek et al., March 2026)
69 ▪ PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints (Huo et al., March 2026)
70 ▪ SKeDA: A Generative Watermarking Framework for Text-to-video Diffusion Models (Yang et al., March 2026)
71 ▪ Hide&Seek: Remove Image Watermarks with Negligible Cost via Pixel-wise Reconstruction (Chen et al., March 2026)
72 ▪ LLM-Text Watermarking based on Lagrange Interpolation (Janas, Morawiecki, and Pieprzyk, February 2026)
73 ▪ WaterVIB: Learning Minimal Sufficient Watermark Representations via Variational Information Bottleneck (He et al., February 2026)
74 ▪ Vanishing Watermarks: Diffusion-Based Image Editing Undermines Robust Invisible Watermarking (Guo et al., February 2026)
75 ▪ Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models (He et al., February 2026)
76 ▪ Can You Tell It's AI? Human Perception of Synthetic Voices in Vishing Scenarios (Bhatti et al., February 2026)
77 ▪ Watermarking LLM Agent Trajectories (Meng et al., February 2026)
78 ▪ MarkSweep: A No-box Removal Attack on AI-Generated Image Watermarking via Noise Intensification and Frequency-aware Denoising (Cao et al., February 2026)
79 ▪ Unforgeable Watermarks for Language Models via Robust Signatures (Lin, Shahabi, and Song, February 2026)
80 ▪ TriniMark: A Robust Generative Speech Watermarking Method for Trinity-Level Traceability (Li et al., February 2026)
81 ▪ MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages (Cui et al., February 2026)
82 ▪ MetaSeal: Defending Against Image Attribution Forgery Through Content-Dependent Cryptographic Watermarks (Zhou et al., February 2026)
83 ▪ More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles (Chen et al., February 2026)
84 ▪ MerkleSpeech: Public-Key Verifiable, Chunk-Localised Speech Provenance via Perceptual Fingerprints and Merkle Commitments (Ono, February 2026)
85 ▪ AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models (Li et al., February 2026)
86 ▪ On Protecting Agentic Systems' Intellectual Property via Watermarking (Wang et al., February 2026)
87 ▪ A Unified Framework for LLM Watermarks (Gloaguen et al., February 2026)
88 ▪ SynthForensics: A Multi-Generator Benchmark for Detecting Synthetic Video Deepfakes (Leotta et al., February 2026)
89 ▪ Origin Lens: A Privacy-First Mobile Framework for Cryptographic Image Provenance and AI Detection (Loth et al., February 2026)
90 ▪ Position: 3D Gaussian Splatting Watermarking Should Be Scenario-Driven and Threat-Model Explicit (Deng, Nakra, and Wu, February 2026)
91 ▪ MIRROR: Manifold Ideal Reference ReconstructOR for Generalizable AI-Generated Image Detection (Liu et al., February 2026)
92 ▪ WorldCup Sampling for Multi-bit LLM Watermarking (Wang et al., February 2026)
93 ▪ MarkCleaner: High-Fidelity Watermark Removal via Imperceptible Micro-Geometric Perturbation (Kong et al., February 2026)
94 ▪ Provenance Verification of AI-Generated Images via a Perceptual Hash Registry Anchored on Blockchain (Mohit, Aggarwal, and Gondhalekar, February 2026)
95 ▪ Color Matters: Demosaicing-Guided Color Correlation Training for Generalizable AI-Generated Image Detection (Zhong, Xu, and Zou, February 2026)
96 ▪ VocBulwark: Towards Practical Generative Speech Watermarking via Additional-Parameter Injection (Liu, Li, and Yin, February 2026)
97 ▪ MirrorMark: A Distortion-Free Multi-Bit Watermark for Large Language Models (Jiang et al., February 2026)
98 ▪ SemBind: Binding Diffusion Watermarks to Semantics Against Black-Box Forgery Attacks (Zhang et al., January 2026)
99 ▪ SWA-LDM: Toward Stealthy Watermarks for Latent Diffusion Models (Yang et al., January 2026)
100 ▪ Watermark-based Attribution of AI-Generated Content (Jiang et al., January 2026)
101 ▪ DeMark: A Query-Free Black-Box Attack on Deepfake Watermarking Defenses (Song et al., January 2026)
102 ▪ Is Your Writing Being Mimicked by AI? Unveiling Imitation with Invisible Watermarks in Creative Writing (Zhang et al., January 2026)
103 ▪ Learning to Watermark in the Latent Space of Generative Models (Rebuffi et al., January 2026)
104 ▪ DRGW: Learning Disentangled Representations for Robust Graph Watermarking (Li et al., January 2026)
105 ▪ GenPTW: Latent Image Watermarking for Provenance Tracing and Tamper Localization (Gan et al., January 2026)
106 ▪ Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction Attacks (Xu et al., January 2026)
107 ▪ Semantic Differentiation for Tackling Challenges in Watermarking Low-Entropy Constrained Generation Outputs (Le, Ritter, and Goyal, January 2026)
108 ▪ Beyond Known Fakes: Generalized Detection of AI-Generated Images via Post-hoc Distribution Alignment (Wang et al., January 2026)
109 ▪ Attack-Resistant Watermarking for AIGC Image Forensics via Diffusion-based Semantic Deflection (Liu et al., January 2026)
110 ▪ Deepfake detectors are DUMB: A benchmark to assess adversarial training robustness under transferability constraints (Serrano, Umlil, and Thomas, January 2026)
111 ▪ AgentMark: Utility-Preserving Behavioral Watermarking for Agents (Huang et al., January 2026)
112 ▪ Vulnerabilities of Audio-Based Biometric Authentication Systems Against Deepfake Speech Synthesis (Hong et al., January 2026)
113 ▪ SoK: Are Watermarks in LLMs Ready for Deployment? (Dang et al., December 2025)
114 ▪ Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform (Zhang, Li, and Gu, December 2025)
115 ▪ Pixel Seal: Adversarial-only training for invisible image and video watermarking (Sou\v{c}ek et al., December 2025)
116 ▪ How Good is Post-Hoc Watermarking With Language Model Rephrasing? (Fernandez et al., December 2025)
117 ▪ Protecting Deep Neural Network Intellectual Property with Chaos-Based White-Box Watermarking (B et al., December 2025)
118 ▪ DualGuard: Dual-stream Large Language Model Watermarking Defense against Paraphrase and Spoofing Attack (Li et al., December 2025)
119 ▪ Remotely Detectable Robot Policy Watermarking (Amir, Flageat, and Prorok, December 2025)
120 ▪ ComMark: Covert and Robust Black-Box Model Watermarking with Compressed Samples (Yang et al., December 2025)
121 ▪ CODE ACROSTIC: Robust Watermarking for Code Generation (Lin et al., December 2025)
122 ▪ SPDMark: Selective Parameter Displacement for Robust Video Watermarking (Fares, Tastan, and Nandakumar, December 2025)
123 ▪ Security and Detectability Analysis of Unicode Text Watermarking Methods Against Large Language Models (Hellmeier, December 2025)
124 ▪ UniMark: Artificial Intelligence Generated Content Identification Toolkit (Li et al., December 2025)
125 ▪ Lightweight Model Attribution and Detection of Synthetic Speech via Audio Residual Fingerprints (Pizarro et al., December 2025)
126 ▪ TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection (Jiang-Lin et al., December 2025)
127 ▪ Watermarks for Language Models via Probabilistic Automata (Wang and Shang, December 2025)
128 ▪ Towards Robust Protective Perturbation against DeepFake Face Swapping (Yao et al., December 2025)
129 ▪ Ideal Attribution and Faithful Watermarks for Language Models (Song and Shahabi, December 2025)
130 ▪ Yours or Mine? Overwriting Attacks Against Neural Audio Watermarking (Yao et al., December 2025)
131 ▪ Detection of AI Deepfake and Fraud in Online Payments Using GAN-Based Models (Ke et al., December 2025)
132 ▪ MarkTune: Improving the Quality-Detectability Trade-off in Open-Weight LLM Watermarking (Zhao, Wu, and Block, December 2025)
133 ▪ Watermarks for Embeddings-as-a-Service Large Language Models (Shetty, December 2025)
134 ▪ HeavyWater and SimplexWater: Distortion-Free LLM Watermarks for Low-Entropy Next-Token Predictions (Tsur et al., December 2025)
135 ▪ JPEGs Just Got Snipped: Croppable Signatures Against Deepfake Images (Perazzo et al., December 2025)
136 ▪ HMARK: Radioactive Multi-Bit Semantic-Latent Watermarking for Diffusion Models (Li et al., December 2025)
137 ▪ TAB-DRW: A DFT-based Robust Watermark for Generative Tabular Data (Zhao et al., November 2025)
138 ▪ Frequency Bias Matters: Diving into Robust and Generalized Deep Image Forgery Detection (Liu et al., November 2025)
139 ▪ The Coding Limits of Robust Watermarking for Generative Models (Francati et al., November 2025)
140 ▪ VIDSTAMP: A Temporally-Aware Watermark for Ownership and Integrity in Video Diffusion Models (Teymoorianfard et al., November 2025)
141 ▪ ForensicFlow: A Tri-Modal Adaptive Network for Robust Deepfake Detection (Romani, November 2025)
142 ▪ Sigil: Server-Enforced Watermarking in U-Shaped Split Federated Learning via Gradient Injection (Dai et al., November 2025)
143 ▪ Video Signature: Implicit Watermarking for Video Diffusion Models (Huang et al., November 2025)
144 ▪ LLM-driven Provenance Forensics for Threat Investigation and Detection (Mukherjee and Kantarcioglu, November 2025)
145 ▪ VideoMark: A Distortion-Free Robust Watermarking Framework for Video Diffusion Models (Hu et al., November 2025)
146 ▪ FLClear: Visually Verifiable Multi-Client Watermarking for Federated Learning (Gu et al., November 2025)
147 ▪ DeiTFake: Deepfake Detection Model using DeiT Multi-Stage Training (Kumar et al., November 2025)
148 ▪ Robust Client-Server Watermarking for Split Federated Learning (Tang et al., November 2025)
149 ▪ Adaptive and Robust Watermark for Generative Tabular Data (Ngo et al., November 2025)
150 ▪ Synthetic Voices, Real Threats: Evaluating Large Text-to-Speech Models in Generating Harmful Audio (Chen et al., November 2025)
151 ▪ SEAL: Subspace-Anchored Watermarks for LLM Ownership (Dai et al., November 2025)
152 ▪ On the Information-Theoretic Fragility of Robust Watermarking under Diffusion Editing (Ni et al., November 2025)
153 ▪ Removal Attack and Defense on AI-generated Content Latent-based Watermarking (Lee et al., November 2025)
154 ▪ DeepTracer: Tracing Stolen Model via Deep Coupled Watermarks (Yang et al., November 2025)
155 ▪ Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction Attacks (Xiao et al., November 2025)
156 ▪ Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation (Zhou et al., November 2025)
157 ▪ Diffusion-Based Image Editing: An Unforeseen Adversary to Robust Invisible Watermarks (Fu et al., November 2025)
158 ▪ Shallow Diffuse: Robust and Invisible Watermarking through Low-Dimensional Subspaces in Diffusion Models (Li, Zhang, and Qu, November 2025)
159 ▪ Watermarking Large Language Models in Europe: Interpreting the AI Act in Light of Technology (Souverain, November 2025)
160 ▪ Watermarking Discrete Diffusion Language Models (Bagchi et al., November 2025)
161 ▪ Optimizing Token Choice for Code Watermarking: An RL Approach (Guo et al., November 2025)
162 ▪ From Evidence to Verdict: An Agent-Based Forensic Framework for AI-Generated Image Detection (Liang et al., November 2025)
163 ▪ Robust GNN Watermarking via Implicit Perception of Topological Invariants (Li and Shen, October 2025)
164 ▪ PVMark: Enabling Public Verifiability for LLM Watermarking Schemes (Duan, Xiang, and Zhang, October 2025)
165 ▪ PRO: Enabling Precise and Robust Text Watermark for Open-Source LLMs (Xue et al., October 2025)
166 ▪ Optimal Detection for Language Watermarks with Pseudorandom Collision (Cai et al., October 2025)
167 ▪ DeepfakeBench-MM: A Comprehensive Benchmark for Multimodal Deepfake Detection (Zhao et al., October 2025)
168 ▪ WMCopier: Forging Invisible Image Watermarks on Arbitrary Images (Dong et al., October 2025)
169 ▪ A Reinforcement Learning Framework for Robust and Secure LLM Watermarking (An et al., October 2025)
170 ▪ Can Current Detectors Catch Face-to-Voice Deepfake Attacks? (Nguyen et al., October 2025)
171 ▪ Watermarking Autoregressive Image Generation (Jovanovi\'c et al., October 2025)
172 ▪ Transferable Black-Box One-Shot Forging of Watermarks via Image Preference Models (Sou\v{c}ek et al., October 2025)
173 ▪ Robustness Assessment and Enhancement of Text Watermarking for Google's SynthID (Han et al., October 2025)
174 ▪ Provenance of AI-Generated Images: A Vector Similarity and Blockchain-based Approach (Sharma, Carvalho, and Bhunia, October 2025)
175 ▪ Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption (Liu et al., October 2025)
176 ▪ EditMark: Watermarking Large Language Models based on Model Editing (Li et al., October 2025)
177 ▪ Learning to Watermark: A Selective Watermarking Framework for Large Language Models via Multi-Objective Optimization (Wang et al., October 2025)
178 ▪ MarkDiffusion: An Open-Source Toolkit for Generative Watermarking of Latent Diffusion Models (Pan et al., October 2025)
179 ▪ Every Language Model Has a Forgery-Resistant Signature (Finlayson, Ren, and Swayamdipta, October 2025)
180 ▪ NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models (Goren et al., October 2025)
181 ▪ SimKey: A Semantically Aware Key Module for Watermarking Language Models (Kodama et al., October 2025)
182 ▪ We Can Hide More Bits: The Unused Watermarking Capacity in Theory and in Practice (Petrov et al., October 2025)
183 ▪ SWIFT: Semantic Watermarking for Image Forgery Thwarting (Evennou et al., October 2025)
184 ▪ SynthID-Image: Image watermarking at internet scale (Gowal et al., October 2025)
185 ▪ STOPA: A Database of Systematic VariaTion Of DeePfake Audio for Open-Set Source Tracing and Attribution (Firc et al., October 2025)
186 ▪ LLM Fingerprinting via Semantically Conditioned Watermarks (Gloaguen et al., October 2025)
187 ▪ Benchmarking Fake Voice Detection in the Fake Voice Generation Arms Race (Mao et al., October 2025)
188 ▪ Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security (Naseh et al., October 2025)
189 ▪ Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code (Kim, Park, and Han, October 2025)
190 ▪ LexiMark: Robust Watermarking via Lexical Substitutions to Enhance Membership Verification of an LLM's Textual Training Data (German et al., October 2025)
191 ▪ DMark: Order-Agnostic Watermarking for Diffusion Large Language Models (Wu et al., October 2025)
192 ▪ Secure and Robust Watermarking for AI-generated Images: A Comprehensive Survey (Cao et al., October 2025)
193 ▪ CATMark: A Context-Aware Thresholding Framework for Robust Cross-Task Watermarking in Large Language Models (Zhang et al., October 2025)
194 ▪ ZK-WAGON: Imperceptible Watermark for Image Generation Models using ZK-SNARKs (Ramakrishnan et al., October 2025)
195 ▪ EditTrack: Detecting and Attributing AI-assisted Image Editing (Jiang et al., October 2025)
196 ▪ Fast, Secure, and High-Capacity Image Watermarking with Autoencoded Text Vectors (Evennou, Chappelier, and Kijak, October 2025)
197 ▪ Watermark under Fire: A Robustness Evaluation of LLM Watermarking (Liang et al., October 2025)
198 ▪ Mitigating Watermark Forgery in Generative Models via Randomized Key Selection (Aremu et al., September 2025)
199 ▪ Watermarking Diffusion Language Models (Gloaguen et al., September 2025)
200 ▪ Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size (Tondi, Costanzo, and Barni, September 2025)
201 ▪ PRIVMARK: Private Large Language Models Watermarking with MPC (Fargues et al., September 2025)
202 ▪ Analyzing and Evaluating Unbiased Language Model Watermark (Wu et al., September 2025)
203 ▪ An Ensemble Framework for Unbiased Language Model Watermarking (Wu et al., September 2025)
204 ▪ LLM Watermark Evasion via Bias Inversion (Hwang, Park, and Ok, September 2025)
205 ▪ FakeIDet: Exploring Patches for Privacy-Preserving Fake ID Detection (Mu\~noz-Haro et al., August 2025)
206 ▪ Efficient and Universal Watermarking for LLM-Generated Code Detection (Li et al., August 2025)
207 ▪ Is It Really You? Exploring Biometric Verification Scenarios in Photorealistic Talking-Head Avatar Videos (Pedrouzo-Rodriguez et al., August 2025)
208 ▪ Towards Privacy-preserving Photorealistic Self-avatars in Mixed Reality (Wilson et al., July 2025)
209 ▪ MaXsive: High-Capacity and Robust Training-Free Generative Image Watermarking in Diffusion Models (Mao, Tsai, and Lu, July 2025)
210 ▪ Unmasking Synthetic Realities in Generative AI: A Comprehensive Review of Adversarially Robust Deepfake Detection Systems (Khan et al., July 2025)
211 ▪ WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes (Pujari and Rattani, July 2025)
212 ▪ Robust Data Watermarking in Language Models by Injecting Fictitious Knowledge (Cui et al., July 2025)
213 ▪ Two Views, One Truth: Spectral and Self-Supervised Features Fusion for Robust Speech Deepfake Detection (Kheir et al., July 2025)
214 ▪ VENENA: A Deceptive Visual Encryption Framework for Wireless Semantic Secrecy (Han, Yuan, and Schotten, July 2025)
215 ▪ AEDR: Training-Free AI-Generated Image Attribution via Autoencoder Double-Reconstruction (Wang et al., July 2025)
216 ▪ NWaaS: Nonintrusive Watermarking as a Service for X-to-Image DNN (An et al., July 2025)
217 ▪ Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering (An et al., July 2025)
218 ▪ Towards Trustworthy AI: Secure Deepfake Detection using CNNs and Zero-Knowledge Proofs (Islam, Vo, and Rane, July 2025)
219 ▪ Watermark Anything with Localized Messages (Sander et al., July 2025)
220 ▪ LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech (Liu et al., July 2025)
221 ▪ Chameleon Channels: Measuring YouTube Accounts Repurposed for Deception and Profit (Cuevas, Ribeiro, and Christin, July 2025)
222 ▪ PhishIntentionLLM: Uncovering Phishing Website Intentions through Multi-Agent Retrieval-Augmented Generation (Li et al., July 2025)
223 ▪ IConMark: Robust Interpretable Concept-Based Watermark For AI Images (Sadasivan, Saberi, and Feizi, July 2025)
224 ▪ How does Watermarking Affect Visual Language Models in Document Understanding? (Xu et al., July 2025)
225 ▪ Dynamic Risk Assessments for Offensive Cybersecurity Agents (Wei et al., July 2025)
226 ▪ A Survey on Speech Deepfake Detection (Li, Ahmadiadli, and Zhang, July 2025)
227 ▪ Hashed Watermark as a Filter: Defeating Forging and Overwriting Attacks in Weight-based Neural Network Watermarking (Yao, Song, and Jin, July 2025)
228 ▪ Watermarking Degrades Alignment in Language Models: Analysis and Mitigation (Verma, Phan, and Trivedi, July 2025)
229 ▪ Mitigating Watermark Stealing Attacks in Generative Models via Multi-Key Watermarking (Aremu et al., July 2025)
230 ▪ Phishing Detection in the Gen-AI Era: Quantized LLMs vs Classical Models (Thapa et al., July 2025)
231 ▪ Enhancing LLM Watermark Resilience Against Both Scrubbing and Spoofing Attacks (Shen, Huang, and Wan, July 2025)
232 https://arxiv.org/abs/2411.05091
233 ▪ Removing Watermarks with Partial Regeneration using Semantic Information<br><br> (Tallam et al., May 2025)
234 ▪ LLM-Text Watermarking based on Lagrange Interpolation<br><br> (Janas, Morawiecki, and Pieprzyk, May 2025)
235 ▪ VoiceCloak: A Multi-Dimensional Defense Framework against Unauthorized Diffusion-based Voice Cloning (Hu et al, May 2025)
236 ▪ AGATE: Stealthy Black-box Watermarking for Multimodal Model Copyright Protection (Gao et al, Apr 2025)
237 ▪ Defending LLM Watermarking Against Spoofing Attacks with Contrastive Representation Learning (An et al, Apr 2025)
238 ▪ Mark Your LLM: Detecting the Misuse of Open-Source Large Language Models via Watermarking (Xu et al, Mar 2025)
239 ▪ Robust Watermarking Using Generative Priors Against Image Editing: From Benchmarking to Advances (Lu et al, Mar 2025)
240 ▪ Theoretically Grounded Framework for LLM Watermarking: A Distribution-Adaptive Approach (He et al, Feb 2025)
241 ▪ ID-Cloak: Crafting Identity-Specific Cloaks Against Personalized Text-to-Image Generation (Teng et al, Feb 2025)
242 ▪ Watermarking Language Models with Error Correcting Codes (Chao et al, Feb 2025)
243 ▪ Is The Watermarking Of LLM-Generated Code Robust? (Suresh et al, Feb 2025)
244 ▪ Provably Robust Multi-bit Watermarking for AI-generated Text (Qu et al, Jan 2025)
245 ▪ Audio-Visual Deepfake Detection With Local Temporal Inconsistencies (Astrid, Ghorbel, and Aouada, Jan 2025)
246 ▪ GaussMark: A Practical Approach for Structural Watermarking of Language Models (Block, Sekhari, and Rakhlin, Jan 2025)
247 ▪ Neural Honeytrace: A Robust Plug-and-Play Watermarking Framework against Model Extraction Attacks (Xu et al, Jan 2025)
248 ▪ ModelShield: Adaptive and Robust Watermark against Model Extraction Attack (Pang et al, Jan 2025)
249 ▪ Can Watermarked LLMs be Identified by Users via Crafted Prompts? (Liu et al, Dec 2024)
250 ▪ Watermarking Graph Neural Networks via Explanations for Ownership Protection (Downer et al, Jan 2025)
251 ▪ AI-generated Image Detection: Passive or Watermark? (Guo et al, Jan 2025)
252 ▪ Mesh Watermark Removal Attack and Mitigation: A Novel Perspective of Function Space (Zhu et al, Dec 2024)
253 ▪ PersonaMark: Personalized LLM watermarking for model protection and user attribution (Zhang et al, Dec 2024)
254 ▪ WaterPark: A Robustness Assessment of Language Model Watermarking (Liang et al, Dec 2024)
255 ▪ BlockDoor: Blocking Backdoor Based Watermarks in Deep Neural Networks (Puah et al, Dec 2024)
256 ▪ Towards Effective User Attribution for Latent Diffusion Models via Watermark-Informed Blending (Pan et al, Dec 2024)
257 ▪ GENIE: Watermarking Graph Neural Networks for Link Prediction (Bachina et al, Dec 2024)
258 ▪ The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis (Wu and Chandrasekaran, Dec 2024)
259 ▪ SoK: Watermarking for AI-Generated Content (Zhao et al, Nov 2024)
260 ▪ Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey (Nguyen-Le et al, Nov 2024)
261 ▪ CLUE-MARK: Watermarking Diffusion Models using CLWE (Shehata, Kolluri, and Saxena, Nov 2024)
262 ▪ UnMarker: A Universal Attack on Defensive Image Watermarking (Kasiss and Hengartner, Nov 2024)
263 ▪ SoK: On the Role and Future of AIGC Watermarking in the Era of Gen-AI (Ren et al, Nov 2024)
264 ▪ Conceptwm: A Diffusion Model Watermark for Concept Protection (Lei et al, Nov 2024)
265 ▪ Watermark-based Detection and Attribution of AI-Generated Content (Jiang et al, Nov 2024)
266 ▪ An undetectable watermark for generative image models (Gunn, Zhao, and Song, Nov 2024)
267 ▪ InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance (Xu et al, Nov 2024)
268 ▪ FoldMark: Protecting Protein Generative Models with Watermarking (Zhang et al, Nov 2024)
269 ▪ Invisible Image Watermarks Are Provably Removable Using Generative AI (Zhao et al, Oct 2024)
270 ▪ Embedding Watermarks in Diffusion Process for Model Intellectual Property Protection (Yang, Peng, and Xia, Nov 2024)
271 ▪ Bileve: Securing Text Provenance in Large Language Models Against Spoofing with Bi-level Signature (Zhou et al, Oct 2024)
272 ▪ Watermarking Large Language Models and the Generated Content: Opportunities and Challenges (Zhou and Koushanfar, Oct 2024)
273 ▪ An undetectable watermark for generative image models (Gunn, Zhao, and Song, Oct 2024)
274 ▪ Deepfake detection in videos with multiple faces using geometric-fakeness features (Vyshegorodtsev et al, Oct 2024)
275 ▪ Universally Optimal Watermarking Schemes for LLMs: from Theory to Practice (He et al, Oct 2024)
276 ▪ Discovering Clues of Spoofed LM Watermarks (Gloaguen et al, Oct 2024)
277 ▪ Multi-Designated Detector Watermarking for Language Models (Huang et al, Oct 2024)
278 ▪ Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection (PN et al, Oct 2024)
279 ▪ Signal Watermark on Large Language Models (Zu and Sheng, Oct 2024)
280 ▪ Diffuse or Confuse: A Diffusion Deepfake Speech Dataset (Firc, Malinka, and Hanáček, Oct 2024)
281 ▪ A Watermark for Black-Box Language Models (Bahri et al, Oct 2024)
282 ▪ Optimizing Adaptive Attacks against Content Watermarks for Language Models (Diaa, Aremu, and Lukas, Oct 2024)
283 ▪ Social Media Authentication and Combating Deepfakes using Semi-fragile Invisible Image Watermarking (Nadimpalli and Rattani, Oct 2024)
284 ▪ PITCH: AI-assisted Tagging of Deepfake Audio Calls using Challenge-Response (Mittal et al, Oct 2024)
285 ▪ Shaking the Fake: Detecting Deepfake Videos in Real Time via Active Probes (Xie and Luo, Sep 2024)
286 ▪ XAI-Based Detection of Adversarial Attacks on Deepfake Detectors (Pinhasov et al, Aug 2024)
Deep Fakes, Content Provenance, and Watermarking

>

<

‍

Shallow Fakes

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Shallow Fakes

cybersecurity_tracker - Google Drive

Shallow Fakes

>

<

‍

Misidentification

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Misidentification

cybersecurity_tracker - Google Drive

Misidentification

>

<

‍

Private Information Used in Training

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Private Information Used in Training

cybersecurity_tracker - Google Drive

2 ▪ Efficient and High-Accuracy Private CNN Inference with Helper-Assisted Malicious Security (Wang et al., April 2026)
3 ▪ Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation (Ma and Rajtmajer, April 2026)
4 ▪ A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset (Hong et al., April 2026)
5 ▪ Opal: Private Memory for Personal AI (Kaviani et al., April 2026)
6 ▪ Beyond Laplace and Gaussian: Exploring the Generalized Gaussian Mechanism for Private Machine Learning (Rinberg et al., April 2026)
7 ▪ Combating Data Laundering in LLM Training (Li et al., April 2026)
8 ▪ Quantifying Memorization and Privacy Risks in Genomic Language Models (Nemecek et al., March 2026)
9 ▪ The Hidden Costs of Domain Fine-Tuning: Pii-Bearing Data Degrades Safety and Increases Leakage (Choudhari and Singh, March 2026)
10 ▪ Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs (Liu et al., January 2026)
11 ▪ UnPII: Unlearning Personally Identifiable Information with Quantifiable Exposure Risk (Jeon, Kwon, and Koo, January 2026)
12 ▪ Towards Benchmarking Privacy Vulnerabilities in Selective Forgetting with Large Language Models (Qian et al., December 2025)
13 ▪ DeepShare: Sharing ReLU Across Channels and Layers for Efficient Private Inference (Bornfeld and Avidan, December 2025)
14 ▪ Towards Privacy-Preserving Code Generation: Differentially Private Code Language Models (Catal, Rani, and Gall, December 2025)
15 ▪ Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMs (Joshi and Smith, December 2025)
16 ▪ Multi-PA: A Multi-perspective Benchmark on Privacy Assessment for Large Vision-Language Models (Zhang et al., November 2025)
17 ▪ How do data owners say no? A case study of data consent mechanisms in web-scraped vision-language AI training datasets (Lee et al., November 2025)
18 ▪ Characterizing the Training Dynamics of Private Fine-tuning with Langevin diffusion (Ke et al., November 2025)
19 ▪ AERO: Entropy-Guided Framework for Private LLM Inference (Jha and Reagen, November 2025)
20 ▪ Toward provably private analytics and insights into GenAI use (Cheu et al., October 2025)
21 ▪ SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority Oversampling (Ganev et al., October 2025)
22 ▪ T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation (Li et al., July 2025)
23 ▪ A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census: Full Technical Report (Abowd et al., July 2025)
24 ▪ IDFace: Face Template Protection for Efficient and Secure Identification (Kim et al., July 2025)
25 ▪ Pantomime: Motion Data Anonymization using Foundation Motion Models (Hanisch, Todt, and Strufe, July 2025)
26 ▪ Predicting memorization within Large Language Models fine-tuned for classification (Dentan et al., July 2025)
27 ▪ "Is it always watching? Is it always listening?" Exploring Contextual Privacy and Security Concerns Toward Domestic Social Robots (Bell et al., July 2025)
28 ▪ Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models (Huang et al., July 2025)
29 ▪ RemoteRAG: A Privacy-Preserving LLM Cloud RAG Service (Cheng, Chow and Li, Dec 2024)
30 ▪ Privacy-hardened and hallucination-resistant synthetic data generation with logic-solvers (Burgess et al, Oct 2024)
31 ▪ Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data (Akkus et al, Sep 2024)
32 ▪ Catch Me if You Can: Detecting Unauthorized Data Use in Deep Learning Models (Chen and Pattabiraman, Sep 2024)
33 ▪ Ethical Challenges in Computer Vision: Ensuring Privacy and Mitigating Bias in Publicly Available Datasets (Tahir, Aug 2024)
34 ▪ Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions (Liu and Yang, Aug 2024)
Private Information Used in Training

>

<

‍

Unsecured Credentials

Covers:

  • MITRE ATLAS Credential Access

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : Unsecured Credentials

cybersecurity_tracker - Google Drive

2 ▪ Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study (Chen et al., April 2026)
3 ▪ Keys on Doormats: Exposed API Credentials on the Web (Demir et al., March 2026)
Unsecured Credentials

>

<

‍

AI-Generated/Augmented Exploits

Added this category to cover instances where generative AI systems are used to generate cybersecurity exploits.

Research:

cybersecurity_tracker - Google Drive

cybersecurity_tracker : AI-Generated/Augmented Exploits

cybersecurity_tracker - Google Drive

2 ▪ The Infinite Mutation Engine? Measuring Polymorphism in LLM-Generated Offensive Code (Hortea and Tapiador, May 2026)
3 ▪ Evaluating LLM-Generated Obfuscated XSS Payloads for Machine Learning-Based Detection (Gabbireddy and Saha, April 2026)
4 ▪ RedShell: A Generative AI-Based Approach to Ethical Hacking (Bessa et al., April 2026)
5 ▪ From Theory to Practice: Code Generation Using LLMs for CAPEC and CWE Frameworks (Shahzad et al., April 2026)
6 ▪ Automatic Attack Script Generation: a MDA Approach (Goux and Lammari, March 2026)
7 ▪ Synergistic Directed Execution and LLM-Driven Analysis for Zero-Day AI-Generated Malware Detection (Edwards and Eslamimehr, March 2026)
8 ▪ Execution-State-Aware LLM Reasoning for Automated Proof-of-Vulnerability Generation (Li et al., February 2026)
9 ▪ KryptoPilot: An Open-World Knowledge-Augmented LLM Agent for Automated Cryptographic Exploitation (Liu et al., January 2026)
10 ▪ LLM-based Vulnerable Code Augmentation: Generate or Refactor? (Ouchebara and Dupont, December 2025)
11 ▪ Security Vulnerabilities in AI-Generated Code: A Large-Scale Analysis of Public GitHub Repositories (Schreiber and Tippe, October 2025)
12 ▪ deepSURF: Detecting Memory Safety Vulnerabilities in Rust Through Fuzzing LLM-Augmented Harnesses (Androutsopoulos and Bianchi, October 2025)
13 ▪ PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of Multi-turn Exploits (Bhuiya, Aggarwal, and Purwar, October 2025)
14 ▪ NATLM: Detecting Defects in NFT Smart Contracts Leveraging LLM (Niu, Li, and Li, August 2025)
15 ▪ Autonomous Penetration Testing: Solving Capture-the-Flag Challenges with LLMs (Bakker and Hastings, August 2025)
16 ▪ Fuzzing: Randomness? Reasoning! Efficient Directed Fuzzing via Large Language Models (Feng et al., July 2025)
17 ▪ SAEL: Leveraging Large Language Models with Adaptive Mixture-of-Experts for Smart Contract Vulnerability Detection (Yu et al., July 2025)
18 ▪ Vulnerability Mitigation System (VMS): LLM Agent and Evaluation Framework for Autonomous Penetration Testing (Abdulzada, July 2025)
19 ▪ Intelligent ARP Spoofing Detection using Multi-layered Machine Learning (ML) Techniques for IoT Networks (Ali, Husain, and Hans, July 2025)
20 ▪ Automated Static Vulnerability Detection via a Holistic Neuro-symbolic Approach (Li et al., July 2025)
21 ▪ Are AI-Generated Fixes Secure? Analyzing LLM and Agent Patches on SWE-bench (Sajadi, Damevski, and Chatterjee, July 2025)
22 ▪ Auto-SGCR: Automated Generation of Smart Grid Cyber Range Using IEC 61850 Standard Models (Roomi et al., July 2025)
23 ▪ SynthCTI: LLM-Driven Synthetic CTI Generation to enhance MITRE Technique Mapping (Ruiz-R\'odenas et al., July 2025)
24 ▪ From Text to Actionable Intelligence: Automating STIX Entity and Relationship Extraction (Lekssays, Sencar, and Yu, July 2025)
25 ▪ LibLMFuzz: LLM-Augmented Fuzz Target Generation for Black-box Libraries (Hardgrove and Hastings, July 2025)
26 ▪ Using Modular Arithmetic Optimized Neural Networks To Crack Affine Cryptographic Schemes Efficiently (Stojanovi\'c, Lesar, and Bohak, July 2025)
27 ▪ LLAMA: Multi-Feedback Smart Contract Fuzzing Framework with LLM-Guided Seed Generation (Gai et al., July 2025)
28 ▪ QLPro: Automated Code Vulnerability Discovery via LLM and Static Code Analysis Integration (Hu et al., July 2025)
29 ▪ MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning (Gajjar, Subramaniakuppusamy, and Kachach, July 2025)
30 ▪ A Mixture of Linear Corrections Generates Secure Code (Yu et al., July 2025)
31 ▪ LLMalMorph: On The Feasibility of Generating Variant Malware using Large-Language-Models (Akil et al., July 2025)
32 ▪ PenTest2.0: Towards Autonomous Privilege Escalation Using GenAI (Al-Sinani and Mitchell, July 2025)
33 ▪ AI Agent Smart Contract Exploit Generation (Gervais and Zhou, July 2025)
34 ▪ Metamorphic Malware Evolution: The Potential and Peril of Large Language Models (Madani, Oct 2024)
35 ▪ Exploring RAG-based Vulnerability Augmentation with LLMs (Daneshvar et al, Aug 2024)
AI-Generated/Augmented Exploits

>

<

‍

‍

‍