| 2 |
▪ Toward Securing AI Agents Like Operating Systems (Pirch et al., May 2026) |
| 3 |
▪ One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries (Zloczower et al., May 2026) |
| 4 |
▪ Defenses at Odds: Measuring and Explaining Defense Conflicts in Large Language Models (Meng et al., May 2026) |
| 5 |
▪ MemLineage: Lineage-Guided Enforcement for LLM Agent Memory (Ouyang and Hou, May 2026) |
| 6 |
▪ ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents (Lee and Brumley, May 2026) |
| 7 |
▪ Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution (Zhang et al., May 2026) |
| 8 |
▪ RACC: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., May 2026) |
| 9 |
▪ QuiLL: An LLM-Based Vulnerability Assessment Framework for the Wild (Safdar et al., May 2026) |
| 10 |
▪ Continuous Discovery of Vulnerabilities in LLM Serving Systems with Fuzzing (Zhao et al., May 2026) |
| 11 |
▪ Benchmarking LLM-Based Static Analysis for Secure Smart Contract Development: Reliability, Limitations, and Potential Hybrid Solutions (Susan, Arusoaie, and Lucanu, May 2026) |
| 12 |
▪ What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers (Zhang et al., May 2026) |
| 13 |
▪ Containment Verification: AI Safety Guarantees Independent of Alignment (Moon and Varshney, May 2026) |
| 14 |
▪ Acceptance Cards:A Four-Diagnostic Standard for Safe Fine-Tuning Defense Claims (Konrad, Tanyel, and Ayvaz, May 2026) |
| 15 |
▪ Operationalizing Cybersecurity Governance for Mitigation Planning with Attack-Path Modeling and Reinforcement Learning (Huff et al., May 2026) |
| 16 |
▪ MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring (Jotautait\.e et al., May 2026) |
| 17 |
▪ SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization (Liu et al., May 2026) |
| 18 |
▪ Seed Hijacking of LLM Sampling and Quantum Random Number Defense (You et al., May 2026) |
| 19 |
▪ Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks (Hu, May 2026) |
| 20 |
▪ GLiGuard: Schema-Conditioned Classification for LLM Safeguard (Zaratiana et al., May 2026) |
| 21 |
▪ FIT to Forget: Robust Continual Unlearning for Large Language Models (Xu et al., May 2026) |
| 22 |
▪ SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering (Lin et al., May 2026) |
| 23 |
▪ Information Theoretic Adversarial Training of Large Language Models (Zhang et al., May 2026) |
| 24 |
▪ Autonomous Adversary: Red-Teaming in the age of LLM (Mamun et al., May 2026) |
| 25 |
▪ Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks (Lu et al., May 2026) |
| 26 |
▪ SafeHarbor: Hierarchical Memory-Augmented Guardrail for LLM Agent Safety (Liu et al., May 2026) |
| 27 |
▪ GLiNER Guard: Unified Encoder Family for Production LLM Safety and Privacy (Minko, Sadiekh, and Kokuykin, May 2026) |
| 28 |
▪ A Systematic Survey of Security Threats and Defenses in LLM-Based AI Agents: A Layered Attack Surface Framework (Chu, May 2026) |
| 29 |
▪ On the (In-)Security of the Shuffling Defense in the Transformer Secure Inference (Li et al., May 2026) |
| 30 |
▪ Large Language Model assisted Hybrid Fuzzing (Meng, Duck, and Roychoudhury, May 2026) |
| 31 |
▪ MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents (Steinberg and Gal, May 2026) |
| 32 |
▪ ARGUS: Defending LLM Agents Against Context-Aware Prompt Injection (Weng et al., May 2026) |
| 33 |
▪ MAGE: Safeguarding LLM Agents against Long-Horizon Threats via Shadow Memory (Wang et al., May 2026) |
| 34 |
▪ Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses (Li et al., May 2026) |
| 35 |
▪ RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs (Asif and Amiri, May 2026) |
| 36 |
▪ Toward a Principled Framework for Agent Safety Measurement (Lin et al., May 2026) |
| 37 |
▪ When Embedding-Based Defenses Fail: Rethinking Safety in LLM-Based Multi-Agent Systems (Zhang, Zheng, and Chen, May 2026) |
| 38 |
▪ Sentra-Guard: A Real-Time Multilingual Defense Against Adversarial LLM Prompts (Hasan et al., May 2026) |
| 39 |
▪ ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models (Zhao et al., May 2026) |
| 40 |
▪ GAVEL: Towards Rule-Based Safety Through Activation Monitoring (Rozenfeld et al., May 2026) |
| 41 |
▪ Machine Unlearning for Class Removal through SISA-based Deep Neural Network Architectures (Mahi et al., May 2026) |
| 42 |
▪ AdaBFL: Multi-Layer Defensive Adaptive Aggregation for Bzantine-Robust Federated Learning (Tang, Liu, and Huang, May 2026) |
| 43 |
▪ Adaptive and AI-Augmented Security Testing: A Systematic Survey of Program Analysis, Feedback-Driven Testing, and Hybrid Learning-Based Approaches (Wienczkowski, May 2026) |
| 44 |
▪ Security Attack and Defense Strategies for Autonomous Agent Frameworks: A Layered Review with OpenClaw as a Case Study (Xu and Chen, May 2026) |
| 45 |
▪ Toward Autonomous SOC Operations: End-to-End LLM Framework for Threat Detection, Query Generation, and Resolution in Security Operations (Saju and Azim, May 2026) |
| 46 |
▪ An Empirical Security Evaluation of LLM-Generated Cryptographic Rust Code (Elsayed, Fulton, and Yang, May 2026) |
| 47 |
▪ Enforcing Benign Trajectories: A Behavioral Firewall for Structured-Workflow AI Agents (Dang, April 2026) |
| 48 |
▪ Large Language Models as Explainable Cyberattack Detectors for Energy Industrial Control Systems (Kong et al., April 2026) |
| 49 |
▪ A Comparative Evaluation of AI Agent Security Guardrails (Li et al., April 2026) |
| 50 |
▪ Mitigating Error Amplification in Fast Adversarial Training (Zhao et al., April 2026) |
| 51 |
▪ AgentVisor: Defending LLM Agents Against Prompt Injection via Semantic Virtualization (Ying et al., April 2026) |
| 52 |
▪ Poster: ClawdGo: Endogenous Security Awareness Training for Autonomous AI Agents (Li et al., April 2026) |
| 53 |
▪ Evaluation of Prompt Injection Defenses in Large Language Models (Deep et al., April 2026) |
| 54 |
▪ UNSEEN: A Cross-Stack LLM Unlearning Defense against AR-LLM Social Engineering Attacks (Yu et al., April 2026) |
| 55 |
▪ Secure eFPGA-Enabled Edge LLM Inference: Architectural and Hardware Countermeasures (Das et al., April 2026) |
| 56 |
▪ AutoRISE: Agent-Driven Strategy Evolution for Red-Teaming Large Language Models (Gautam, Bahramali, and Atluri, April 2026) |
| 57 |
▪ SecureVibeBench: Benchmarking Secure Vibe Coding of AI Agents via Reconstructing Vulnerability-Introducing Scenarios (Chen et al., April 2026) |
| 58 |
▪ Secure LLM Fine-Tuning via Safety-Aware Probing (Wu et al., April 2026) |
| 59 |
▪ Adaptive Instruction Composition for Automated LLM Red-Teaming (Zymet et al., April 2026) |
| 60 |
▪ Adaptive Defense Orchestration for RAG: A Sentinel-Strategist Architecture against Multi-Vector Attacks (Pallerla et al., April 2026) |
| 61 |
▪ Verification of Machine Unlearning is Fragile (Zhang et al., April 2026) |
| 62 |
▪ AVISE: Framework for Evaluating the Security of AI Systems (Lempinen, Kemppainen, and Raesalmi, April 2026) |
| 63 |
▪ Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing (Talluri, April 2026) |
| 64 |
▪ Towards Certified Malware Detection: Provable Guarantees Against Evasion Attacks (Giri et al., April 2026) |
| 65 |
▪ Taint-Style Vulnerability Detection and Confirmation for Node.js Packages Using LLM Agent Reasoning (Ni, Christodorescu, and Jia, April 2026) |
| 66 |
▪ Benchmarking Misuse Mitigation Against Covert Adversaries (Brown et al., April 2026) |
| 67 |
▪ ARES: Adaptive Red-Teaming and End-to-End Repair of Policy-Reward System (Liang et al., April 2026) |
| 68 |
▪ Malicious ML Model Detection by Learning Dynamic Behaviors (Nambiar, Pradhan, and Soremekun, April 2026) |
| 69 |
▪ SAGE: Signal-Amplified Guided Embeddings for LLM-based Vulnerability Detection (Shan et al., April 2026) |
| 70 |
▪ Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents (Xu, April 2026) |
| 71 |
▪ ReGA: Model-Based Safeguard for LLMs via Representation-Guided Abstraction (Wei, Wu, and Sun, April 2026) |
| 72 |
▪ SDLLMFuzz: Dynamic-static LLM-assisted greybox fuzzing for structured input programs (Zou et al., April 2026) |
| 73 |
▪ Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives (Son, April 2026) |
| 74 |
▪ enclawed: A Configurable, Sector-Neutral Hardening Framework for Single-User AI Assistant Gateways (Metere, April 2026) |
| 75 |
▪ SafeLM: Unified Privacy-Aware Optimization for Trustworthy Federated Large Language Models (Mohammad and Bayaz{\i}t, April 2026) |
| 76 |
▪ TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts (Chu, Wang, and Huang, April 2026) |
| 77 |
▪ Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility (Hong et al., April 2026) |
| 78 |
▪ Privacy-Preserving LLMs Routing (Wu et al., April 2026) |
| 79 |
▪ SafeHarness: Lifecycle-Integrated Security Architecture for LLM-based Agent Deployment (Lin et al., April 2026) |
| 80 |
▪ Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory (Fu and Wang, April 2026) |
| 81 |
▪ SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents (Begimher et al., April 2026) |
| 82 |
▪ Neuro-symbolic Static Analysis with LLM-generated Vulnerability Patterns (Li et al., April 2026) |
| 83 |
▪ Towards Automated Pentesting with Large Language Models (Bessa et al., April 2026) |
| 84 |
▪ QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits (Azimi et al., April 2026) |
| 85 |
▪ PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification (Gong and Deng, April 2026) |
| 86 |
▪ Towards Effective Offensive Security LLM Agents: Hyperparameter Tuning, LLM as a Judge, and a Lightweight CTF Benchmark (Shao et al., April 2026) |
| 87 |
▪ Robustness via Referencing: Defending against Prompt Injection Attacks by Referencing the Executed Instruction (Chen et al., April 2026) |
| 88 |
▪ Vulnerability Detection with Interprocedural Context in Multiple Languages: Assessing Effectiveness and Cost of Modern LLMs (Lira et al., April 2026) |
| 89 |
▪ Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions (Xu et al., April 2026) |
| 90 |
▪ Towards Identification and Intervention of Safety-Critical Parameters in Large Language Models (Qi et al., April 2026) |
| 91 |
▪ MCP-DPT: A Defense-Placement Taxonomy and Coverage Analysis for Model Context Protocol Security (Rostamzadeh et al., April 2026) |
| 92 |
▪ Can Drift-Adaptive Malware Detectors Be Made Robust? Attacks and Defenses Under White-Box and Black-Box Threats (Li, Akil, and Bertino, April 2026) |
| 93 |
▪ The Defense Trilemma: Why Prompt Injection Defense Wrappers Fail? (Bhatt et al., April 2026) |
| 94 |
▪ Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization (Maljkovic et al., April 2026) |
| 95 |
▪ SALLIE: Safeguarding Against Latent Language & Image Exploits (Azov, Rivlin, and Shtar, April 2026) |
| 96 |
▪ CritBench: A Framework for Evaluating Cybersecurity Capabilities of Large Language Models in IEC 61850 Digital Substation Environments (Keppler, Gst\"ur, and Hagenmeyer, April 2026) |
| 97 |
▪ A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms (Acharya and Gupta, April 2026) |
| 98 |
▪ Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts (Uenal, April 2026) |
| 99 |
▪ BodhiPromptShield: Pre-Inference Prompt Mediation for Suppressing Privacy Propagation in LLM/VLM Agents (Ma, Wu, and Yan, April 2026) |
| 100 |
▪ Broken by Default: A Formal Verification Study of Security Vulnerabilities in AI-Generated Code (Blain and Noiseux, April 2026) |
| 101 |
▪ SoSBench: Benchmarking Safety Alignment on Six Scientific Domains (Jiang et al., April 2026) |
| 102 |
▪ Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework (Zhou et al., April 2026) |
| 103 |
▪ Explainable Autonomous Cyber Defense using Adversarial Multi-Agent Reinforcement Learning (Zhang, Goel, and Ahmad, April 2026) |
| 104 |
▪ CoopGuard: Stateful Cooperative Agents Safeguarding LLMs Against Evolving Multi-Round Attacks (Li et al., April 2026) |
| 105 |
▪ Automating Cloud Security and Forensics Through a Secure-by-Design Generative AI Framework (Alharthi and Garcia, April 2026) |
| 106 |
▪ SecPI: Secure Code Generation with Reasoning Models via Security Reasoning Internalization (Wang et al., April 2026) |
| 107 |
▪ Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AI (Cheruvu et al., April 2026) |
| 108 |
▪ Seclens: Role-specific Evaluation of LLM's for security vulnerablity detection (Halder et al., April 2026) |
| 109 |
▪ Assertain: Automated Security Assertion Generation Using Large Language Models (Tarek et al., April 2026) |
| 110 |
▪ Certifiably Robust RAG against Retrieval Corruption (Xiang et al., April 2026) |
| 111 |
▪ Secure Forgetting: A Framework for Privacy-Driven Unlearning in Large Language Model (LLM)-Based Agents (Ye et al., April 2026) |
| 112 |
▪ VibeGuard: A Security Gate Framework for AI-Generated Code (Xie, April 2026) |
| 113 |
▪ Automated Framework to Evaluate and Harden LLM System Instructions against Encoding Attacks (Sahu, Samanta, and Soosahabi, April 2026) |
| 114 |
▪ Quantum-Safe Code Auditing: LLM-Assisted Static Analysis and Quantum-Aware Risk Scoring for Post-Quantum Cryptography Migration (Shaw, April 2026) |
| 115 |
▪ SecureVibeBench: Evaluating Secure Coding Capabilities of Code Agents with Realistic Vulnerability Scenarios (Chen et al., April 2026) |
| 116 |
▪ Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks (Xiang et al., April 2026) |
| 117 |
▪ CivicShield: A Cross-Domain Defense-in-Depth Framework for Securing Government-Facing AI Chatbots Against Multi-Turn Adversarial Attacks (Patil, April 2026) |
| 118 |
▪ Design Principles for the Construction of a Benchmark Evaluating Security Operation Capabilities of Multi-agent AI Systems (Cai et al., April 2026) |
| 119 |
▪ Attesting LLM Pipelines: Enforcing Verifiable Training and Release Claims (Tan, Singer, and Anagnostopoulos, April 2026) |
| 120 |
▪ GUARD-SLM: Token Activation-Based Defense Against Jailbreak Attacks for Small Language Models (Mia et al., April 2026) |
| 121 |
▪ SkillTester: Benchmarking Utility and Security of Agent Skills (Wang, Wang, and Xu, April 2026) |
| 122 |
▪ VulInstruct: Teaching LLMs Root-Cause Reasoning for Vulnerability Detection via Security Specifications (Zhu et al., March 2026) |
| 123 |
▪ Safeguarding LLMs Against Misuse and AI-Driven Malware Using Steganographic Canaries (Raz et al., March 2026) |
| 124 |
▪ A Systematic Taxonomy of Security Vulnerabilities in the OpenClaw AI Agent Framework (Suwansathit, Zhang, and Gu, March 2026) |
| 125 |
▪ Protecting User Prompts Via Character-Level Differential Privacy (Arachchige et al., March 2026) |
| 126 |
▪ Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models (Wang et al., March 2026) |
| 127 |
▪ Policy-Guided Threat Hunting: An LLM enabled Framework with Splunk SOC Triage (Sahay et al., March 2026) |
| 128 |
▪ Agent Audit: A Security Analysis System for LLM Agent Applications (Zhang, Nian, and Zhao, March 2026) |
| 129 |
▪ Does Teaming-Up LLMs Improve Secure Code Generation? A Comprehensive Evaluation with Multi-LLMSecCodeEval (Sabir et al., March 2026) |
| 130 |
▪ BioShield: A Context-Aware Firewall for Securing Bio-LLMs (Das et al., March 2026) |
| 131 |
▪ T-MAP: Red-Teaming LLM Agents with Trajectory-aware Evolutionary Search (Lee et al., March 2026) |
| 132 |
▪ Agentproof: Static Verification of Agent Workflow Graphs (Xavier et al., March 2026) |
| 133 |
▪ SecureBreak -- A dataset towards safe and secure models (Arazzi, Kembu, and Nocera, March 2026) |
| 134 |
▪ Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks (Mu et al., March 2026) |
| 135 |
▪ DeepXplain: XAI-Guided Autonomous Defense Against Multi-Stage APT Campaigns (Phan and Bauschert, March 2026) |
| 136 |
▪ LISAA: A Framework for Large Language Model Information Security Awareness Assessment (Cohen et al., March 2026) |
| 137 |
▪ A Framework for Formalizing LLM Agent Security (Siu et al., March 2026) |
| 138 |
▪ The Autonomy Tax: Defense Training Breaks LLM Agents (Li and Zhao, March 2026) |
| 139 |
▪ Network and Device Level Cyber Deception for Contested Environments Using RL and LLMs (Sahu, Paul, and Macwan, March 2026) |
| 140 |
▪ Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety (Chen et al., March 2026) |
| 141 |
▪ Security awareness in LLM agents: the NDAI zone case (Bottazzi and Park, March 2026) |
| 142 |
▪ Toward Reliable, Safe, and Secure LLMs for Scientific Applications (Chaturvedi, Bergerson, and Mallick, March 2026) |
| 143 |
▪ Guardrails as Infrastructure: Policy-First Control for Tool-Orchestrated Workflows (Sigdel and Baral, March 2026) |
| 144 |
▪ Differential Privacy in Generative AI Agents: Analysis and Optimal Tradeoffs (Yang and Zhu, March 2026) |
| 145 |
▪ Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework (Onitiju and Vakilinia, March 2026) |
| 146 |
▪ DeepStage: Learning Autonomous Defense Policies Against Multi-Stage APT Campaigns (Phan, Nguyen, and Bauschert, March 2026) |
| 147 |
▪ Evolving Contextual Safety in Multi-Modal Large Language Models via Inference-Time Self-Reflective Memory (Zhang et al., March 2026) |
| 148 |
▪ Rotated Robustness: A Training-Free Defense against Bit-Flip Attacks on Large Language Models (Liu and Chen, March 2026) |
| 149 |
▪ TrinityGuard: A Unified Framework for Safeguarding Multi-Agent Systems (Wang et al., March 2026) |
| 150 |
▪ SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration (Pan et al., March 2026) |
| 151 |
▪ Architecture-Agnostic Feature Synergy for Universal Defense Against Heterogeneous Generative Threats (Zhang et al., March 2026) |
| 152 |
▪ Governing Dynamic Capabilities: Cryptographic Binding and Reproducibility Verification for AI Agent Tool Use (Zhou, March 2026) |
| 153 |
▪ SecDTD: Dynamic Token Drop for Secure Transformers Inference (Cai et al., March 2026) |
| 154 |
▪ CTI-REALM: Benchmark to Evaluate Agent Performance on Security Detection Rule Generation Capabilities (Chakraborty et al., March 2026) |
| 155 |
▪ Why Neural Structural Obfuscation Can't Kill White-Box Watermarks for Good! (Jiang et al., March 2026) |
| 156 |
▪ Uncovering Security Threats and Architecting Defenses in Autonomous Agents: A Case Study of OpenClaw (Ying et al., March 2026) |
| 157 |
▪ AEGIS: No Tool Call Left Unchecked -- A Pre-Execution Firewall and Audit Layer for AI Agents (Yuan, Su, and Zhao, March 2026) |
| 158 |
▪ Security Considerations for Artificial Intelligence Agents (Li et al., March 2026) |
| 159 |
▪ OpenClaw PRISM: A Zero-Fork, Defense-in-Depth Runtime Security Layer for Tool-Augmented LLM Agents (Li, March 2026) |
| 160 |
▪ Security-by-Design for LLM-Based Code Generation: Leveraging Internal Representations for Concept-Driven Steering Mechanisms (Wendlinger et al., March 2026) |
| 161 |
▪ Differential Privacy in Machine Learning: A Survey from Symbolic AI to LLMs (Aguilera-Mart\'inez and Berzal, March 2026) |
| 162 |
▪ TOSSS: a CVE-based Software Security Benchmark for Large Language Models (Damie et al., March 2026) |
| 163 |
▪ Enhancing Network Intrusion Detection Systems: A Multi-Layer Ensemble Approach to Mitigate Adversarial Attacks (Soltani et al., March 2026) |
| 164 |
▪ AutoControl Arena: Synthesizing Executable Test Environments for Frontier AI Risk Evaluation (Li et al., March 2026) |
| 165 |
▪ SCAFFOLD-CEGIS: Preventing Latent Security Degradation in LLM-Driven Iterative Code Refinement (Chen et al., March 2026) |
| 166 |
▪ DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation (Jiang, March 2026) |
| 167 |
▪ Trusting What You Cannot See: Auditable Fine-Tuning and Inference for Proprietary AI (Jin et al., March 2026) |
| 168 |
▪ Where Do LLM-based Systems Break? A System-Level Security Framework for Risk Assessment and Treatment (Nagaraja and Bahsi, March 2026) |
| 169 |
▪ ESAA-Security: An Event-Sourced, Verifiable Architecture for Agent-Assisted Security Audits of AI-Generated Code (Filho, March 2026) |
| 170 |
▪ Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Models (Wu et al., March 2026) |
| 171 |
▪ Secure human oversight of AI: Threat modeling in a socio-technical context (Ditz et al., March 2026) |
| 172 |
▪ EVMbench: Evaluating AI Agents on Smart Contract Security (Wang et al., March 2026) |
| 173 |
▪ Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks (Chu et al., March 2026) |
| 174 |
▪ Goal-Driven Risk Assessment for LLM-Powered Systems: A Healthcare Case Study (Nagaraja and Bahsi, March 2026) |
| 175 |
▪ WARP: Weight Teleportation for Attack-Resilient Unlearning Protocols (Maheri et al., March 2026) |
| 176 |
▪ LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities (Nakka et al., March 2026) |
| 177 |
▪ Contextualized Privacy Defense for LLM Agents (Wen et al., March 2026) |
| 178 |
▪ SOSecure: Safer Code Generation with RAG and StackOverflow Discussions (Mukherjee and Hellendoorn, March 2026) |
| 179 |
▪ ReVision : A Post-Hoc, Vision-Based Technique for Replacing Unacceptable Concepts in Image Generation Pipeline (Singh et al., March 2026) |
| 180 |
▪ Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision (Mukherjee and Hellendoorn, March 2026) |
| 181 |
▪ Token-level Data Selection for Safe LLM Fine-tuning (Li et al., March 2026) |
| 182 |
▪ MPU: Towards Secure and Privacy-Preserving Knowledge Unlearning for Large Language Models (Wang et al., March 2026) |
| 183 |
▪ Learning to Generate Secure Code via Token-Level Rewards (Quan et al., March 2026) |
| 184 |
▪ Automated Vulnerability Detection in Source Code Using Deep Representation Learning (Seas et al., February 2026) |
| 185 |
▪ A Lightweight Defense Mechanism against Next Generation of Phishing Emails using Distilled Attention-Augmented BiLSTM (Eskandarian et al., February 2026) |
| 186 |
▪ Secure Semantic Communications via AI Defenses: Fundamentals, Solutions, and Future Directions (Zhang et al., February 2026) |
| 187 |
▪ Resilient Federated Chain: Transforming Blockchain Consensus into an Active Defense Layer for Federated Learning (Garc\'ia-M\'arquez et al., February 2026) |
| 188 |
▪ A Systematic Review of Algorithmic Red Teaming Methodologies for Assurance and Security of AI Applications (Srivastava, Janardhan, and Jauhari, February 2026) |
| 189 |
▪ AttestLLM: Efficient Attestation Framework for Billion-scale On-device LLMs (Zhang et al., February 2026) |
| 190 |
▪ Dynamic Probabilistic Noise Injection for Membership Inference Defense (Forough and Haddadi, February 2026) |
| 191 |
▪ LLM-enabled Applications Require System-Level Threat Monitoring (Zhang et al., February 2026) |
| 192 |
▪ CIBER: A Comprehensive Benchmark for Security Evaluation of Code Interpreter Agents (Ba, Li, and Li, February 2026) |
| 193 |
▪ KUDA: Knowledge Unlearning by Deviating Representation for Large Language Models (Fang et al., February 2026) |
| 194 |
▪ MANATEE: Inference-Time Lightweight Diffusion Based Safety Defense for LLMs (Kan et al., February 2026) |
| 195 |
▪ DAVE: A Policy-Enforcing LLM Spokesperson for Secure Multi-Document Data Sharing (Brinkhege and Menon, February 2026) |
| 196 |
▪ NeST: Neuron Selective Tuning for LLM Safety (Behrouzi et al., February 2026) |
| 197 |
▪ NESSiE: The Necessary Safety Benchmark -- Identifying Errors that should not Exist (Bertram and Geiping, February 2026) |
| 198 |
▪ Secure Coding with AI -- From Detection to Repair (Belozerov, Barclay, and Sami, February 2026) |
| 199 |
▪ PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models (Yuan et al., February 2026) |
| 200 |
▪ Large Language Models for Secure Code Assessment: A Multi-Language Empirical Study (Dozono, Gasiba, and Stocco, February 2026) |
| 201 |
▪ ForesightSafety Bench: A Frontier Risk Evaluation and Governance Framework towards Safe AI (Tong et al., February 2026) |
| 202 |
▪ MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agents (Zhou et al., February 2026) |
| 203 |
▪ From SFT to RL: Demystifying the Post-Training Pipeline for LLM-based Vulnerability Detection (Li, Yu, and Wang, February 2026) |
| 204 |
▪ Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents (Li et al., February 2026) |
| 205 |
▪ Neighborhood Blending: A Lightweight Inference-Time Defense Against Membership Inference Attacks (Zafar et al., February 2026) |
| 206 |
▪ TensorCommitments: A Lightweight Verifiable Inference for Language Models (Baser et al., February 2026) |
| 207 |
▪ Sparse Autoencoders are Capable LLM Jailbreak Mitigators (Assogba et al., February 2026) |
| 208 |
▪ Poly-Guard: Massive Multi-Domain Safety Policy-Grounded Guardrail Dataset (Kang et al., February 2026) |
| 209 |
▪ DeepSight: An All-in-One LM Safety Toolkit (Zhang et al., February 2026) |
| 210 |
▪ The PBSAI Governance Ecosystem: A Multi-Agent AI Reference Architecture for Securing Enterprise AI Estates (Willis, February 2026) |
| 211 |
▪ SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models (Thornton, February 2026) |
| 212 |
▪ GoodVibe: Security-by-Vibe for LLM-Based Code Generation (Thang et al., February 2026) |
| 213 |
▪ MAPS: A Multilingual Benchmark for Agent Performance and Security (Hofman et al., February 2026) |
| 214 |
▪ Selective KV-Cache Sharing to Mitigate Timing Side-Channels in LLM Inference (Chu et al., February 2026) |
| 215 |
▪ Stop Testing Attacks, Start Diagnosing Defenses: The Four-Checkpoint Framework Reveals Where LLM Safety Breaks (Dhabhi and Thimmaraju, February 2026) |
| 216 |
▪ Dashed Line Defense: Plug-And-Play Defense Against Adaptive Score-Based Query Attacks (Fu, Guo, and Luo, February 2026) |
| 217 |
▪ MemPot: Defending Against Memory Extraction Attack with Optimized Honeypots (Wang et al., February 2026) |
| 218 |
▪ Secure Code Generation via Online Reinforcement Learning with Vulnerability Reward Model (Wu et al., February 2026) |
| 219 |
▪ Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention (Singh et al., February 2026) |
| 220 |
▪ TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering (Hossain et al., February 2026) |
| 221 |
▪ Evaluating and Enhancing the Vulnerability Reasoning Capabilities of Large Language Models (Lu et al., February 2026) |
| 222 |
▪ How Catastrophic is Your LLM? Certifying Risk in Conversation (Wang et al., February 2026) |
| 223 |
▪ Persistent Human Feedback, LLMs, and Static Analyzers for Secure Code Generation and Vulnerability Detection (Firouzi and Ghafari, February 2026) |
| 224 |
▪ Spider-Sense: Intrinsic Risk Sensing for Efficient Agent Defense with Hierarchical Adaptive Screening (Yu et al., February 2026) |
| 225 |
▪ Comparative Insights on Adversarial Machine Learning from Industry and Academia: A User-Study Approach (Kakkad et al., February 2026) |
| 226 |
▪ Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models (Ding et al., February 2026) |
| 227 |
▪ Constitutional Spec-Driven Development: Enforcing Security by Construction in AI-Assisted Code Generation (Marri, February 2026) |
| 228 |
▪ GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, and Video (Zhu et al., February 2026) |
| 229 |
▪ TinyGuard:A lightweight Byzantine Defense for Resource-Constrained Federated Learning via Statistical Update Fingerprints (Mahdavi et al., February 2026) |
| 230 |
▪ To Defend Against Cyber Attacks, We Must Teach AI Agents to Hack (Zhuo et al., February 2026) |
| 231 |
▪ MAS-Shield: A Defense Framework for Secure and Efficient LLM MAS (Wang et al., February 2026) |
| 232 |
▪ When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models (Ou et al., February 2026) |
| 233 |
▪ RACA: Representation-Aware Coverage Criteria for LLM Safety Testing (Wei et al., February 2026) |
| 234 |
▪ Co-RedTeam: Orchestrated Security Discovery and Exploitation with LLM Agents (He et al., February 2026) |
| 235 |
▪ SMCP: Secure Model Context Protocol (Hou et al., February 2026) |
| 236 |
▪ Tri-LLM Cooperative Federated Zero-Shot Intrusion Detection with Semantic Disagreement and Trust-Aware Aggregation (Jamshidi et al., February 2026) |
| 237 |
▪ Detecting Instruction Fine-tuning Attacks using Influence Function (Li, February 2026) |
| 238 |
▪ No More, No Less: Least-Privilege Language Models (Rauba et al., February 2026) |
| 239 |
▪ RealSec-bench: A Benchmark for Evaluating Secure Code Generation in Real-World Repositories (Wang et al., February 2026) |
| 240 |
▪ FraudShield: Knowledge Graph Empowered Defense for LLMs against Fraud Attacks (Xu et al., February 2026) |
| 241 |
▪ A Systematic Literature Review on LLM Defenses Against Prompt Injection and Jailbreaking: Expanding NIST Taxonomy (Correia et al., February 2026) |
| 242 |
▪ ShellForge: Adversarial Co-Evolution of Webshell Generation and Multi-View Detection for Robust Webshell Defense (Ding, February 2026) |
| 243 |
▪ SafeSearch: Automated Red-Teaming of LLM-Based Search Agents (Dong et al., January 2026) |
| 244 |
▪ FIT: Defying Catastrophic Forgetting in Continual LLM Unlearning (Xu et al., January 2026) |
| 245 |
▪ RerouteGuard: Understanding and Mitigating Adversarial Risks for LLM Routing (Zhang et al., January 2026) |
| 246 |
▪ Securing AI Agents in Cyber-Physical Systems: A Survey of Environmental Interactions, Deepfake Threats, and Defenses (Hatami et al., January 2026) |
| 247 |
▪ Benchmarking LLAMA Model Security Against OWASP Top 10 For LLM Applications (Shahin and Alsmadi, January 2026) |
| 248 |
▪ Towards Secure MLOps: Surveying Attacks, Mitigation Strategies, and Research Challenges (Patel et al., January 2026) |
| 249 |
▪ GAVEL: Towards rule-based safety through activation monitoring (Rozenfeld et al., January 2026) |
| 250 |
▪ RvB: Automating AI System Hardening via Iterative Red-Blue Games (Huang et al., January 2026) |
| 251 |
▪ LLM-Assisted Authentication and Fraud Detection (Chan and Chan, January 2026) |
| 252 |
▪ SHIELD: An Auto-Healing Agentic Defense Framework for LLM Resource Exhaustion Attacks (Sivaroopan et al., January 2026) |
| 253 |
▪ Proactive Hardening of LLM Defenses with HASTE (Chen et al., January 2026) |
| 254 |
▪ Attributing and Exploiting Safety Vectors through Global Optimization in Large Language Models (Chu et al., January 2026) |
| 255 |
▪ $\alpha^3$-SecBench: A Large-Scale Evaluation Suite of Security, Resilience, and Trust for LLM-based UAV Agents over 6G Networks (Ferrag, Lakas, and Debbah, January 2026) |
| 256 |
▪ Mitigating the OWASP Top 10 For Large Language Models Applications using Intelligent Agents (Fasha et al., January 2026) |
| 257 |
▪ PatchIsland: Orchestration of LLM Agents for Continuous Vulnerability Repair (Kim et al., January 2026) |
| 258 |
▪ ProveRAG: Provenance-Driven Vulnerability Analysis with Automated Retrieval-Augmented LLMs (Fayyazi et al., January 2026) |
| 259 |
▪ Evaluating the Defense Potential of Machine Unlearning against Membership Inference Attacks (Tsiolakis et al., January 2026) |
| 260 |
▪ PAL*M: Property Attestation for Large Generative Models (Chantasantitam et al., January 2026) |
| 261 |
▪ Introducing the Generative Application Firewall (GAF) (Farreny et al., January 2026) |
| 262 |
▪ The Good, the Bad and the Ugly: Meta-Analysis of Watermarks, Transferable Attacks and Adversarial Defenses (G{\l}uch et al., January 2026) |
| 263 |
▪ Guardrails for trust, safety, and ethical development and deployment of Large Language Models (LLM) (Biswas and Talukdar, January 2026) |
| 264 |
▪ Beyond the Safety Tax: Mitigating Unsafe Text-to-Image Generation via External Safety Rectification (Meng et al., January 2026) |
| 265 |
▪ HardSecBench: Benchmarking the Security Awareness of LLMs for Hardware Code Generation (Chen et al., January 2026) |
| 266 |
▪ Serverless AI Security: Attack Surface Analysis and Runtime Protection Mechanisms for FaaS-Based Machine Learning (Pathade et al., January 2026) |
| 267 |
▪ SecMLOps: A Comprehensive Framework for Integrating Security Throughout the MLOps Lifecycle (Zhang et al., January 2026) |
| 268 |
▪ Be Your Own Red Teamer: Safety Alignment via Self-Play and Reflective Experience Replay (Wang et al., January 2026) |
| 269 |
▪ Blue Teaming Function-Calling Agents (Dolcetti, Zizzo, and Maffeis, January 2026) |
| 270 |
▪ Evaluating Implicit Regulatory Compliance in LLM Tool Invocation via Logic-Guided Synthesis (Song et al., January 2026) |
| 271 |
▪ AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation (Kumar, January 2026) |
| 272 |
▪ How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape (Kelley et al., January 2026) |
| 273 |
▪ BlindU: Blind Machine Unlearning without Revealing Erasing Data (Wang et al., January 2026) |
| 274 |
▪ Defenses Against Prompt Attacks Learn Surface Heuristics (Li et al., January 2026) |
| 275 |
▪ Safe-FedLLM: Delving into the Safety of Federated Large Language Models (Tao et al., January 2026) |
| 276 |
▪ United We Defend: Collaborative Membership Inference Defenses in Federated Learning (Bai et al., January 2026) |
| 277 |
▪ Cyber Threat Detection and Vulnerability Assessment System using Generative AI and Large Language Model (M et al., January 2026) |
| 278 |
▪ AutoVulnPHP: LLM-Powered Two-Stage PHP Vulnerability Detection and Automated Localization (Wang et al., January 2026) |
| 279 |
▪ VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit (Lin et al., January 2026) |
| 280 |
▪ HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense (Li et al., January 2026) |
| 281 |
▪ AI-Driven Cybersecurity Threats: A Survey of Emerging Risks and Defensive Strategies (Erukude, Marella, and Veluru, January 2026) |
| 282 |
▪ Autonomous Threat Detection and Response in Cloud Security: A Comprehensive Survey of AI-Driven Strategies (Sarraf and Pal, January 2026) |
| 283 |
▪ Automated Post-Incident Policy Gap Analysis via Threat-Informed Evidence Mapping using Large Language Models (Oh et al., January 2026) |
| 284 |
▪ SLIM: Stealthy Low-Coverage Black-Box Watermarking via Latent-Space Confusion Zones (Wu and Cao, January 2026) |
| 285 |
▪ How to make Medical AI Systems safer? Simulating Vulnerabilities, and Threats in Multimodal Medical RAG System (Zuo et al., January 2026) |
| 286 |
▪ Red-Teaming Coding Agents from a Tool-Invocation Perspective: An Empirical Security Assessment (Xie et al., January 2026) |
| 287 |
▪ Temporal Attack Pattern Detection in Multi-Agent AI Workflows: An Open Framework for Training Trace-Based Security Models (Rosario, January 2026) |
| 288 |
▪ Structural Representations for Cross-Attack Generalization in AI Agent Threat Detection (Iyer, January 2026) |
| 289 |
▪ One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models (Gu et al., January 2026) |
| 290 |
▪ Improving LLM-Assisted Secure Code Generation through Retrieval-Augmented-Generation and Multi-Tool Feedback (Sriram et al., January 2026) |
| 291 |
▪ PatchBlock: A Lightweight Defense Against Adversarial Patches for Embedded EdgeAI Devices (Chattopadhyay et al., January 2026) |
| 292 |
▪ Large Empirical Case Study: Go-Explore adapted for AI Red Team Testing (Bhatt et al., January 2026) |
| 293 |
▪ Towards Provably Secure Generative AI: Reliable Consensus Sampling (Cui et al., January 2026) |
| 294 |
▪ How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench (Sajadi, Damevski, and Chatterjee, December 2025) |
| 295 |
▪ Pruning Graphs by Adversarial Robustness Evaluation to Strengthen GNN Defenses (Wang, December 2025) |
| 296 |
▪ Multi-Agent Framework for Threat Mitigation and Resilience in AI-Based Systems (Foundjem et al., December 2025) |
| 297 |
▪ Toward Secure and Compliant AI: Organizational Standards and Protocols for NLP Model Lifecycle Management (Arora and Hastings, December 2025) |
| 298 |
▪ Assessing the Software Security Comprehension of Large Language Models (Siddiq et al., December 2025) |
| 299 |
▪ AIAuditTrack: A Framework for AI Security system (Luo et al., December 2025) |
| 300 |
▪ AegisAgent: An Autonomous Defense Agent Against Prompt Injection Attacks in LLM-HARs (Wang et al., December 2025) |
| 301 |
▪ On the Effectiveness of Instruction-Tuning Local LLMs for Identifying Software Vulnerabilities (Park, Ko, and Cho, December 2025) |
| 302 |
▪ IoT-based Android Malware Detection Using Graph Neural Network With Adversarial Defense (Yumlembam et al., December 2025) |
| 303 |
▪ Certified Defense on the Fairness of Graph Neural Networks (Dong et al., December 2025) |
| 304 |
▪ DREAM: Dynamic Red-teaming across Environments for AI Models (Lu et al., December 2025) |
| 305 |
▪ Explainable and Fine-Grained Safeguarding of LLM Multi-Agent Systems via Bi-Level Graph Anomaly Detection (Pan et al., December 2025) |
| 306 |
▪ AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs (Gaikwad, December 2025) |
| 307 |
▪ Prefix Probing: Lightweight Harmful Content Detection for Large Language Models (Yang et al., December 2025) |
| 308 |
▪ Adversarial Robustness in Financial Machine Learning: Defenses, Economic Impact, and Governance Evidence (Baviskar, December 2025) |
| 309 |
▪ PrivateXR: Defending Privacy Attacks in Extended Reality Through Explainable AI-Guided Differential Privacy (Kundu, Ahmed, and Hoque, December 2025) |
| 310 |
▪ Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks (Shaheer et al., December 2025) |
| 311 |
▪ Autoencoder-based Denoising Defense against Adversarial Attacks on Object Detection (Song et al., December 2025) |
| 312 |
▪ Auto-Tuning Safety Guardrails for Black-Box Large Language Models (Abdulkadir, December 2025) |
| 313 |
▪ Quantifying Return on Security Controls in LLM Systems (Moulton, O'Brien, and Hastings, December 2025) |
| 314 |
▪ GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients (Razmjoo, Sharifian, and Shouraki, December 2025) |
| 315 |
▪ Behavior-Aware and Generalizable Defense Against Black-Box Adversarial Attacks for ML-Based IDS (Ennaji, Benkhelifa, and Mancini, December 2025) |
| 316 |
▪ SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models (Afane et al., December 2025) |
| 317 |
▪ Toward Intelligent and Secure Cloud: Large Language Model Empowered Proactive Defense (Zhou et al., December 2025) |
| 318 |
▪ CloudFix: Automated Policy Repair for Cloud Access Control Policies Using Large Language Models (Hall, Ungaro, and Eiers, December 2025) |
| 319 |
▪ Adaptive Intrusion Detection System Leveraging Dynamic Neural Models with Adversarial Learning for 5G/6G Networks (Neha and Bhatia, December 2025) |
| 320 |
▪ From Lab to Reality: A Practical Evaluation of Deep Learning Models and LLMs for Vulnerability Detection (Lu and Lagaisse, December 2025) |
| 321 |
▪ Chasing Shadows: Pitfalls in LLM Security Research (Evertz et al., December 2025) |
| 322 |
▪ AEIOU: A Unified Defense Framework against NSFW Prompts in Text-to-Image Models (Wang et al., December 2025) |
| 323 |
▪ Evaluating the robustness of adversarial defenses in malware detection systems (Jafari and Shameli-Sendi, December 2025) |
| 324 |
▪ Toward Patch Robustness Certification and Detection for Deep Learning Systems Beyond Consistent Samples (Zhou et al., December 2025) |
| 325 |
▪ VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection (Nie et al., December 2025) |
| 326 |
▪ Securing the Model Context Protocol: Defending LLMs Against Tool Poisoning and Adversarial Attacks (Jamshidi et al., December 2025) |
| 327 |
▪ IF-GUIDE: Influence Function-Guided Detoxification of LLMs (Coalson et al., December 2025) |
| 328 |
▪ Analyzing PDFs like Binaries: Adversarially Robust PDF Malware Analysis via Intermediate Representation and Language Model (Liu et al., December 2025) |
| 329 |
▪ ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior (Lu et al., December 2025) |
| 330 |
▪ Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models (Cretu et al., December 2025) |
| 331 |
▪ AutoGuard: A Self-Healing Proactive Security Layer for DevSecOps Pipelines Using Reinforcement Learning (Anugula et al., December 2025) |
| 332 |
▪ Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMs (Ma et al., December 2025) |
| 333 |
▪ Ensemble Privacy Defense for Knowledge-Intensive LLMs against Membership Inference Attacks (Fu et al., December 2025) |
| 334 |
▪ VeriLoRA: Fine-Tuning Large Language Models with Verifiable Security via Zero-Knowledge Proofs (Liao et al., December 2025) |
| 335 |
▪ OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning (Zhu et al., December 2025) |
| 336 |
▪ COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers (Wang et al., December 2025) |
| 337 |
▪ Factor(T,U): Factored Cognition Strengthens Monitoring of Untrusted AI (Sandoval and Rushing, December 2025) |
| 338 |
▪ Large Language Model based Smart Contract Auditing with LLMBugScanner (Yuan et al., December 2025) |
| 339 |
▪ Teleportation-Based Defenses for Privacy in Approximate Machine Unlearning (Maheri et al., December 2025) |
| 340 |
▪ Benchmarking and Understanding Safety Risks in AI Character Platforms (Wei, Zhang, and Tyson, December 2025) |
| 341 |
▪ Red Teaming Large Reasoning Models (Chen et al., December 2025) |
| 342 |
▪ An Empirical Study on the Security Vulnerabilities of GPTs (Wu, Wu, and Zheng, December 2025) |
| 343 |
▪ ShieldAgent: Shielding Agents via Verifiable Safety Policy Reasoning (Chen, Kang, and Li, December 2025) |
| 344 |
▪ A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures (Kong et al., December 2025) |
| 345 |
▪ Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities (Garg et al., December 2025) |
| 346 |
▪ Evaluating the Robustness of Large Language Model Safety Guardrails Against Adversarial Attacks (Young, December 2025) |
| 347 |
▪ Standardized Threat Taxonomy for AI Security, Governance, and Regulatory Compliance (Huwyler, December 2025) |
| 348 |
▪ GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision (Xiang et al., November 2025) |
| 349 |
▪ DUALGUAGE: Automated Joint Security-Functionality Benchmarking for Secure Code Generation (Pathak et al., November 2025) |
| 350 |
▪ VULSOLVER: Vulnerability Detection via LLM-Driven Constraint Solving (Li et al., November 2025) |
| 351 |
▪ SPQR: A Standardized Benchmark for Modern Safety Alignment Methods in Text-to-Image Diffusion Models (Alam et al., November 2025) |
| 352 |
▪ EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering (Gungor et al., November 2025) |
| 353 |
▪ Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization (Li et al., November 2025) |
| 354 |
▪ Shadows in the Code: Exploring the Risks and Defenses of LLM-based Multi-Agent Software Development Systems (Wang et al., November 2025) |
| 355 |
▪ SHIELD: Secure Hypernetworks for Incremental Expansion Learning Defense (Krukowski et al., November 2025) |
| 356 |
▪ Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks (Ji et al., November 2025) |
| 357 |
▪ DRIP: Defending Prompt Injection via Token-wise Representation Editing and Residual Instruction Fusion (Liu et al., November 2025) |
| 358 |
▪ N-GLARE: An Non-Generative Latent Representation-Efficient LLM Safety Evaluator (Lin et al., November 2025) |
| 359 |
▪ Certified but Fooled! Breaking Certified Defences with Ghost Certificates (Vo et al., November 2025) |
| 360 |
▪ ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning (Guan et al., November 2025) |
| 361 |
▪ AI Kill Switch for malicious web-based LLM agent (Lee and Park, November 2025) |
| 362 |
▪ DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt (Zhang et al., November 2025) |
| 363 |
▪ Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning (Zimmer and Karame, November 2025) |
| 364 |
▪ SGuard-v1: Safety Guardrail for Large Language Models (Lee et al., November 2025) |
| 365 |
▪ Defending Unauthorized Model Merging via Dual-Stage Weight Protection (Chen et al., November 2025) |
| 366 |
▪ On the Trade-Off Between Transparency and Security in Adversarial Machine Learning (Fenaux, Srinivasa, and Kerschbaum, November 2025) |
| 367 |
▪ TZ-LLM: Protecting On-Device Large Language Models with Arm TrustZone (Wang et al., November 2025) |
| 368 |
▪ InfoDecom: Decomposing Information for Defending against Privacy Leakage in Split Inference (Deng, Lu, and Duan, November 2025) |
| 369 |
▪ DualTAP: A Dual-Task Adversarial Protector for Mobile MLLM Agents (Zhang et al., November 2025) |
| 370 |
▪ SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization (Rong et al., November 2025) |
| 371 |
▪ AI Bill of Materials and Beyond: Systematizing Security Assurance through the AI Risk Scanning (AIRS) Framework (Nathanson et al., November 2025) |
| 372 |
▪ VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization (Li, Yu, and Wang, November 2025) |
| 373 |
▪ Securing Generative AI in Healthcare: A Zero-Trust Architecture Powered by Confidential Computing on Google Cloud (Amanna and Shinde, November 2025) |
| 374 |
▪ PATCHEVAL: A New Benchmark for Evaluating LLMs on Patching Real-World Vulnerabilities (Wei et al., November 2025) |
| 375 |
▪ Rethinking the Evaluation of Secure Code Generation (Dai, Xu, and Tao, November 2025) |
| 376 |
▪ AdaptDel: Adaptable Deletion Rate Randomized Smoothing for Certified Robustness (Huang et al., November 2025) |
| 377 |
▪ iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verification (Xiong et al., November 2025) |
| 378 |
▪ Provable Repair of Deep Neural Network Defects by Preimage Synthesis and Property Refinement (Ma et al., November 2025) |
| 379 |
▪ Oblivionis: A Lightweight Learning and Unlearning Framework for Federated Large Language Models (Zhang et al., November 2025) |
| 380 |
▪ Mitigating Sexual Content Generation via Embedding Distortion in Text-conditioned Diffusion Models (Ahn and Jung, November 2025) |
| 381 |
▪ Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks (Chen et al., November 2025) |
| 382 |
▪ LiteUpdate: A Lightweight Framework for Updating AI-Generated Image Detectors (Lu et al., November 2025) |
| 383 |
▪ Efficient LLM Safety Evaluation through Multi-Agent Debate (Lin et al., November 2025) |
| 384 |
▪ Preserving security in a world with powerful AI Considerations for the future Defense Architecture (Generous, Cook, and Pruet, November 2025) |
| 385 |
▪ Amulet: a Python Library for Assessing Interactions Among ML Defenses and Risks (Waheed et al., November 2025) |
| 386 |
▪ ConVerse: Benchmarking Contextual Safety in Agent-to-Agent Conversations (Gomaa, Salem, and Abdelnabi, November 2025) |
| 387 |
▪ Specification-Guided Vulnerability Detection with Large Language Models (Zhu et al., November 2025) |
| 388 |
▪ SME-TEAM: Leveraging Trust and Ethics for Secure and Responsible Use of AI and LLMs in SMEs (Sarker et al., November 2025) |
| 389 |
▪ LLM-Driven SAST-Genius: A Hybrid Static Analysis Framework for Comprehensive and Actionable Security (Agrawal and Ahi, November 2025) |
| 390 |
▪ SecRepoBench: Benchmarking Code Agents for Secure Code Completion in Real-World Repositories (Shen et al., November 2025) |
| 391 |
▪ LM-Fix: Lightweight Bit-Flip Detection and Rapid Recovery Framework for Language Models (Tahmasivand et al., November 2025) |
| 392 |
▪ RepoMark: A Data-Usage Auditing Framework for Code Large Language Models (Qu et al., November 2025) |
| 393 |
▪ AgentBnB: A Browser-Based Cybersecurity Tabletop Exercise with Large Language Model Support and Retrieval-Aligned Scaffolding (Anwar and Liu, November 2025) |
| 394 |
▪ Keys in the Weights: Transformer Authentication Using Model-Bound Latent Representations (Okatan et al., November 2025) |
| 395 |
▪ DRIP: Defending Prompt Injection via De-instruction Training and Residual Fusion Model Architecture (Liu, Lin, and Dong, November 2025) |
| 396 |
▪ SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents (Yin et al., November 2025) |
| 397 |
▪ On Selecting Few-Shot Examples for LLM-based Code Vulnerability Detection (Hannan et al., November 2025) |
| 398 |
▪ SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation (Su et al., November 2025) |
| 399 |
▪ LLM-based Multi-class Attack Analysis and Mitigation Framework in IoT/IIoT Networks (Ikbarieh, Gupta, and Mahalal, November 2025) |
| 400 |
▪ The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs (Liu et al., October 2025) |
| 401 |
▪ Improving LLM Safety Alignment with Dual-Objective Optimization (Zhao et al., October 2025) |
| 402 |
▪ IRCopilot: Automated Incident Response with Large Language Models (Lin et al., October 2025) |
| 403 |
▪ ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models (Jin et al., October 2025) |
| 404 |
▪ Who Grants the Agent Power? Defending Against Instruction Injection via Task-Centric Access Control (Cai et al., October 2025) |
| 405 |
▪ SIRAJ: Diverse and Efficient Red-Teaming for LLM Agents via Distilled Structured Reasoning (Zhou et al., October 2025) |
| 406 |
▪ SoK: Honeypots & LLMs, More Than the Sum of Their Parts? (Bridges et al., October 2025) |
| 407 |
▪ OpenGuardrails: A Configurable, Unified, and Scalable Guardrails Platform for Large Language Models (Wang and Li, October 2025) |
| 408 |
▪ Secure Retrieval-Augmented Generation against Poisoning Attacks (Cheng et al., October 2025) |
| 409 |
▪ FaRAccel: FPGA-Accelerated Defense Architecture for Efficient Bit-Flip Attack Resilience in Transformer Models (Nazari et al., October 2025) |
| 410 |
▪ SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability (Xu et al., October 2025) |
| 411 |
▪ Adversarially-Aware Architecture Design for Robust Medical AI Systems (Gerhart and Iyangar, October 2025) |
| 412 |
▪ Cybersecurity AI Benchmark (CAIBench): A Meta-Benchmark for Evaluating Cybersecurity AI Agents (Sanz-G\'omez et al., October 2025) |
| 413 |
▪ SAGE: A Generic Framework for LLM Safety Evaluation (Jindal et al., October 2025) |
| 414 |
▪ T2I-RiskyPrompt: A Benchmark for Safety Evaluation, Attack, and Defense on Text-to-Image Model (Zhang et al., October 2025) |
| 415 |
▪ SecureLearn - An Attack-agnostic Defense for Multiclass Machine Learning Against Data Poisoning Attacks (Paracha et al., October 2025) |
| 416 |
▪ Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs (Davies et al., October 2025) |
| 417 |
▪ SBASH: a Framework for Designing and Evaluating RAG vs. Prompt-Tuned LLM Honeypots (Adebimpe, Neukirchen, and Welsh, October 2025) |
| 418 |
▪ FLAMES: Fine-tuning LLMs to Synthesize Invariants for Smart Contract Security (Eshghie et al., October 2025) |
| 419 |
▪ Soft Instruction De-escalation Defense (Walter et al., October 2025) |
| 420 |
▪ Towards Strong Certified Defense with Universal Asymmetric Randomization (Hong et al., October 2025) |
| 421 |
▪ Enhancing Security in Deep Reinforcement Learning: A Comprehensive Survey on Adversarial Attacks and Defenses (Yichao et al., October 2025) |
| 422 |
▪ SAID: Empowering Large Language Models with Self-Activating Internal Defense (Chen et al., October 2025) |
| 423 |
▪ SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks (Lee et al., October 2025) |
| 424 |
▪ AegisMCP: Online Graph Intrusion Detection for Tool-Augmented LLMs on Edge Devices (Zhan et al., October 2025) |
| 425 |
▪ Monitoring LLM-based Multi-Agent Systems Against Corruptions via Node Evaluation (Wu et al., October 2025) |
| 426 |
▪ Defending Against Prompt Injection with DataFilter (Wang et al., October 2025) |
| 427 |
▪ OpenGuardrails: An Open-Source Context-Aware AI Guardrails Platform (Wang and Li, October 2025) |
| 428 |
▪ Prompting the Priorities: A First Look at Evaluating LLMs for Vulnerability Triage and Prioritization (Haddad et al., October 2025) |
| 429 |
▪ SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal Steering (Lin et al., October 2025) |
| 430 |
▪ Breaking and Fixing Defenses Against Control-Flow Hijacking in Multi-Agent Systems (Jha et al., October 2025) |
| 431 |
▪ VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search (Li, Zhao, and Liu, October 2025) |
| 432 |
▪ Watermark Robustness and Radioactivity May Be at Odds in Federated Learning (Huang, Shao, and Baluta, October 2025) |
| 433 |
▪ Verifiable Fine-Tuning for LLMs: Zero-Knowledge Training Proofs Bound to Data Provenance and Policy (Akgul et al., October 2025) |
| 434 |
▪ SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detection (Feng and Pan, October 2025) |
| 435 |
▪ Breaking Guardrails, Facing Walls: Insights on Adversarial AI for Defenders & Researchers (Bertollo, Bodemir, and Burgess, October 2025) |
| 436 |
▪ A Framework for Rapidly Developing and Deploying Protection Against Large Language Model Attacks (Swanda et al., October 2025) |
| 437 |
▪ GuardReasoner: Towards Reasoning-based LLM Safeguards (Liu et al., October 2025) |
| 438 |
▪ Towards Proactive Defense Against Cyber Cognitive Attacks (Rushing, Umeokolo, and Xu, October 2025) |
| 439 |
▪ Generalist++: A Meta-learning Framework for Mitigating Trade-off in Adversarial Training (Wang et al., October 2025) |
| 440 |
▪ GRIDAI: Generating and Repairing Intrusion Detection Rules via Collaboration among Multiple LLM-based Agents (Li et al., October 2025) |
| 441 |
▪ Countermind: A Multi-Layered Security Architecture for Large Language Models (Schwarz, October 2025) |
| 442 |
▪ TraceAegis: Securing LLM-Based Agents via Hierarchical and Behavioral Anomaly Detection (Liu et al., October 2025) |
| 443 |
▪ Pharmacist: Safety Alignment Data Curation for Large Language Models against Harmful Fine-tuning (Liu et al., October 2025) |
| 444 |
▪ SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents (Ying et al., October 2025) |
| 445 |
▪ Fortifying LLM-Based Code Generation with Graph-Based Reasoning on Secure Coding Practices (Patir et al., October 2025) |
| 446 |
▪ Toward a Unified Security Framework for AI Agents: Trust, Risk, and Liability (Mo et al., October 2025) |
| 447 |
▪ Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models (Meng et al., October 2025) |
| 448 |
▪ From Defender to Devil? Unintended Risk Interactions Induced by LLM Defenses (Meng et al., October 2025) |
| 449 |
▪ RedTWIZ: Diverse LLM Red Teaming via Adaptive Attack Planning (Horal et al., October 2025) |
| 450 |
▪ A Generative Approach to LLM Harmfulness Mitigation with Red Flag Tokens (Dobre et al., October 2025) |
| 451 |
▪ DoomArena: A framework for Testing AI Agents Against Evolving Security Threats (Boisvert et al., October 2025) |
| 452 |
▪ VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation (Miculicich et al., October 2025) |
| 453 |
▪ Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling (Llewellyn et al., October 2025) |
| 454 |
▪ SafeGuider: Robust and Practical Content Safety Control for Text-to-Image Models (Qi et al., October 2025) |
| 455 |
▪ AutoBnB-RAG: Enhancing Multi-Agent Incident Response with Retrieval-Augmented Generation (Liu and Anwar, October 2025) |
| 456 |
▪ Thought Purity: A Defense Framework For Chain-of-Thought Attack (Xue et al., October 2025) |
| 457 |
▪ SteerDiff: Steering towards Safe Text-to-Image Diffusion Models (Zhang, He, and Chen, October 2025) |
| 458 |
▪ Proactive defense against LLM Jailbreak (Zhao et al., October 2025) |
| 459 |
▪ MulVuln: Enhancing Pre-trained LMs with Shared and Language-Specific Knowledge for Multilingual Vulnerability Detection (Nguyen et al., October 2025) |
| 460 |
▪ SVDefense: Effective Defense against Gradient Inversion Attacks via Singular Value Decomposition (Luo, Yau, and Song, October 2025) |
| 461 |
▪ Permissioned LLMs: Enforcing Access Control in Large Language Models (Jayaraman et al., October 2025) |
| 462 |
▪ A-MemGuard: A Proactive Defense Framework for LLM-Based Agent Memory (Wei et al., October 2025) |
| 463 |
▪ Defend LLMs Through Self-Consciousness (Huang and Paula, October 2025) |
| 464 |
▪ UpSafe$^\circ$C: Upcycling for Controllable Safety in Large Language Models (Sun et al., October 2025) |
| 465 |
▪ TAIBOM: Bringing Trustworthiness to AI-Enabled Systems (Safronov et al., October 2025) |
| 466 |
▪ Adaptive Federated Learning Defences via Trust-Aware Deep Q-Networks (Palit, October 2025) |
| 467 |
▪ A Multi-Agent LLM Defense Pipeline Against Prompt Injection Attacks (Hossain et al., October 2025) |
| 468 |
▪ A Call to Action for a Secure-by-Design Generative AI Paradigm (Alharthi and Garcia, October 2025) |
| 469 |
▪ MAVUL: Multi-Agent Vulnerability Detection via Contextual Reasoning and Interactive Refinement (Li et al., October 2025) |
| 470 |
▪ Detecting Instruction Fine-tuning Attacks on Language Models using Influence Function (Li, October 2025) |
| 471 |
▪ Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety (Zhang et al., October 2025) |
| 472 |
▪ Adversarial Defense in Cybersecurity: A Systematic Review of GANs for Threat Detection and Mitigation (Ndayipfukamiye et al., October 2025) |
| 473 |
▪ QGuard:Question-based Zero-shot Guard for Multi-modal LLM Safety (Lee et al., October 2025) |
| 474 |
▪ Federated Learning Resilient to Byzantine Attacks and Data Heterogeneity (Zuo et al., September 2025) |
| 475 |
▪ SafeSearch: Automated Red-Teaming for the Safety of LLM-Based Search Agents (Dong et al., September 2025) |
| 476 |
▪ GSPR: Aligning LLM Safeguards as Generalizable Safety Policy Reasoners (Li et al., September 2025) |
| 477 |
▪ Benchmarking LLM-Assisted Blue Teaming via Standardized Threat Hunting (Meng et al., September 2025) |
| 478 |
▪ ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search (Shen et al., September 2025) |
| 479 |
▪ Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment (Kim et al., September 2025) |
| 480 |
▪ Think Broad, Act Narrow: CWE Identification with Multi-Agent Large Language Models (Sayagh and Ghafari, August 2025) |
| 481 |
▪ ConfGuard: A Simple and Effective Backdoor Detection for Large Language Models (Wang et al., August 2025) |
| 482 |
▪ Provably Secure Retrieval-Augmented Generation (Zhou, Feng, and Yang, August 2025) |
| 483 |
▪ FedGuard: A Diverse-Byzantine-Robust Mechanism for Federated Learning with Major Malicious Clients (Jiang et al., August 2025) |
| 484 |
▪ CEE: An Inference-Time Jailbreak Defense for Embodied Intelligence via Subspace Concept Rotation (Yang et al., August 2025) |
| 485 |
▪ Bridging Privacy and Robustness for Trustworthy Machine Learning (Zhang and Chen, July 2025) |
| 486 |
▪ Large Language Model-Based Framework for Explainable Cyberattack Detection in Automatic Generation Control Systems (Sharshar et al., July 2025) |
| 487 |
▪ Strategic Deflection: Defending LLMs from Logit Manipulation (Rachidy et al., July 2025) |
| 488 |
▪ Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security (Dai et al., July 2025) |
| 489 |
▪ GUARD-CAN: Graph-Understanding and Recurrent Architecture for CAN Anomaly Detection (Kim and Kim, July 2025) |
| 490 |
▪ SDD: Self-Degraded Defense against Malicious Fine-tuning (Chen et al., July 2025) |
| 491 |
▪ OneShield -- the Next Generation of LLM Guardrails (DeLuca et al., July 2025) |
| 492 |
▪ Quantifying Security Vulnerabilities: A Metric-Driven Security Analysis of Gaps in Current AI Standards (Madhavan et al., July 2025) |
| 493 |
▪ Repairing vulnerabilities without invisible hands. A differentiated replication study on LLMs (Camporese and Massacci, July 2025) |
| 494 |
▪ PurpCode: Reasoning for Safer Code Generation (Liu et al., July 2025) |
| 495 |
▪ SAVANT: Vulnerability Detection in Application Dependencies through Semantic-Guided Reachability Analysis (Lingxiang et al., July 2025) |
| 496 |
▪ Layer-Aware Representation Filtering: Purifying Finetuning Data to Preserve LLM Safety Alignment (Li et al., July 2025) |
| 497 |
▪ CASCADE: LLM-Powered JavaScript Deobfuscator at Google (Jiang et al., July 2025) |
| 498 |
▪ LLM Meets the Sky: Heuristic Multi-Agent Reinforcement Learning for Secure Heterogeneous UAV Networks (Zheng et al., July 2025) |
| 499 |
▪ CP-uniGuard: A Unified, Probability-Agnostic, and Adaptive Framework for Malicious Agent Detection and Defense in Multi-Agent Embodied Perception Systems (Hu et al., July 2025) |
| 500 |
▪ DP-TLDM: Differentially Private Tabular Latent Diffusion Model (Zhu et al., July 2025) |
| 501 |
▪ Recent Advances in Malware Detection: Graph Learning and Explainability (Shokouhinejad et al., July 2025) |
| 502 |
▪ "We Need a Standard": Toward an Expert-Informed Privacy Label for Differential Privacy (Dibia et al., July 2025) |
| 503 |
▪ ACFIX: Guiding LLMs with Mined Common RBAC Practices for Context-Aware Repair of Access Control Vulnerabilities in Smart Contracts (Zhang et al., July 2025) |
| 504 |
▪ OMNISEC: LLM-Driven Provenance-based Intrusion Detection via Retrieval-Augmented Behavior Prompting (Cheng et al., July 2025) |
| 505 |
▪ Defending Against Unforeseen Failure Modes with Latent Adversarial Training (Casper et al., July 2025) |
| 506 |
▪ FedStrategist: A Meta-Learning Framework for Adaptive and Robust Aggregation in Federated Learning (Haque, Kamal, and Hossain, July 2025) |
| 507 |
▪ PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants (Liu et al., July 2025) |
| 508 |
▪ PromptArmor: Simple yet Effective Prompt Injection Defenses (Shi et al., July 2025) |
| 509 |
▪ A Privacy-Centric Approach: Scalable and Secure Federated Learning Enabled by Hybrid Homomorphic Encryption (Nguyen, Khan, and Michalas, July 2025) |
| 510 |
▪ PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training (Du, July 2025) |
| 511 |
▪ Defense Against Prompt Injection Attack by Leveraging Attack Techniques (Chen et al., July 2025) |
| 512 |
▪ GIFT: Gradient-aware Immunization of diffusion models against malicious Fine-Tuning with safe concepts retention (Abdalla et al., July 2025) |
| 513 |
▪ Risks of ignoring uncertainty propagation in AI-augmented security pipelines (Mezzi et al., July 2025) |
| 514 |
▪ JailDAM: Jailbreak Detection with Adaptive Memory for Vision-Language Model (Nian et al., July 2025) |
| 515 |
▪ SHIELD: A Secure and Highly Enhanced Integrated Learning for Robust Deepfake Detection against Adversarial Attacks (Uddin et al., July 2025) |
| 516 |
▪ Safeguarding Federated Learning-based Road Condition Classification (Liu and Papadimitratos, July 2025) |
| 517 |
▪ Thought Purity: Defense Paradigm For Chain-of-Thought Attack (Xue et al., July 2025) |
| 518 |
▪ Expanding ML-Documentation Standards For Better Security (Appel, July 2025) |
| 519 |
▪ A Generative Approach to LLM Harmfulness Detection with Special Red Flag Tokens (Xhonneux et al., July 2025) |
| 520 |
▪ A Review of Privacy Metrics for Privacy-Preserving Synthetic Data Generation (Trudslev et al., July 2025) |
| 521 |
▪ From Alerts to Intelligence: A Novel LLM-Aided Framework for Host-based Intrusion Detection (Sun et al., July 2025) |
| 522 |
▪ Contrastive-KAN: A Semi-Supervised Intrusion Detection Framework for Cybersecurity with scarce Labeled Data (Alikhani and Kazemi, July 2025) |
| 523 |
▪ Defense-as-a-Service: Black-box Shielding against Backdoored Graph Models (Yang et al., July 2025) |
| 524 |
▪ Securing Transformer-based AI Execution via Unified TEEs and Crypto-protected Accelerators (Xue et al., July 2025) |
| 525 |
▪ SynthGuard: Redefining Synthetic Data Generation with a Scalable and Privacy-Preserving Workflow Framework (Brito et al., July 2025) |
| 526 |
▪ AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions (Liu et al., July 2025) |
| 527 |
▪ Entangled Threats: A Unified Kill Chain Model for Quantum Machine Learning Security (Debus et al., July 2025) |
| 528 |
▪ ADAPT: A Pseudo-labeling Approach to Combat Concept Drift in Malware Detection (Alam, Piplai, and Rastogi, July 2025) |
| 529 |
▪ Beyond the Worst Case: Extending Differential Privacy Guarantees to Realistic Adversaries (Swanberg et al., July 2025) |
| 530 |
▪ A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning (Gluch and Goldwasser, July 2025) |
| 531 |
▪ Adaptive Randomized Smoothing: Certified Adversarial Robustness for Multi-Step Defences (Lyu et al., July 2025) |
| 532 |
▪ Adversarial Defenses via Vector Quantization (Dong and Mao, July 2025) |
| 533 |
▪ Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking (Wu et al., July 2025) |
| 534 |
▪ Defending Against Prompt Injection With a Few DefensiveTokens (Chen et al., July 2025) |
| 535 |
▪ Can Large Language Models Improve Phishing Defense? A Large-Scale Controlled Experiment on Warning Dialogue Explanations (Cau et al., July 2025) |
| 536 |
▪ Hybrid LLM-Enhanced Intrusion Detection for Zero-Day Threats in IoT Networks (Al-Hammouri et al., July 2025) |
| 537 |
▪ Saffron-1: Safety Inference Scaling (Qiu et al., July 2025) |
| 538 |
▪ LDP$^3$: An Extensible and Multi-Threaded Toolkit for Local Differential Privacy Protocols and Post-Processing Methods (Balioglu, Khodaie, and Gursoy, July 2025) |
| 539 |
▪ TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data (Cheruvu et al., July 2025) |
| 540 |
▪ PROTEAN: Federated Intrusion Detection in Non-IID Environments through Prototype-Based Knowledge Sharing (Chennoufi et al., July 2025) |
| 541 |
▪ MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teamin<br><br> (Guo et al., May 2025) |
| 542 |
▪ Improving LLM Outputs Against Jailbreak Attacks with Expert Model Integration<br><br> (Tsmindashvili et al., May 2025) |
| 543 |
▪ LLM-Based Threat Detection and Prevention Framework for IoT Ecosystems<br><br> (Otoum, Asad, and Nayak, May 2025) |
| 544 |
▪ Securing RAG: A Risk Assessment and Mitigation Framework<br><br> (Ammann et al., May 2025) |
| 545 |
▪ Large Language Model Sentinel: LLM Agent for Adversarial Purification (Lin, Tanaka, and Zhao, Apr 2025) |
| 546 |
▪ JailGuard: A Universal Detection Framework for LLM Prompt-based Attacks (Zhang et al, Mar 2025) |