Continuation of Nicholas Carlini's list, covering arXiv submissions since 2025-09-15.
TF-IDF + logistic regression trained on the original 13,697-entry list (held-out AUC 0.998); papers scoring ≥40% shown, grouped by arXiv ID month.
Total: 3179 papers.
Last update: 2026-10-11.


Full list (3179 papers, all months): https://nsl.cs.waseda.ac.jp/papers/advex_papers.html
Download: CSV
Download: CSV
Most recent month
2026-10 (106 papers)
- [99%] Boosting Transferable Adversarial Attacks against Deep Reinforcement Learning
Zexin Li; Ruili Yao; Yiming Zeng; Xiaoxue Gao - [99%] ATLAS-AL: Adaptive Trust-Region for Latent Adversarial Searches via Active Learning
Marsalis Gibson; Claire Tomlin; Shankar Sastry - [99%] Adaptive Model Inversion Attacks Generalize a Privacy-Robustness Tradeoff
Shailen Smith; Rasmus Torp; Adam Breuer - [99%] Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Driving
Heyam Bin Jahlan Areej Alhothali Abeer Alhothali - [99%] Detection-Guided Adaptive Purification with Diffusion Models for Robust Audio Deepfake Detection
Muhammed Salih Kayhan; Qiben Yan - [99%] Collaboratively Guided Adversarial Robust Distillation with Teacher-Favorable Examples
Zhi Li; Haowei Liu; Hongchen Yang; Xiaoxuan Wang; Song Gao; Shaowen Yao; Wei Zhou - [98%] Breaking Adversarial Transferability in Fine-Tuned Speech Recognition
Mojtaba Nafez; Aref Mousavi; Mohammad Ebrahim Mahdavi; Mobina Poulaei; Kiarash Kiani Feriz; Mohammad Hossein Rohban - [95%] CalCErt: Bin-wise Certification of Confidence Calibration in Medical Image Classification
Leo Fillioux; Stergios Christodoulidis; Stergios Christodoulidis; Maria Vakalopoulou; Jose Dolz - [95%] VCR-Bench: A Modular Open-Source Benchmark for Video Classification Robustness
Maksim Plinskiy; Aleksandr Gushchin; Sergey Lavrushkin; Dmitriy S. Vatolin; Anastasia Antsiferova - [95%] Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks
Yannick Lunk; Atell Yehor Krasnopolsky; Damien Garreau; Leon Bungert - [94%] Target-free Latent Safety Alignment
Luoyu Chen; Weiqi Wang; Chenhan Zhang; Zhiyi Tian; Yuxian Huang; Shui Yu - [94%] Pareto-Improving Adversarial Attacks with Primal-Dual Regularization
Yang Dai; Longfei Zhang; Wei Tao; Li Shen; Jincai Huang; Qing Tao - [94%] Transferable Adversarial Robustness for Speech Foundation Models via Hierarchical Stabilization
Aref Mousavi; Shahab Sherafat; Kiarash Kiani Feriz; Amirparsa Safari; Raoof Zare Moayedi; Mohammad Hossein Rohban; Mohammad Sabokrou - [93%] Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMs
Xiaojun Jia; Simeng Qin; Yiming Li; Jie Liao; Sensen Gao; Ke Ma; Yang Liu; Xiaochun Cao - [93%] Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitation
Xiaoyi Pang; Haoyue Feng; Quanxin Shou; Yikun Miao; Zhengyang Yan; Song Guo - [92%] Detect and Suppress: A Mechanistic Defense against Adversarial Patches in VLA Models
Yukiya Horiba; Koshiro Aoki; Shunsuke Yasuki; Bum Jun Kim; Taiki Miyanishi - [92%] GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks
Arash Vashagh; Roozbeh Razavi-Far - [91%] Online AutoML: Evaluating Poisoning Attacks on Adversarial Training Defense Strategy in IoT Networks
Chukwunonso Henry Nwokoye; Khalil El-Khatib; Li Yang - [89%] Adversarially Trained Linear Transformers Are Optimal Robust In-Context Learners for Gaussian Mixtures
Soichiro Kumano - [88%] MRCert: Towards Post-deployment Patch Robustness Certification for Adversarially Patched Samples via Type-specific Masking
Qilin Zhou; Zhengyuan Wei; Haipeng Wang; Zhuo Wang; Shuo Liu; W. K. Chan - [85%] When Normalization Selects the Sign: Auditing Robustness Ablations in Quantum Attention
Owen Friedewald; Srikar Alla; Ali Shiri Sichani; Chi-Ren Shyu - [84%] SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
Chaeeun Han; Soodeh Atefi; Yevgeniy Vorobeychik; Aron Laszka - [84%] Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
Narek Maloyan - [84%] JASPER: Special Session on Joint Reliability And Security Assessment of SPlit Computing for Edge Robustness
Enrico Magliano; Giuseppe Esposito; Amir Hossein Shahdadian; Rama Mounika Kodamanchili; Juan David Guerrero Balaguera; Juan Esteban Rodriguez Condia; Annachiara Ruospo; Roberta Siciliano; Stefano Di Carlo; Maksim Jenihhin; Marco Levorato; Alessandro Savino; Matteo Sonza Reorda; Christian Herglotz; Michael Hübner; Mahdi Taheri - [82%] Reward Stealing Attack on Large Language Models
Jiaming Qian; Pengyang Zhou; Jiahe Xu; Chaochao Chen - [80%] Detecting Adversarial Images through Response Profiles of Vision-Language Models
Arash Vashagh; Roozbeh Razavi-Far - [78%] TwinViT-DeepJSCC: Adversarially Robust Semantic Image Communication
Maedeh Fallahreyhani; Paeiz Azmi; Nader Mokari; M. Reza Abedi; Melike Erol-Kantarci; Eduard A. Jorswieck - [78%] Visual Memory Attacks Can Persist Through The KV Cache
David Dobre; Leo Schwinn; Gauthier Gidel; Spandana Gella; Perouz Taslakian; Pierre-André Noël - [77%] The Fragility of Trigger-Tag Mechanisms for Misuse Detection in Open-Weight LLMs
Toluwani Aremu; Manit Baser; Mohan Gurusamy; Nils Lukas; Dinil Mon Divakaran - [76%] The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
Yibo Zhang; Tianrong Guan; Liang Lin; Puze Wang; Jin Wang; Qingsong Wen - [75%] TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
Ruizhi Xu; Wei Xu; Sibo Zhu - [75%] Adversarial RL for Port-Scan Evasion: Attacker Feature Visibility in Edge-Deployed IDS
Logan Andrew North; Priya Sanjay Kaluskar; Shasi Kumar Ramachandran Prabhu; Peilong Li; Suman Saha - [74%] HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
Xiqiao Xiong; Moxin Li; Zhixin Ma; Ouxiang Li; Wenjie Wang; Fuli Feng; Xiangnan He - [74%] Beyond Reward Suppression: Near-Optimal Offline Attacks on Warm-Start Bandits with Bounded Rewards
Qirun Zeng; Manhin Poon; Xiangxiang Dai; Qixin Zhang; Jinhang Zuo - [72%] Reactivating Alignment: Defending LLMs from Jailbreaks via Intention-Aware Input-Output Matching
Luoyu Chen; Weiqi Wang; Chenhan Zhang; Zhiyi Tian; Shui Yu - [72%] Reflections and Fragments: Securing LLMs Against Sequential Mosaic Attacks
Emanuele La Malfa; Saar Cohen; Gabriele La Malfa; Mickel Liu; Christian Schroeder de Witt; Natasha Jaques; Michael J. Wooldridge - [72%] Lipschitz Thinking: Ten Years of Certifiable-by-Design Robust Neural Networks
Fabio Brau; Giorgio Piras; Maura Pintor; Battista Biggio - [72%] TAPDreamer: Transferable Adversarial Patches for World Action Models
Xuanyu Lu; Fengqing Jiang; Kaiyuan Zheng; Yichen Feng; Yaorui Ding; Yuetai Li; Zhen Xiang; Bhaskar Ramasubramanian; Basel Alomair; Luyao Niu; Radha Poovendran - [70%] MLCommons Jailbreak Benchmark v1.0
Carsten Maple; Cagatay Yucel; Isaac Holeman; Chris Knotz; Peter Mattson; James Goel; Jonathan Petit; Sean McGregor; James Ezick; Abhishek Kumar; Alicia Parrish; Murali Emani; Kashyap Iyer; Faiza Khan Khattak; Washington Mbonu; Daniel Machlab; Eileen Long; Shaona Ghosh; Jibin Varghese; Roman Lutz; Andrew Gruen; Bennett Hillenbrand; Prabal Gupta; Mohammed Serrhini; Dhivya Nagasubramanian; Aakash Gupta; Jun; Lu; Kurt Bollacker; Chang Liu; Jonathan Petit; Cong Chen; Jean-Philippe Monteuuis; Brent Miller; Apurv Verma; Roman Eng; Armstrong Foundjem; Mohammed Serrhini - [70%] The TellTail of Embeddings: Fingerprinting Retrievers in Black-Box Systems
Abdullah Garra; Matan Ben-Tov; Mahmood Sharif - [68%] Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Minoo Kim; Vasileios Lampos; George Drayson - [67%] DeBERTa-ConPara: Attack-Aware and Deployment-Realistic Detection of AI-Generated Text
Mohamed Mady; Yupei Li; Johannes Reschke; Björn W. Schuller - [66%] Exploring Weaknesses of Generative Image Watermarks against Latent Frequency Masking
Kirill Aistov; Khaled Abud; Irina Serzhenko; Egor Kovalev; Aleksey Yakushev; Aleksandr Akimenkov; Dmitry Obydenkov; Yury Markin; Sergey Lavrushkin; Dmitriy Vatolin; Anastasia Antsiferova - [65%] Memetic Trojans: Social Contagions as Carriers of Adversarial Payloads in Agent Networks
Birk Torpmann-Hagen; Finn Schwall; Leon Moonen - [65%] Curvature Under Attack in hZACH-ViT: Gauge Symmetry, Boundary Saturation, and Adversarial Failure
Athanasios Angelakis; Marta Gomez-Barrero - [65%] COPEX: Benchmarking LLM Robustness to Adversarial Context Across Model Context Protocol Layers
Nahom Birhan; Mehrdad Rostamzadeh; Sidhant Narula; Mahmoud Nazzal; Mohammad Ghasemigol; Daniel Takabi - [65%] What Response Marginals Miss: Adaptive Query Complexity of Functional Backdoor Recovery
Yunjae Hwang; Byoungjin Seok - [65%] BRANCH: Bypassing Multi-Scanner AI Guardrails
William Hackett; Peter Garraghan - [65%] Does Target Alignment Mean Target Recovery? An Evidence-Ladder Study of Adversarial Claims on Contrastive Encoders
Tao Yang; Jianying Zhou - [65%] Neural Network Verification for Deep Joint Source-Channel Coding
Thanh Le; Hai Duong; Takeshi Matsumura; ThanhVu Nguyen - [64%] Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally
Arun Josephraj Arokiaraj; Zekun Wu; Adriano Koshiyama - [64%] Does On-Policy Distillation for Safety Pose Backdoor Risks?
Jian Luo; Kehan Qi; Qingqiao Hu; Meilong Xu; Jiacheng Qiu; Weimin Lyu; Jiawei Zhou; Chao Chen - [63%] Certification of Real Images through Calibrated Content Authentication
Sarim Hashmi; Abdelrahman Elsayed; Mohammed Talha Alam; Samuele Poppi; Nils Lukas - [62%] Adversarial Training for Deep Hedging in Nonstationary Markets
Philipp J. Schneider; Lukas Looser; Antoine Garin; Shuhan Liu; Daniel Kuhn - [61%] Safe Image Generation via Reinforcement Learning
Eungyeol Han; Jong-Seok Lee - [60%] From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model
Yuelin Han - [60%] Towards Robust Numerical Claim Verification
Peter Røysland Aarnes; Vinay Setty - [60%] RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents
Mohamed Dhouib; Clement Elliker; Alexi Canesse; Maël Jenny; Lucas-Andrei Thil; Mahammed El-Sharkawy; Sonia Vanier; Elie Bursztein - [60%] AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
Sarim Hashmi; Mukul Ranjan; Kshitij Mishra; Mikhail Kuznetsov; Praneeth Vepakomma; Nils Lukas - [58%] Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure Experience
Zixuan Wang; Hao Li; Fengyu Gao; G. Edward Suh; Yi Zeng; Yevgeniy Vorobeychik; Ning Zhang; Chaowei Xiao - [58%] APEX: Active Protection at Execution Boundaries for LLM Agents
Xinran Zheng; Xin Fan Guo; Zhiqiang Hao; Fan Yang; Xingzhi Qian; Jiawei Du; Jinfeng Xu; Zheng Xing; Shuo Yang; Xingjun Wang - [57%] Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction
Shuze Liu; Kaixiang Zhao; Runyang Xu; Jingzhi Chen; Nathan Wu; Yu Wang; Yushun Dong - [57%] FoSeRL: Formal Sequential Robustness Certification for Reinforcement Learning Policies
Sara Taheri; Deep Kumar Ganguly; Jan KÅetÃnský; Majid Zamani - [57%] Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security
Shaswata Mitra; Raj Patel; Subash Neupane; Sudip Mittal; Md Rayhanur Rahman; Shahram Rahimi - [56%] Don't Waste the Noise: Importance-Guided Perturbation Allocation under Joint Global and Local Constraints
Melika Shirian; Kianoosh Vadaei - [56%] Jailbreaking Open-Weight LLMs via Random Embedding Perturbations
Abhinav Sudhakar Dubey; Scott Sirri; Vaggos Chatziafratis; C. Seshadhri - [56%] Speedbumps: Rejection Attacks on Speculative Decoding
Adam Y. J. Jones; Yu Yuan; Sergio Maffeis - [55%] Localisation-Aware Uncertainty for Pretrained Object Detection
Charmaine Barker; Daniel Bethell; Simos Gerasimou - [55%] SkillPoison: Progressive Skill Poisoning via Successful Experiences
Lizhi Zhang; Xin He; Dianxuan Fu; Yuyuan Feng; Jiatong Li; Qi Wang; Xin Wang; Qinggang Zhang - [54%] Robust Evidential Learning Through Latent Consistency
Charmaine Barker; Daniel Bethell; Simos Gerasimou - [53%] AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models
Pooneh Mousavi; Amir Ivry; Mirco Ravanelli; Cem Subakan - [53%] Adversarial Images Hijack Web Agents from Visual Grounding to Browser Execution
Wanjing Han; Levi Taiji Li; Mu Zhang; Yue Jiang; Guanhong Tao - [52%] Surviving the Router: Optimizing Skill Injections for Retrieval and Execution
Haneen Najjar; Luca Scionis; Haritz Puerto; Sahar Abdelnabi - [52%] BARE-AI: Bit-Flip Attack Resilience in AI Hardware through Built-in Performance Monitors
Habibur Rahaman; Swastik Bhattacharya; Sanjay Das; Kanad Basu; Swarup Bhunia - [52%] LTBD: Learnable Trust-Boundary Delimiters for Prompt Injection Defense
Luman Zhao; Minghui Xu; Yue Zhang; Yijun Yang - [52%] RobustLDS: Learning linear dynamical systems under adversarial corruptions
Aravinda Kanchana Ruwanpathirana; Hemant Tyagi - [51%] Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Files
Yupu Wang; Zhengyuan Jiang; Reachal Wang; Neil Zhenqiang Gong - [50%] Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning
Jungseob Lee; Dongyub Jude Lee; Sugyeong Eo; Seongtae Hong; Seungyoon Lee; Heuiseok Lim - [50%] Out of Sync, Out of Sight: Phantom State Attacks against IIoT Intrusion Detection
Sabrine Ennaji; Elhadj Benkhelifa; Nadia Kabachi - [50%] Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directions
Bojun Yang; Haochen Zhou; Zhifang Zhang; Haobo Wang; Songze Li; Lei Feng - [50%] Safe at One Loop, Risky at Another: Aligning Safety Across Recurrent Depths in Looped Language Models
Yi Wang; Xiuyuan Qi; Dongqi Han; Dongsheng Li; Wenjie Wang - [48%] Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
Md Sazid Uddin; Md. Khairul Alam Mazumder; M. F. Mridha - [48%] Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Models
Luze Sun; Cristina Nita-Rotaru; Alina Oprea - [48%] Hidden Risks of Jev: An Empirical Study of Security, Privacy, and Dual Use
Shang Wang; Tianqing Zhu; Huajie Chen; Jiayang Li; Meng Yang; Bo Liu - [47%] Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks
Neeraj Karamchandani; Piyush Nagasubramaniam; Xinhong Xie; Sencun Zhu; Dinghao Wu - [46%] PhaseAT: Fourier Phase Adversarial Training for Medical Image Domain Generalization
Ahmed Sharshar; Asif Hanif; Naveen Kumar Kummari; Mohammad Yaqub; Mohsen Guizan - [46%] Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Maria Carmen Jica; Ali Satvaty; Suzan Verberne; Fatih Turkmen - [45%] Bounded Reachability & Jailbreak Detection via Contraction-Constrained State Space Models
Omanshu Thapliyal - [45%] Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Qiusi Zhan; Nian Lyu; Stephanie Ding; Arnav Mehta; Xander Davies; Daniel Kang - [45%] Don't Let One Lie Survive A Hundred Truths: A Selective Bayesian Trust Estimator for Collaborative Perception
Yutong Liu; Chenyi Wang; Ming F. Li; Qingzhao Zhang - [45%] Dynamical low-rank equilibrium computation for stochastic games between advanced persistent threats and moving target defense
Tian Zijian; Zhang He; Chen Xinjie; Wang Wenhai; Liu Xinggao - [44%] Persona Guardrail: A Production-Grade Defense Framework for Agentic Systems
Bijeeta Pal; Sridhar Reddy Maddireddy; Muhaimin Bin Munir; Zoltan Puha; Max Zhurovich; Adi Raghavendra; Sean Tout - [44%] Backdooring Acoustic Foundation Models for Physically Realizable Triggers
Zebin Yun; Eyal Ronen; Mahmood Sharif - [43%] PerSpectron: Detecting Invariant Footprints of Microarchitectural Attacks with Perceptron
Samira Mirbagher-Ajorpaz; Gilles Pokam; Esmaeil Mohammadian-Koruyeh; Elba Garza; Nael Abu-Ghazaleh; Daniel A. Jiménez - [42%] Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
Haoyu Wang; Wei Zhao; Yedi Zhang; Christopher M. Poskitt; Jun Sun - [42%] Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge
Wonjun Lee; Kyungsik Yang; Gaeun Ji; Vaidehi Patil; Haon Park; Bumsub Ham; Mohit Bansal; Suhyun Kim - [42%] MARCO: The Radioactive Watermark for Protein Generative Models
Huajie Chen; Xin Guo; Yuchen Shi; Yuchen Zhong; Minhui Xue; Chi Liu; Congcong Zhu; Kun Gao; Minfeng Qi; Tianqing Zhu - [42%] Automotive Hardware Attacks: An Architect's Guide to TARA
Jakub Breier; Xiaolu Hou - [42%] How Hackable Is Your Speech Quality Metric? A Corrected Protocol, a Benchmark, and What Patching Buys
Ali Alavi; Donald S. Williamson - [41%] Backdoor Containment via Expert Quarantine and Shutdown in LLMs
Jianwei Li; Min-Seon Kim; Jung-Eun Kim - [41%] Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulation
Thanh Dong; An Ngo; Minh Dau; Rajesh Kumar - [41%] When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning
Shutong Zheng; Sijia Chen - [40%] Transferable Graph Metanetworks
Yuxin Ma; Adir Dayan; Yam Eitan; Haggai Maron; Soledad Villar - [40%] High-quality Data Do not Mean Safe! Poisoning LLMs after Data Selection
Kaiyang Li; Jiahao Chen; Yuwen Pu; Chunyi Zhou; Tong Zhang; Bin Cai; Chunqiang Hu; Haibo Hu - [40%] Mitigating Private Data Leakage in LLMs with Whiteout
Anna Yoo Jeong Ha; Ronik Bhaskar; Haitao Zheng; Ben Y. Zhao - [40%] SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populations
Vahid Tavakkoli; Kabeh Mohsenzadegan; Kyandoghere Kyamakya