TY - GEN
T1 - Deep Learning Backdoor Defense via Adaptive Trigger Collisions in Latent Space
AU - Xiong, Zixun
AU - Wang, Hao
AU - Li, Jian
AU - Hua, Yang
AU - Pan, Miao
AU - Du, Xiaojiang
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/4
Y1 - 2026/6/4
N2 - Backdoor attacks in data outsourcing settings pose severe risks to deep neural networks. Specifically, adversaries can manipulate externally sourced training data to implant hidden behaviors in target models (e.g., incorrect predictions on triggered samples). Existing defenses are either pre-processing or post-processing. Since the two approaches are orthogonal and either one can independently strengthen real-world defenses, we focus on the latter in this paper. Yet current post-processing defenses face one or more of the following issues: overemphasis on output logits while overlooking rich information in intermediate layers, injection of uncertain new triggers while requiring alignment with the original triggers, and underuse of poisoned model representations. To overcome the aforementioned limitations, we propose ATClean, an adaptive post-processing defense based on feature collisions in latent space. Specifically, it leverages all layers rather than only output logits to capture backdoor-affected regions using an adaptive loss function, relaxes the need for exact trigger reconstruction by generating adversarial samples that only enforce feature collisions with a theoretical guarantee, and fully exploits poisoned representations with feature-collision-based fine-tuning. Experiments across benchmark datasets, multiple architectures, and seven representative attacks show that ATClean achieves state-of-the-art defense effectiveness with the lowest drop on clean data, including about a 20% improvement in DER, which measures the accuracy-defense trade-off.
AB - Backdoor attacks in data outsourcing settings pose severe risks to deep neural networks. Specifically, adversaries can manipulate externally sourced training data to implant hidden behaviors in target models (e.g., incorrect predictions on triggered samples). Existing defenses are either pre-processing or post-processing. Since the two approaches are orthogonal and either one can independently strengthen real-world defenses, we focus on the latter in this paper. Yet current post-processing defenses face one or more of the following issues: overemphasis on output logits while overlooking rich information in intermediate layers, injection of uncertain new triggers while requiring alignment with the original triggers, and underuse of poisoned model representations. To overcome the aforementioned limitations, we propose ATClean, an adaptive post-processing defense based on feature collisions in latent space. Specifically, it leverages all layers rather than only output logits to capture backdoor-affected regions using an adaptive loss function, relaxes the need for exact trigger reconstruction by generating adversarial samples that only enforce feature collisions with a theoretical guarantee, and fully exploits poisoned representations with feature-collision-based fine-tuning. Experiments across benchmark datasets, multiple architectures, and seven representative attacks show that ATClean achieves state-of-the-art defense effectiveness with the lowest drop on clean data, including about a 20% improvement in DER, which measures the accuracy-defense trade-off.
KW - Backdoor defense
KW - Trustworthy AI
UR - https://www.scopus.com/pages/publications/105042452535
UR - https://www.scopus.com/pages/publications/105042452535#tab=citedBy
U2 - 10.1145/3779208.3806081
DO - 10.1145/3779208.3806081
M3 - Conference contribution
AN - SCOPUS:105042452535
T3 - ASIA CCS 2026 - Proceedings of the 21st ACM ASIA Conference on Computer and Communications Security
SP - 591
EP - 607
BT - ASIA CCS 2026 - Proceedings of the 21st ACM ASIA Conference on Computer and Communications Security
T2 - 21st ACM Asia Conference on Computer and Communications Security, AsiaCCS 2026
Y2 - 1 June 2026 through 5 June 2026
ER -