Skip to main navigation Skip to search Skip to main content

Deep Learning Backdoor Defense via Adaptive Trigger Collisions in Latent Space

  • Stevens Institute of Technology
  • Stony Brook University
  • Queen's University Belfast
  • University of Houston

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Backdoor attacks in data outsourcing settings pose severe risks to deep neural networks. Specifically, adversaries can manipulate externally sourced training data to implant hidden behaviors in target models (e.g., incorrect predictions on triggered samples). Existing defenses are either pre-processing or post-processing. Since the two approaches are orthogonal and either one can independently strengthen real-world defenses, we focus on the latter in this paper. Yet current post-processing defenses face one or more of the following issues: overemphasis on output logits while overlooking rich information in intermediate layers, injection of uncertain new triggers while requiring alignment with the original triggers, and underuse of poisoned model representations. To overcome the aforementioned limitations, we propose ATClean, an adaptive post-processing defense based on feature collisions in latent space. Specifically, it leverages all layers rather than only output logits to capture backdoor-affected regions using an adaptive loss function, relaxes the need for exact trigger reconstruction by generating adversarial samples that only enforce feature collisions with a theoretical guarantee, and fully exploits poisoned representations with feature-collision-based fine-tuning. Experiments across benchmark datasets, multiple architectures, and seven representative attacks show that ATClean achieves state-of-the-art defense effectiveness with the lowest drop on clean data, including about a 20% improvement in DER, which measures the accuracy-defense trade-off.

Original languageEnglish
Title of host publicationASIA CCS 2026 - Proceedings of the 21st ACM ASIA Conference on Computer and Communications Security
Pages591-607
Number of pages17
ISBN (Electronic)9798400723568
DOIs
StatePublished - 4 Jun 2026
Event21st ACM Asia Conference on Computer and Communications Security, AsiaCCS 2026 - Bangalore, India
Duration: 1 Jun 20265 Jun 2026

Publication series

NameASIA CCS 2026 - Proceedings of the 21st ACM ASIA Conference on Computer and Communications Security

Conference

Conference21st ACM Asia Conference on Computer and Communications Security, AsiaCCS 2026
Country/TerritoryIndia
CityBangalore
Period1/06/265/06/26

Keywords

  • Backdoor defense
  • Trustworthy AI

Fingerprint

Dive into the research topics of 'Deep Learning Backdoor Defense via Adaptive Trigger Collisions in Latent Space'. Together they form a unique fingerprint.

Cite this