Pruning Backdoors via Inducive Trigger-Based Detection

Abstract

Backdoor attacks pose a serious threat to the security of deep neural networks by embedding hidden malicious behaviors that are triggered only by specific inputs. This paper presents a novel defense framework that detects and mitigates such attacks without requiring access to poisoned data or prior knowledge of the trigger. The proposed method introduces localized perturbations to clean samples to induce decoy behaviors, enabling the identification of infected models through output consistency analysis. Suspicious neurons are then identified using a Backdoor Loss Change (BLC) and Clean Loss Change (CLC) scoring mechanism, allowing fine-grained pruning of backdoor-relevant neurons while preserving those critical to clean performance. Experiments conducted on CIFAR-10 under various backdoor scenarios, including BadNet and Trojan attacks, demonstrate that the proposed approach significantly reduces attack success rates while maintaining high classification accuracy. Post-pruning finetuning further improves performance, establishing the framework as an effective and practical defense for real-world applications.

Publication
2025 International Conference on Wavelet Analysis and Pattern Recognition (ICWAPR), 23-28

Lab Members

Pan Jiahao 潘嘉灏
Pan Jiahao 潘嘉灏

My research interests include distributed robotics, mobile computing and programmable matter.

Patrick Chan
Patrick Chan
Associate Professor, Vice Dean

Patrick Chan works on machine learning, deep learning, image processing, adversarial learning, and secure machine learning.