| Author: | Liang, Kaisheng |
| Title: | Enhancing adversarial transferability across neural networks |
| Advisors: | Xiao, Bin (COMP) |
| Degree: | Ph.D. |
| Year: | 2026 |
| Department: | Department of Computing |
| Pages: | xix, 155 pages : color illustrations |
| Language: | English |
| Abstract: | Despite their widespread success in numerous applications, deep neural networks (DNNs) remain susceptible to adversarial attacks. These attacks involve imperceptible input modifications that can mislead models into producing erroneous outputs. A particular concern is the phenomenon known as adversarial transferability, where adversarial inputs crafted using one white-box model can deceive unknown black-box models. This transferability raises serious security concerns, especially in scenarios where model access is restricted but reliability is critical. Studying adversarial transferability is key to understanding and improving the robustness of DNNs. Numerous methods have been introduced to improve adversarial transferability, spanning both non-targeted and targeted attacks. Although non-targeted attacks have received more attention, many existing methods lack a deep understanding of the underlying mechanisms. In addition, they fail to transfer effectively across models with different architectures, primarily due to overfitting to the surrogate model. In contrast, targeted attacks seek to manipulate model predictions toward a specific class, posing a much greater challenge due to their stricter success criteria. Current targeted attacks are relatively limited in number and suffer from high computational cost and low success rates. In this thesis, we first present a novel perspective on the underlying mechanism of adversarial transferability. We observe that the non-robust style features of input images can hinder transferability by inducing overfitting to the surrogate model. Motivated by this insight, we propose Style-less Perturbation (StyLess), which suppresses non-robust style features by attacking a set of stylized surrogate models. These stylized models are constructed by injecting adaptive instance normalization layers into the surrogate and perturbing intermediate features in a training-free manner. StyLess enhances transferability by replacing the original surrogate model with stylized variants that reduce overfitting to the original image style. Beyond non-targeted attacks, we address the more challenging task of generating targeted adversarial examples. We introduce Feature Tuning Mixup (FTM), a novel method that augments the feature space instead of the input space. FTM introduces learnable attack-specific perturbations across multiple layers and employs a momentum-based stochastic update strategy to iteratively tune these feature perturbations. Empirical results show that FTM achieves state-of-the-art adversarial transferability with high efficiency in targeted attacks. We further explore two general strategies to boost adversarial transferability. The first, Soft Dropout Adaptation (SODA), introduces structured, learnable randomness in the input space to enhance adversarial transferability. Our analysis shows that a patch-based soft mask outperforms traditional random dropout. Based on this insight, SODA jointly adapts internal image content and surrounding context through a soft mask and a visual prompt. The second, Perspective-Invariant Attack (PIA), addresses the limitations of fixed perspective in transformations-based attacks. By applying randomized perspective transformations during the attack process, PIA improves the transferability of adversarial examples across models and input variations. Overall, this thesis contributes both conceptual insights and practical techniques for enhancing adversarial transferability in deep learning. |
| Rights: | All rights reserved |
| Access: | open access |
Copyright Undertaking
As a bona fide Library user, I declare that:
- I will abide by the rules and legal ordinances governing copyright regarding the use of the Database.
- I will use the Database for the purpose of my research or private study only and not for circulation or further reproduction or any other purpose.
- I agree to indemnify and hold the University harmless from and against any loss, damage, cost, liability or expenses arising from copyright infringement or unauthorized usage.
By downloading any item(s) listed above, you acknowledge that you have read and understood the copyright undertaking as stated above, and agree to be bound by all of its terms.
Please use this identifier to cite or link to this item:
https://theses.lib.polyu.edu.hk/handle/200/14473

