Correlating Adversarial Robustness with the Energy of Untrained Classifiers

Abstract

Neural networks are known to be vulnerable to imperceptible perturbations to their inputs, causing unwarranted behavior in the model’s predictions and attracting the scientific community’s interest with the aim of designing architectures and training methods that are robust to such attacks. By reinterpreting a discriminant classifier as an energy-based model, the following work further studies the connection between robustness, architectural configuration, and energy landscape. Energy is leveraged as a tool to study a network’s intrinsic robustness by analyzing models with randomized weights, suggesting a connection between a classifier’s untrained energy landscape and its final robustness after adversarial training. Notably, the energy profile of an untrained network appears to align with its final robustness metric. This correspondence holds true across three distinct model scales, each encompassing a diverse range of architectural configurations, suggesting a universal link independent of specific structural choices.

Publication
International Conference on Machine Learning and Cybernetics (ICMLC)
Mirza Mujtaba Hussain
Mirza Mujtaba Hussain
PhD Student

Hi there! 👋 I’m Hussain, a Ph.D. student at Sapienza University. Currently I’m diving into Adversarial Machine Learning and Explainable AI to find practical solutions for real-world challenges. My goal is to use AI to make a positive impact on our society.

Iacopo Masi
Iacopo Masi
Associate Professor (PI)

My research interests include computer vision, biometrics, AI.