OmnAI Lab
OmnAI Lab
News
People
Research Group
Thesis
Publications
Projects
Photos
Contact
Light
Dark
Automatic
Arxiv-Preprint
Steering Vectors are an Adversarial Attack Surface
Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning. Since the technique is …
Abzal Aidakhmetov
,
Donato Crisostomi
,
Tommaso Mencattini
,
Robert Adrian Minut
,
Iacopo Masi
,
Emanuele Rodolà
PDF
Cite
Code
Not All Latent Spaces Are Flat: Hyperbolic Concept Control
As modern text-to-image (T2I) models draw closer to synthesizing highly realistic content, the threat of unsafe content generation …
Maria Rosaria Briglia
,
Simone Facchiano
,
Paolo Cursi
,
Alessio Sampieri
,
Emanuele Rodolà
,
Guido Maria D'Amely Di Melendugno
,
Luca Franco
,
Fabio Galasso
,
Iacopo Masi
PDF
Cite
Understanding Adversarial Training with Energy-based Models
We aim at using Energy-based Model (EBM) framework to better understand adversarial training (AT) in classifiers, and additionally to …
Mirza Mujtaba Hussain
,
Maria Rosaria Briglia
,
Filippo Bartolucci
,
Senad Beadini
,
Giuseppe Lisanti
,
Iacopo Masi
PDF
Cite
Evaluating the Robustness of Geometry-Aware Instance-Reweighted Adversarial Training
In this technical report, we evaluate the adversarial robustness of a very recent method called "Geometry-aware …
Dorjan Hitaj
,
Giulio Pagnotta
,
Iacopo Masi
,
Luigi v Mancini
Cite
Cite
×