conference paper

On Enhancing Adversarial Robustness of Large Pre-trained Vision-Language Models

Abstract

Large pre-trained vision-language models (VLMs), such as CLIP, have demonstrated significant potential in acquiring transferable representations for various downstream tasks. These models have the ability to comprehend visual information, making them valuable for applications in radar, electronic warfare, and cognitive warfare domains. By integrating VLMs into these frameworks, enhanced visual perception and cognitive capabilities can be achieved. However, it is important to address the vulnerability of VLMs to adversarial attacks originating from the visual modality, which can compromise their robustness. In this study, we propose an adversarial fine-tuning technique with a self-distillation mechanism to improve the visual robustness of VLMs while preserving their acquired representations. Our experimental results on various tasks, including zero-shot image classification, image captioning, and ScienceQA, demonstrate the effectiveness of our approach in enhancing the visual robustness of VLMs against adversarial perturbations. © 2024 Copyright held by the owner/author(s).

Author keywords

adversarial fine-tuning; clip models; vision-language models