Modern vision–language models (VLMs) such as CLIP, BLIP-2, and LLaVA have shown remarkable capabilities in understanding images and text jointly. However, these multimodal AI systems face critical security vulnerabilities along two intertwined dimensions: robustness and privacy. In high-stakes applications like cyber-defence, such vulnerabilities could be disastrous. An adversary might subtly perturb an image to mislead an AI system, or extract private information from a model. This project is motivated by the need for trustworthy multimodal AI in security-sensitive contexts, with a focus on distributed AI systems, whether physical or logical, because these setups are often adopted to protect privacy. The main objectives of this projects are: evaluating the state-of-the-art VLMs performance under a variety of realistic threat scenarios by creating new types of attacks that address the two intertwined dimensions (robustness and privacy), exploring possible defences to strengthen the VLMs against found vulnerabilities and not downgrade the initial performance of the model (better trade-off between performance and security) and deploying a benchmark for VLMs, a framework of new found attacks to challenge VLMs.
Security of Vision-Language Models (CYD Doctoral Fellowship)
| Date | 01/05/2026 - 30/04/2030 |
| Type | Machine Learning |
| Partner | armasuisse |
| Partner contact | Ljiljana Dolamic |
| EPFL Laboratory | Distributed Computing Laboratory |