Name: JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models
Keywords: FOS: Computer and information sciences, Computer Science - Cryptography and Security, Computer Science - Computation and Language, Computer Science - Human-Computer Interaction, Cryptography and Security (cs.CR), Computation and Language (cs.CL), Human-Computer Interaction (cs.HC)

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 01 Oct 2025Embargo end date: 01 Jan 2024Publisher:Institute of Electrical and Electronics Engineers (IEEE)Journal:IEEE Transactions on Visualization and Computer Graphics, volume 31, pages 8,668-8,682 (issn: 1077-2626, eissn: 2160-9306,

Authors: Yingchaojie Feng; Zhizhang Chen; Zhining Kang; Sijia Wang; Haoyu Tian; Wei Zhang; Minfeng Zhu; +1 Authors

doi: 10.1109/tvcg.2025.3575694 , 10.48550/arxiv.2404.08793

arXiv: http://arxiv.org/abs/2404.08793

JailbreakLens: Visual Analysis of Jailbreak Attacks Against Large Language Models

- Summary
- Subjects
- Metrics

Abstract

The proliferation of large language models (LLMs) has underscored concerns regarding their security vulnerabilities, notably against jailbreak attacks, where adversaries design jailbreak prompts to circumvent safety mechanisms for potential misuse. Addressing these concerns necessitates a comprehensive analysis of jailbreak prompts to evaluate LLMs' defensive capabilities and identify potential weaknesses. However, the complexity of evaluating jailbreak performance and understanding prompt characteristics makes this analysis laborious. We collaborate with domain experts to characterize problems and propose an LLM-assisted framework to streamline the analysis process. It provides automatic jailbreak assessment to facilitate performance evaluation and support analysis of components and keywords in prompts. Based on the framework, we design JailbreakLens, a visual analysis system that enables users to explore the jailbreak performance against the target model, conduct multi-level analysis of prompt characteristics, and refine prompt instances to verify findings. Through a case study, technical evaluations, and expert interviews, we demonstrate our system's effectiveness in helping users evaluate model security and identify model weaknesses.

Related Organizations

Zhejiang Ocean University
China (People's Republic of)

Keywords

FOS: Computer and information sciences, Computer Science - Cryptography and Security, Computer Science - Computation and Language, Computer Science - Human-Computer Interaction, Cryptography and Security (cs.CR), Computation and Language (cs.CL), Human-Computer Interaction (cs.HC)

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	1
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Average

Green