NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
Jump to Discussion
Presentation
- Session
- Let's figure out how things work behind the scenes
- Time
- Wednesday, Nov 11, 13:48 – 14:00 (US/Eastern) · session 13:00 – 14:30
- Room
- Hall America north
- Presenting from
- Boston
Links
Sign in to access the Paper PDF.
- Preprint PDF(opens in new tab)

Abstract
Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging. We present NeuroBreak, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors. A layer-wise probing pipeline traces how harmful representations evolve across layers, while a dual-dimensional character--behavior categorization reveals each safety-related neuron's inherent tendency and contextual contribution. These analyses are made interpretable through tailored visualization designs: a task-driven probing projection that reveals safety decision boundaries, a dual-stream semantic evolution flow that traces cross-layer semantic shifts, and a character--behavior chord graph that unifies neuron roles, attribution scores, and collaborative relations in a single view with in-situ causal verification. Quantitative evaluations and case studies show that NeuroBreak uncovers safety failure causes and provides actionable insights for strengthening LLM defenses.
For Practitioners
Practitioners in LLM safety, model alignment, and mechanistic interpretability may find this work useful. NeuroBreak provides a workflow for diagnosing jailbreak failures by tracing harmful semantic changes across layers and examining the roles and interactions of safety-related neurons. Practitioners can use these analyses to identify vulnerable internal mechanisms, compare failure patterns across attack contexts, and guide more targeted model interventions and safety fine-tuning.