Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

sbulaev2 pts0 comments

[2608.07430] Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Skip to main content

Search arXiv

Press Enter to search · Advanced search

-->

Computer Science > Machine Learning

arXiv:2608.07430 (cs)

[Submitted on 7 Aug 2026]

Title:Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Authors:Elena Dumitrescu, Gert Lek, Lydia Y. Chen, Jérémie Decouchant<br>View a PDF of the paper titled Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits, by Elena Dumitrescu and 3 other authors

View PDF<br>HTML (experimental)

Abstract:Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment.

We first show that safety alignment in DLLMs remains sparse and transferable across architectures. DLLMs initialized from autoregressive predecessors inherit the same mechanistic safety footprint as their source models, enabling transfer attacks via direct safety neuron mapping and pruning. Self-pruning increases attack success rates (ASR) from 2.6% to 73.8% on LLaDA and from 1.9% to 86.6% on Dream, while transfer pruning from Qwen2.5 increases ASR from 1.9% to 73.2% on Dream and from 7.0% to 86.3% on Fast-dLLM.

Building on these findings, we introduce SN-Guided Diffusion, a fully offline black-box jailbreak framework that steers the diffusion process away from safety-triggering regions using a weighted safety neuron loss, which achieves near-perfect prompt separability (AUROC = 1.0 for benign-vs-jailbreak discrimination). Across multiple open and proprietary targets, our method achieves a transfer ASR of up to 77.1% on Llama-3-8B-Instruct, 86.9% on Qwen2.5-7B-Instruct, and 74.3% against Gemini-2.5-Flash-Lite, while requiring only 20 generation episodes per prompt. Compared to prior jailbreaking frameworks, our method achieves competitive transferability with orders-of-magnitude lower generation cost.

Our codebase is available at this https URL.

Subjects:

Machine Learning (cs.LG); Artificial Intelligence (cs.AI)

Cite as:<br>arXiv:2608.07430 [cs.LG]

(or<br>arXiv:2608.07430v1 [cs.LG] for this version)

https://doi.org/10.48550/arXiv.2608.07430

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

Submission history<br>From: Jérémie Decouchant [view email]<br>[v1]<br>Fri, 7 Aug 2026 17:17:18 UTC (504 KB)

Full-text links:<br>Access Paper:

View a PDF of the paper titled Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits, by Elena Dumitrescu and 3 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source

view license

Current browse context:

cs.LG

next >

new<br>recent<br>| 2026-08

Change to browse by:

cs<br>cs.AI

References & Citations

NASA ADS<br>Google Scholar

Semantic Scholar

export BibTeX citation<br>Loading...

BibTeX formatted citation

&times;

loading...

Data provided by:

Bookmark

Bibliographic Tools

Bibliographic and Citation Tools

Bibliographic Explorer Toggle

Bibliographic Explorer (What is the Explorer?)

Connected Papers Toggle

Connected Papers (What is Connected Papers?)

Litmaps Toggle

Litmaps (What is Litmaps?)

scite.ai Toggle

scite Smart Citations (What are Smart Citations?)

Code, Data, Media

Code, Data and Media Associated with this Article

alphaXiv Toggle

alphaXiv (What is alphaXiv?)

Links to Code Toggle

CatalyzeX Code Finder for Papers (What is CatalyzeX?)

DagsHub Toggle

DagsHub (What is DagsHub?)

GotitPub Toggle

Gotit.pub (What is GotitPub?)

Huggingface Toggle

Hugging Face (What is Huggingface?)

ScienceCast Toggle

ScienceCast (What is ScienceCast?)

Demos

Demos

Replicate Toggle

Replicate (What is Replicate?)

Spaces Toggle

Hugging Face Spaces (What is Spaces?)

Spaces Toggle

TXYZ.AI (What is TXYZ.AI?)

Related Papers

Recommenders and Search Tools

Link to Influence Flower

Influence Flower (What are Influence Flowers?)

Core recommender toggle

CORE Recommender (What is CORE?)

IArxiv recommender toggle

IArxiv Recommender<br>(What is IArxiv?)

Author

Venue

Institution

Topic

About arXivLabs

arXivLabs: experimental projects with community collaborators

arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.

Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.

Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs .

Which authors of this paper are endorsers? |<br>Disable MathJax (What is MathJax?)

Major funding support from

toggle safety diffusion arxiv from targets

Related Articles