Cross-modal intent obfuscation
Rewrite the explicit query into a visually grounded proxy prompt and place the removed semantics in a composite image.
↗A study of how visual conditions are repeatedly re-applied during discrete diffusion generation, and a diffusion-aware framework for analyzing this vulnerability.

DIVA isolates a generation-process property that is easy to miss when visual attacks are designed for autoregressive models.
Discrete diffusion vision-language models (dVLMs) generate responses through iterative masked-token denoising. Unlike autoregressive models, they re-use the visual embedding at every reverse denoising step. We identify this repeated interaction as cross-step conditional propagation: adversarial visual semantics can be propagated and amplified across the generation trajectory.
DIVA (Discrete-diffusion Vision-language model Attack) combines cross-modal intent obfuscation with diffusion-aware multi-timestep optimization. The resulting framework targets the intermediate masked states instead of optimizing a single snapshot, providing a direct way to study safety behavior in dVLMs.
The method follows the model's own generation loop: move intent into the visual channel, observe how it re-enters the model, then optimize across the states that matter.
Rewrite the explicit query into a visually grounded proxy prompt and place the removed semantics in a composite image.
↗The visual embedding conditions masked-token prediction repeatedly during reverse denoising, creating a persistent pathway.
↗Sample multiple timesteps and backpropagate the aggregated masked cross-entropy to learn one bounded perturbation.
↗The figures below summarize the complete pipeline and the architectural difference that motivates it.



Paper-reported aggregate results are presented as responsive HTML tables. DIVA rows are highlighted for quick comparison.
| Model / method | Animal | Financial | Privacy | Self-harm | Violence | Average |
|---|---|---|---|---|---|---|
| LLaDA-V | ||||||
| LLaDA-V | 53.33 | 60.00 | 45.33 | 34.67 | 76.67 | 54.00 |
| + DIJA | 67.33 | 55.33 | 47.33 | 34.00 | 76.00 | 56.00 |
| + PAD | 69.33 | 54.00 | 50.00 | 39.33 | 65.33 | 55.60 |
| + DIVA (ours) | 61.33 | 66.00 | 39.33 | 50.00 | 77.33 | 58.80 |
| MMaDA | ||||||
| MMaDA | 32.00 | 32.00 | 30.67 | 18.67 | 40.67 | 30.80 |
| + DIJA | 70.67 | 66.00 | 56.67 | 39.33 | 72.67 | 61.07 |
| + PAD | 53.33 | 64.00 | 64.67 | 32.67 | 70.67 | 57.07 |
| + DIVA (ours) | 70.67 | 71.33 | 72.00 | 45.33 | 79.33 | 67.73 |
| LLaDA2.0-Uni | ||||||
| LLaDA2.0-Uni | 14.00 | 37.33 | 23.33 | 6.00 | 34.67 | 23.07 |
| + DIJA | 55.33 | 36.67 | 58.00 | 30.67 | 68.67 | 49.87 |
| + PAD | 47.33 | 53.33 | 61.33 | 30.00 | 76.67 | 53.73 |
| + DIVA (ours) | 83.33 | 81.33 | 70.00 | 36.00 | 74.67 | 69.07 |
Values transcribed from Table 1 in the paper. Higher ASR indicates more unsafe completions under the Beaver metric.
| Model / method | Safe rate ↓ | Safe refuse ↓ | Safe warning ↓ | ASR ↑ |
|---|---|---|---|---|
| LLaDA-V | ||||
| Plain text-only | 49.53 | 0.00 | 49.53 | 50.47 |
| LLaDA-V | 50.42 | 0.00 | 50.42 | 49.58 |
| + DIJA | 49.64 | 0.05 | 49.59 | 50.36 |
| + PAD | 48.94 | 0.05 | 48.89 | 51.06 |
| + DIVA (ours) | 47.93 | 0.00 | 47.93 | 52.07 |
| MMaDA | ||||
| Plain text-only | 27.93 | 0.00 | 27.93 | 72.07 |
| MMaDA | 29.90 | 0.04 | 29.85 | 70.10 |
| + DIJA | 26.95 | 0.00 | 26.95 | 73.05 |
| + PAD | 27.31 | 0.49 | 26.82 | 72.69 |
| + DIVA (ours) | 25.84 | 0.04 | 25.79 | 74.16 |
Values transcribed from Table 3. Safe-rate columns are lower-is-better; ASR is higher-is-better.
The ablation isolates the contribution of the composite carrier and the optimized perturbation. Bars are normalized to the DIVA result.
These paper figures show the before/after response comparison and the visual footprint of the learned perturbation.



Qualitative figures are included for research documentation. Please verify redistribution rights before mirroring the release publicly.
The curated release keeps the three model adapters, paper figures, and portable checks together in one folder.
Each adapter follows the target model's official preprocessing and generation path.
diva_lladav/LLaDA-V / LaViDadiva_mmada/MMaDA MMUdiva_llada_uni/LLaDA2.0-Uni$ python -m venv .venv
$ python -m pip install -r requirements.txt
$ python scripts/check_release.py
# portable release checks
Camera-ready paper PDF and the complete figure archive are included.
Accepted to the EMNLP 2026 Main Conference. Add the ACL Anthology URL and page numbers when the proceedings record is published.
@inproceedings{song-etal-2026-diva,
title = {DIVA: Exploiting Cross-Step Conditional Propagation for Visual Jailbreaks in Discrete Diffusion Vision-Language Models},
author = {Song, Guorui and Tang, Runqing and Zhang, Jingye and Zhang, Luyuan and Huang, Feice and Ray, Cong and Wang, Guocun and Zhong, Dake and Wai, Choo Sin and Dai, Bingquan and Wang, Chuming and Lin, Tongxu and Guo, Wanyu and Wang, Haoqian},
booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year = {2026},
publisher = {Association for Computational Linguistics},
note = {Accepted to the EMNLP 2026 Main Conference}
}