arXiv cs.AIOctober 7, 2026
Harmful Content Generation in Text-to-Image Models: Capabilities and Moderation Limitations
Excerpt
arXiv:2610.06503v1 Announce Type: cross Abstract: Text-to-image generative models can produce highly realistic imagery but also raise concerns about harmful misuse. While safety mechanisms exist, systematic evaluations of their effectiveness against realistic attacks remain limited. We present a systematic evaluation of harmful content generation across five open text-to-image models using an automated pipeline that transforms legitimate news captions into unsafe prompts targeting sexually expli