Robustness of Mask-Conditioned Image Editors Under Ambiguous or Adversarial Prompts

Authors
  • Suman Thapa

    Department of Computer Science, Pokhara University, Dhungepatan Road, Pokhara 33700, Nepal
    Author
  • Kiran Shahi

    Department of Information Technology, Nepal Open University, Manbhawan Road, Lalitpur 44700, Nepal
    Author
Abstract

Mask-conditioned image editors based on large generative models are widely used to manipulate local image regions while preserving unedited content. In these systems, a user provides an input image, a binary mask, and a natural language prompt that describes the intended modification inside the masked region. Despite their practical relevance, the behavior of such editors under ambiguous or adversarial prompts has not been characterized in a systematic way. Ambiguity naturally arises when language descriptions are under-specified, conflicting, or stylistically overloaded, while adversarial prompts are crafted to cause mislocalization of edits, content leakage beyond the mask, or unexpected semantic transformations. Understanding the robustness of masked editors to these classes of prompts is important for assessing reliability, safety, and predictability in interactive workflows. This paper examines robustness properties of generic mask-conditioned image editors by formulating a mathematical model of the editing operator, defining explicit robustness metrics over masked and unmasked regions, and analyzing the response to both ambiguous and adversarial prompts. The study considers discrete masks, continuous-valued mask relaxations, and different ways of injecting prompt information into the generative pipeline. A probabilistic framework is introduced to capture distributions over prompts, masks, and images, allowing robustness to be treated as an expected or worst-case risk. Several mitigation strategies are discussed, including regularization of cross-mask interactions, robust optimization over prompt neighborhoods, and training objectives that penalize unintended changes outside the mask. The analysis highlights characteristic failure modes and trade-offs between edit strength, semantic controllability, and spatial robustness without assuming any specific model architecture.

Downloads
Published
2025-09-04
Section
Articles
License

Copyright (c) 2025 authors

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.

How to Cite

Thapa, Suman, and Kiran Shahi. 2025. “Robustness of Mask-Conditioned Image Editors Under Ambiguous or Adversarial Prompts”. Transactions on General Science, Evidence Synthesis, and Interdisciplinary Methods 15 (9): 1-14. https://grovesocieties.com/index.php/TGSESIM/article/view/Robustness-of-Mask-Conditioned.