Skip to main navigation Skip to search Skip to main content

Point2pix-Zero: Point-driven refined diffusion for multi-object image editing

    • University of Liverpool
    • University of Liverpool
    • Digital Innovation Research Center
    • Duke Kunshan University

    Research output: Contribution to journalArticlepeer-review

    6 Citations (Scopus)

    Abstract

    Semantic image editing methods employing large-scale diffusion models have made significant strides in precise and controlled image editing with text prompts as guidance. However, these models struggle to handle complex images containing hard-described objects and/or multiple objects. In this work, we introduce a novel inference-time multi-object image editing strategy, Point2pix-Zero, editing a single object with the simple guidance of clicked points and the text of target objects. We employ an interactive methodology, point-discovery, as text-free guidance to identify the semantic information of intended edited objects and generate text prompts automatically. Instead of exploiting internal cross-attention maps of diffusion models as a guide, we inject external attention maps to rectify the visual-and-semantic pairing mismatches in cross-attention maps during the denoising process. Extensive empirical evaluations demonstrate the effectiveness of our proposed inference-time method in ensuring precise editing while maintaining image fidelity. Our method showcases superior performance in single- and multi-object image editing, positioning it as a new state-of-the-art.

    Original languageEnglish
    Article number112041
    JournalPattern Recognition
    Volume170
    DOIs
    Publication statusPublished - Feb 2026

    Keywords

    • Diffusion model
    • Image editing

    Fingerprint

    Dive into the research topics of 'Point2pix-Zero: Point-driven refined diffusion for multi-object image editing'. Together they form a unique fingerprint.

    Cite this