Showing 1–2 of 2 results for author: Palaev, A

Search v0.5.6 released 2020-02-24

arXiv:2501.14046 [pdf, other]

cs.CV

LLM-guided Instance-level Image Manipulation with Diffusion U-Net Cross-Attention Maps

Authors: Andrey Palaev, Adil Khan, Syed M. Ahsan Kazmi

Abstract: The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance level. While existing methods offer some control through fine-tuning or auxiliary information, they often face limitations in flexibility and accuracy. To addres… ▽ More The advancement of text-to-image synthesis has introduced powerful generative models capable of creating realistic images from textual prompts. However, precise control over image attributes remains challenging, especially at the instance level. While existing methods offer some control through fine-tuning or auxiliary information, they often face limitations in flexibility and accuracy. To address these challenges, we propose a pipeline leveraging Large Language Models (LLMs), open-vocabulary detectors, cross-attention maps and intermediate activations of diffusion U-Net for instance-level image manipulation. Our method detects objects mentioned in the prompt and present in the generated image, enabling precise manipulation without extensive training or input masks. By incorporating cross-attention maps, our approach ensures coherence in manipulated images while controlling object positions. Our method enables precise manipulations at the instance level without fine-tuning or auxiliary information such as masks or bounding boxes. Code is available at https://github.com/Palandr123/DiffusionU-NetLLM △ Less

Submitted 23 January, 2025; originally announced January 2025.

Comments: Presented at BMVC 2024
arXiv:2305.14551 [pdf, other]

cs.CV cs.LG

Exploring Semantic Variations in GAN Latent Spaces via Matrix Factorization

Authors: Andrey Palaev, Rustam A. Lukmanov, Adil Khan

Abstract: Controlled data generation with GANs is desirable but challenging due to the nonlinearity and high dimensionality of their latent spaces. In this work, we explore image manipulations learned by GANSpace, a state-of-the-art method based on PCA. Through quantitative and qualitative assessments we show: (a) GANSpace produces a wide range of high-quality image manipulations, but they can be highly ent… ▽ More Controlled data generation with GANs is desirable but challenging due to the nonlinearity and high dimensionality of their latent spaces. In this work, we explore image manipulations learned by GANSpace, a state-of-the-art method based on PCA. Through quantitative and qualitative assessments we show: (a) GANSpace produces a wide range of high-quality image manipulations, but they can be highly entangled, limiting potential use cases; (b) Replacing PCA with ICA improves the quality and disentanglement of manipulations; (c) The quality of the generated images can be sensitive to the size of GANs, but regardless of their complexity, fundamental controlling directions can be observed in their latent spaces. △ Less

Submitted 23 May, 2023; originally announced May 2023.

Comments: Accepted at ICLR 2023 Tiny Papers

Search v0.5.6 released 2020-02-24