SCAM! Transferring humans between images with Semantic Cross Attention Modulation

Dufour, Nicolas; Picard, David; Kalogeiton, Vicky

Computer Science > Computer Vision and Pattern Recognition

arXiv:2210.04883 (cs)

[Submitted on 10 Oct 2022]

Title:SCAM! Transferring humans between images with Semantic Cross Attention Modulation

Authors:Nicolas Dufour, David Picard, Vicky Kalogeiton

View PDF

Abstract:A large body of recent work targets semantically conditioned image generation. Most such methods focus on the narrower task of pose transfer and ignore the more challenging task of subject transfer that consists in not only transferring the pose but also the appearance and background. In this work, we introduce SCAM (Semantic Cross Attention Modulation), a system that encodes rich and diverse information in each semantic region of the image (including foreground and background), thus achieving precise generation with emphasis on fine details. This is enabled by the Semantic Attention Transformer Encoder that extracts multiple latent vectors for each semantic region, and the corresponding generator that exploits these multiple latents by using semantic cross attention modulation. It is trained only using a reconstruction setup, while subject transfer is performed at test time. Our analysis shows that our proposed architecture is successful at encoding the diversity of appearance in each semantic region. Extensive experiments on the iDesigner and CelebAMask-HD datasets show that SCAM outperforms SEAN and SPADE; moreover, it sets the new state of the art on subject transfer.

Comments:	Accepted at ECCV 2022
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2210.04883 [cs.CV]
	(or arXiv:2210.04883v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2210.04883

Submission history

From: Nicolas Dufour [view email]
[v1] Mon, 10 Oct 2022 17:54:47 UTC (17,962 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:SCAM! Transferring humans between images with Semantic Cross Attention Modulation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:SCAM! Transferring humans between images with Semantic Cross Attention Modulation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators