descriptionPublicationkeyboard_double_arrow_right Article , Preprint , Conference object , Research , Contribution for newspaper or weekly magazine 01 Jun 2022Embargo end date: 01 Jan 2021Publisher:IEEEJournal:2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Funded by:DFG | unidentified

Authors: Liao, Wentong; Hu, Kai; Yang, Michael Ying; Rosenhahn, Bodo;

doi: 10.1109/cvpr52688.2022.01765 , 10.48550/arxiv.2104.00567

arXiv: 2104.00567

Text to Image Generation with Semantic-Spatial Aware GAN

- Summary
- Subjects
- Related research
  (1)
- Metrics

Abstract

Text-to-image synthesis (T2I) aims to generate photo-realistic images which are semantically consistent with the text descriptions. Existing methods are usually built upon conditional generative adversarial networks (GANs) and initialize an image from noise with sentence embedding, and then refine the features with fine-grained word embedding iteratively. A close inspection of their generated images reveals a major limitation: even though the generated image holistically matches the description, individual image regions or parts of somethings are often not recognizable or consistent with words in the sentence, e.g. "a white crown". To address this problem, we propose a novel framework Semantic-Spatial Aware GAN for synthesizing images from input text. Concretely, we introduce a simple and effective Semantic-Spatial Aware block, which (1) learns semantic-adaptive transformation conditioned on text to effectively fuse text features and image features, and (2) learns a semantic mask in a weakly-supervised way that depends on the current text-image fusion process in order to guide the transformation spatially. Experiments on the challenging COCO and CUB bird datasets demonstrate the advantage of our method over the recent state-of-the-art approaches, regarding both visual fidelity and alignment with input text description.

arXiv admin note: text overlap with arXiv:1711.10485 by other authors

Related Organizations

View all View all

Keywords

FOS: Computer and information sciences, Vision + language, Computer Science - Machine Learning, Computer Vision and Pattern Recognition (cs.CV), cs.LG, Computer Science - Computer Vision and Pattern Recognition, Image and video synthesis and generation, /dk/atira/pure/subjectarea/asjc/1700/1712; name=Software, /dk/atira/pure/subjectarea/asjc/1700/1707; name=Computer Vision and Pattern Recognition, cs.CV, Machine Learning (cs.LG)

1 Research products, page 1 of 1

text2image software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	83
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 1%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Top 10%
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Top 1%